World Action Models Explained: Black Forest Labs Says Its Robot Brain Ranks First. NVIDIA’s Leaderboard Says Second

A world action model predicts robot actions and the video that follows. What FLUX 3 Action and Runway’s Praxis-1 are, why NVIDIA’s RoboLab leaderboard lists FLUX second at 42.9%, what the license allows and what the real-arm tests show.
White SO-101 robot arm gripping a white block on a cork tabletop, next to a grey container and a blue container holding dark blocks White SO-101 robot arm gripping a white block on a cork tabletop, next to a grey container and a blue container holding dark blocks
An SO-101 robot arm controlled by Black Forest Labs' FLUX 3 Action picks up a white block, in a still from a demo video the company published with the model. Image: Black Forest Labs.

Black Forest Labs says its new robot model is number one on NVIDIA’s RoboLab benchmark. NVIDIA’s own leaderboard, when we checked it on October 11, lists another entry ahead of it.

That’s the short version of where world action models stand. They’re the newest challenger to the vision-language-action (VLA) models that run most of today’s robot demos, and they’re arriving as open weights you can download. Black Forest Labs, the company behind the FLUX image generators, released FLUX 3 Action on September 22, and Runway has announced a rival called Praxis-1. Here’s what a world action model is, what the numbers do and don’t show, and what you’d need to run one.

Number one, then number two

Black Forest Labs says FLUX 3 Action, a 7-billion-parameter model, “places first” on RoboLab-120 with 42.92% task success, ahead of NVIDIA’s 16-billion-parameter Cosmos 3 Nano policy (36.8%) and Physical Intelligence’s π0.5 (28.0%). The DeepLearning.AI newsletter The Batch repeated the ranking on October 9.

Advertisement

NVIDIA’s RoboLab leaderboard confirms the 42.9% (515 of 1,200 trials), but it lists FLUX 3 Action second. First place goes to an entry called “Lakeside,” labeled LWM, at 51.0% (1,531 of 3,000 trials, 5 billion parameters). We couldn’t find a public write-up of that entry. The leaderboard also attaches a margin of about 5.7 percentage points to the FLUX score, which means the next three entries, at 39.0% to 39.9%, sit inside it.

RoboLab is also a simulation. It’s 120 tabletop tasks in NVIDIA’s Isaac Sim, ten trials each, on a simulated Franka arm. NVIDIA says its rankings track real-world results on a separate benchmark (a rank correlation of 0.94), which is encouraging, not proof about any particular robot in your building.

What a world action model does differently

A VLA looks at camera images and an instruction and outputs actions. Physical Intelligence’s π0.5 and NVIDIA’s GR00T work that way (see our explainer on VLAs). A world action model, or WAM, predicts the actions and the video of what should happen next, together, in a single pass. Black Forest Labs says the term comes from a system called DreamZero.

The bet is that learning from enormous amounts of video teaches a model how objects move and what a task looks like halfway through, so it needs fewer robot demonstrations. Black Forest Labs says video made up more than 95% of FLUX 3 Action’s pretraining tokens. Robot data came later: its report lists a mix that includes game recordings, human hands filmed from the first-person view, handheld grippers and teleoperation from 14 robot types. (For why robot data is the bottleneck, see our training-data explainer.)

Imagining the future costs compute

The catch is speed and hardware. Black Forest Labs says Cosmos 3 Nano needs about 4.7 times as much processing time per second of robot motion as π0.5 on an NVIDIA B200. Its own model closes much of that gap. In the company’s tests the single-step version runs 1.34 to 2.28 times faster than π0.5 on workstation and datacenter GPUs, though it’s slower than π0.5 on a consumer RTX 5090, by 1.15 to 2.42 times depending on precision. It predicts about 2.1 seconds of motion per call, against 1.0 second for π0.5. The speed has a price in accuracy, by the company’s own report: the single-step version scores 38.3% on RoboLab, and the 42.2% to 42.9% results come from slower setups.

The RoboLab leaderboard lists FLUX 3 Action at 69 GB of video memory, against “more than 8 GB” for π0.5. That’s a datacenter-class card, not something you tuck inside a robot, which is why budget robots such as the Flourish 1 lean on the cloud (see our Flourish 1 price report).

What you can actually download

The weights are on Hugging Face, along with ready-made versions fine-tuned for a Franka arm (DROID) and the low-cost SO-101 arm, both wired into Hugging Face’s LeRobot toolkit. The base checkpoint is not a working policy: its model card calls it “an adaptation component, not a complete robot policy,” and says a new robot needs its own action heads. Black Forest Labs says it adapted the model to an SO-101 with about 200 teleoperated demonstrations.

The license is the FLUX Kommunity License. It’s free for research and other non-commercial use. Commercial use is allowed only for a “Qualifying User,” meaning a company whose gross annualized revenue, with its affiliates, is under US$5 million. Everyone else has to ask Black Forest Labs for a commercial license. The model card also carries a warning every robotics team should read: “Nothing in the model bounds joint velocity, force or workspace; the application must enforce those limits and keep a hardware stop within reach.”

On real arms, the samples are small

Black Forest Labs reports that Positronic Robotics, which it describes as an independent third party, ran four policies on a Franka arm: ten tasks, three attempts each. FLUX 3 Action completed 28 of 30, Cosmos 3 Nano 27, DreamZero 20 and π0.5 13. With 30 attempts per model, one extra success separates the top two. It’s a useful signal that the simulation lead isn’t only a simulation effect, but it isn’t a verdict.

The company also ran a simulated hybrid in which a reasoning model, GPT 6 Astra, supervises FLUX 3 Action. By Black Forest Labs’ numbers it solved 90% of episodes at $8.77 and about eight minutes per success, against $13.47 and 16 minutes for Astra alone, which succeeded on all 50. These are the company’s own runs.

Runway, and a field that moves weekly

Runway’s Praxis-1 is billed as its first open-weight world action model. Runway says Noble Machines, Standard Bots and Ultra are testing it on their own hardware and that it will release the model publicly “in the coming months.” Runway says simulating robot policies inside its world model predicts real-world results with a 0.95 correlation. The page we read didn’t give a task-success rate, so there’s nothing yet to compare with FLUX 3 Action.

Academic work is moving in the same direction. A paper posted in October, UNITAS, describes a 1.7-billion-parameter model that works in 3D space and reports 99.8% on the LIBERO benchmark and an 85% average across real-world tasks (we read the abstract only; the code is on GitHub). Another, Being-M0.7, applies the idea to humanoids and says it matches the strongest baseline on real Unitree G1 tasks. Unitree’s open release, UnifoLM-WLA-1.0, is a VLA-style design that also learns to predict what changes in the scene next. Unitree hasn’t published its success rate, so there’s no head-to-head with any of these.

What to watch

No customer humanoid is known to run a world action model today. The tested hardware is a lab arm, a hobby arm, a drone and video games. The things to watch are the Praxis-1 release, a RoboLab leaderboard that now has at least one entry ahead of the open leader, and the first WAM that a humanoid maker says it has put on a robot working in the field.

Frequently asked questions

What is a world action model?

A world action model, or WAM, is a robot-control model that predicts future video frames and the robot’s next actions together. Black Forest Labs, Runway and NVIDIA’s Cosmos policies are all built this way, using video pretraining to reduce the robot data they need.

How is a world action model different from a VLA?

A VLA maps camera images and an instruction straight to actions. A WAM also predicts what the cameras should see next. That can help it generalize, but Black Forest Labs says it makes the models slower and bigger unless they’re distilled.

Is FLUX 3 Action free to use commercially?

Only for small companies. The FLUX Kommunity License allows commercial use for users whose gross annualized revenue, with affiliates, is under US$5 million. Larger companies need a separate license from Black Forest Labs.

Is FLUX 3 Action the best robot model?

Not by NVIDIA’s RoboLab leaderboard on October 11, which lists it second at 42.9% behind an entry called Lakeside at 51.0%. Black Forest Labs calls it the best open-weight model on the benchmark, and we couldn’t confirm whether Lakeside’s weights are public.

Can FLUX 3 Action run on a humanoid robot?

Black Forest Labs has shown it on a Franka arm, an SO-101 arm, a simulated drone and two video games. It hasn’t announced a humanoid, and a new body needs its own action heads and fine-tuning data.

Sources

All accessed October 11, 2026.

  • Black Forest Labs, “FLUX 3 Action: A 7B World Action Model for Robot Control” (research page, September 2026): bfl.ai
  • Black Forest Labs, “FLUX 3 Action: a world action model you can fine-tune,” Hugging Face blog (September 23, 2026): huggingface.co
  • Hugging Face model cards and license, black-forest-labs/flux-3-action-base and flux-3-action-droid (repositories created September 22, 2026): flux-3-action-base
  • NVIDIA, RoboLab-120 leaderboard: research.nvidia.com
  • DeepLearning.AI, “Black Forest Labs’ Flux 3 Action Applies Video Expertise to a World-Action Model,” The Batch (October 9, 2026): deeplearning.ai
  • Runway, “Introducing Praxis-1” (research page, September 2026): runway.com
  • Wang et al., “UNITAS: A 3D-Native World Action Model for Embodied Manipulation,” arXiv 2610.12099 (abstract): arxiv.org
  • Yue et al., “Being-M0.7: A Latent World-Action Model for Humanoid Robots,” arXiv 2610.11283 (abstract): arxiv.org

Related: What is a robot foundation model? · NVIDIA GR00T explained · What is Physical Intelligence’s π0?

Last updated: October 11, 2026. To report an error, see our corrections page. Articles are drafted with AI assistance and reviewed and edited by an editor; see our editorial policy.

Keep Up to Date with the Most Important News

By pressing the Subscribe button, you confirm that you have read and are agreeing to our Privacy Policy and Terms of Use
Advertisement