When Physical Intelligence released π0 (pronounced “pi-zero”), it did something most robot-AI companies won’t: it put the model’s code and weights on GitHub. That decision is a big part of why π0 became one of the most-cited open robot models of 2025 and 2026. It’s also why the name now turns up in plenty of places where it doesn’t belong, including vague claims that a robot “runs π0.”
So what is Physical Intelligence’s π0, exactly? It’s a generalist vision-language-action (VLA) model. It takes camera images, a language instruction and the robot’s own state, and outputs continuous robot actions. It was trained across many different robots, not just humanoids. It’s a research and fine-tuning starting point, not a finished brain for a commercial humanoid. Here’s how it works, what the family looks like now, and what it isn’t.
How π0 turns a picture and a sentence into motion
π0 starts from a pretrained vision-language model, the kind of AI that already understands images and text from the web. Physical Intelligence then bolted on what it calls an “action expert.” This smaller network uses a technique called flow matching to generate smooth, continuous actions instead of the coarse text tokens earlier VLAs used. According to the π0 paper, the result is about 3.3 billion parameters: a 3-billion-parameter language model plus a 300-million-parameter action expert. It outputs chunks of 50 future actions at a time.
That design is built for speed. Company materials describe control at up to about 50 Hz on dexterous tasks, fast enough for fiddly, contact-rich manipulation. And the training data is deliberately mixed. The paper describes pretraining on more than 10,000 hours of robot data from single-arm robots, bimanual setups and mobile manipulators. The bet is that a model that has seen many bodies learns something general about manipulation.
If the VLA idea is new to you, our VLA explainer covers the basics.
The π0 family so far
π0 is now a lineage, released through the openpi repository:
| Release | What Physical Intelligence and openpi describe | Caveat |
|---|---|---|
| π0 | The original flow-matching VLA, pretrained for fine-tuning on new robots; the paper reports large multi-robot data mixtures | Not a drop-in brain for Digit, Atlas or Figure without serious embodiment work |
| π0-FAST | An autoregressive variant that uses “FAST” action tokenization | Different inference trade-offs; read the checkpoint card |
| π0.5 | Better open-world generalization, per openpi notes; flow head supported in the repo | Still a research and fine-tuning stack, not a plant-floor service agreement |
| π0.7 | Described in a company blog and paper as a “steerable” generalist that takes richer context (language, metadata, visual subgoals), framed at about 5 billion parameters | Performance claims are company evaluations until someone replicates them |
Before you build on any of these, check the current model card for the licence, intended use and supported robots. “Open” is a spectrum, and terms change between releases.
How π0 differs from Helix and GR00T
The biggest difference is how much you get to see. π0’s checkpoints are public through openpi (verify the licence), and its demos and datasets lean on research arms such as Franka. Figure’s Helix, by contrast, isn’t available to third parties at all, and it runs on Figure’s own humanoids. NVIDIA’s Isaac GR00T N1.7 sits in between: it’s described as open and commercially licensed, and it’s tied to a reference humanoid design that Unitree markets as the H2 Plus.
None of that says which model is “best.” Nobody has published a fair head-to-head on the same robot and tasks. What it does tell you is how much you can inspect and test yourself. We cover Helix and DeepMind’s Gemini Robotics in AI models in humanoid robots, and the bigger category question in what is a robot foundation model.
What π0 isn’t
π0 being open doesn’t mean any particular warehouse humanoid ships with it in production. A vendor saying it “uses π0” should be able to tell you which checkpoint, fine-tuned on what data, running on which robot. It’s also not a substitute for site safety certification, or for change control when the model gets updated. And it doesn’t succeed every time. Published success rates depend heavily on how the evaluation was set up.
What to watch next is whether π0.7’s “steerable” approach shows up in openpi the way earlier versions did. The open-weights habit is what made π0 matter. If Physical Intelligence keeps its newest models closed, the most useful open baseline in robot learning gets frozen at an older generation.
Frequently asked questions
What is Physical Intelligence’s π0?
It’s a generalist vision-language-action model that maps camera images and language instructions to robot actions, using a flow-matching “action expert.” It was released with open weights through the openpi project.
Can I download π0?
Yes. Physical Intelligence published openpi with π0 and related checkpoints. Check the current GitHub model card for the licence, supported robots and fine-tuning instructions before you deploy anything.
Is π0 a humanoid-only model?
No. Its training and demos focus on manipulation across many robot types, including single arms, two-arm setups and mobile manipulators, not one humanoid model.
Sources
- Physical Intelligence, “π0: A Vision-Language-Action Flow Model for General Robot Control,” arXiv:2410.24164: arxiv.org
- Physical Intelligence blog: “Open Sourcing π0” (openpi); π0.7 steerable model post
- GitHub: Physical-Intelligence/openpi
Last updated: October 7, 2026. To report an error, see our corrections page. Articles are drafted with AI assistance and reviewed and edited by an editor; see our editorial policy.
