A vision-language-action (VLA) model is an AI model that takes camera images and a written or spoken instruction as input and outputs the actions a robot should take. Google DeepMind introduced the term in 2023 with RT-2. Since then, VLAs from NVIDIA, Physical Intelligence, Figure and Google DeepMind have become the main approach for teaching robots general tasks, though they remain limited by data, speed, and dexterity.
VLA in one minute
- Vision: reads images from the robot’s cameras.
- Language: understands an instruction such as “put the cup in the sink.”
- Action: outputs motor commands or joint targets, usually a short sequence at a time.
- Why it matters: one model can handle many tasks and objects, instead of a separate program for each.
How a VLA works
Most VLAs start from a vision-language model (VLM), an AI trained on images and text from the web. Robot data is then added so the model learns to output actions. RT-2 did this by writing actions as text tokens that the language model could generate, and training on web data and robot data together. Later models added dedicated components for actions.
Many recent designs split the work in two, which the field often calls a dual-system design:
| Part | Job | Speed (as reported at release) |
|---|---|---|
| System 2 (slow) | Understands the scene and the instruction; decides what to do | About 7–10 times per second |
| System 1 (fast) | Turns that decision into smooth motor control | About 120–200 times per second |
The speed gap matters because a large model is too slow to balance and move limbs on its own. Separately, humanoids typically use a lower-level whole-body controller for balance and walking. Boston Dynamics says its whole-body controllers are trained with reinforcement learning in simulation, and a VLA sits above that layer rather than replacing it. See our explainer on reinforcement learning in robots.
Timeline: from RT-2 to Gemini Robotics 2
| Model | Who | When | Key published details |
|---|---|---|---|
| RT-2 | Google DeepMind | July 2023 | Coined “VLA.” Actions written as text tokens; trained on web and robot data. Models of 5B and 55B parameters ran at roughly 5 Hz and 1–3 Hz. Roughly doubled generalization to unseen situations versus RT-1, per the authors. Could not learn motions beyond those in robot data. |
| OpenVLA | Stanford, Berkeley and others | June 2024 | Open 7B-parameter model trained on about 970,000 robot demonstrations. Authors report 16.5 percentage points higher success than the 55B RT-2-X across 29 tasks. About 6 Hz on one high-end GPU; can be fine-tuned on consumer GPUs. |
| π0 | Physical Intelligence | October 2024 | 3.3B parameters: a 3B language model plus a 300M-parameter “action expert” using flow matching. Outputs action chunks of 50 steps at up to 50 Hz. Pretrained on more than 10,000 hours of robot data. |
| Helix | Figure | February 2025 | Two-part system for a humanoid upper body: a 7B model at 7–9 Hz and an 80M-parameter controller at 200 Hz controlling 35 degrees of freedom. Trained on about 500 hours of teleoperated data, per Figure. Runs on embedded onboard GPUs. |
| GR00T N1 | NVIDIA | March 2025 | Open foundation model for humanoids with a slow VLM (about 10 Hz) and a fast diffusion-transformer controller (about 120 Hz). The 2B-class version has 2.2B parameters. Tested on the GR-1 humanoid. |
| Gemini Robotics 2 | Google DeepMind | July 2026 | Three models: a VLA for whole-humanoid control, an embodied-reasoning model (ER 2) available through Google AI Studio, and an on-device model that DeepMind says adapts to new two-arm robots with fewer than 200 examples. VLA models are available to early-access partners. Demonstrated on Apptronik Apollo 2. |
Parameter counts and speeds are as published at each release and may have changed. Benchmark comparisons are the developers’ own results, not independent tests.
Where the training data comes from
VLAs need examples of robots doing tasks. Common sources are:
- Teleoperation: a person controls the robot while it records. Figure’s Helix used about 500 hours of this.
- Shared datasets: OpenVLA used roughly 970,000 demonstrations from the Open X-Embodiment collection of many labs.
- Simulation: useful for control and motion practice, harder for realistic contact.
- Human video: abundant, but it does not show the robot’s own forces and joints.
The amount and quality of this data is often cited by developers as a main bottleneck, and several companies say they are building robot fleets partly to collect it. We explain one case in how humanoid robots learn new tasks.
Limits to keep in mind
- Tasks are still narrow and short. NVIDIA’s GR00T N1 paper describes short-horizon tabletop tasks as a limitation.
- Dexterity. DeepMind itself says multi-finger dexterous manipulation “remains challenging” and that movement speed needs work.
- Cost and latency. RT-2’s largest model ran at 1–3 Hz; newer designs address this by splitting fast and slow parts.
- Demos versus reliability. Success rates in papers are often below 100 percent. OpenVLA’s authors report typical success below 90 percent on many tasks.
- Safety. DeepMind released a safety benchmark, ASIMOV-Agentic, with Gemini Robotics 2; independent evaluation is still developing.
How to read a VLA claim
- Is the task shown autonomously, or with a remote operator?
- Was it tested in the training environment or a new one?
- How many trials, and what was the success rate?
- Is the model available (open weights, API, early access) or only shown in a video?
For the language-model side, see large language models in robotics and AI models in humanoid robots. For the hardware that these models run on, see edge AI in robots and our humanoid robot explainer.
Frequently asked questions
Is a VLA the same as a large language model?
No. An LLM works with text. A VLA also takes images and produces robot actions, and is usually built on a vision-language model with an added action output.
Do VLAs run on the robot or in the cloud?
Both exist. Figure says Helix runs on onboard embedded GPUs, and DeepMind offers an on-device Gemini Robotics model. Larger reasoning models may run elsewhere.
Are VLA models open source?
Some are. OpenVLA and NVIDIA’s GR00T N1 have been released openly; others, such as Gemini Robotics 2, are available through early access or limited programs.
Related reading
Sources
- Brohan et al., “RT-2: Vision-Language-Action Models Transfer Web Knowledge to Robotic Control,” arXiv:2307.15818: arxiv.org/abs/2307.15818
- Kim et al., “OpenVLA: An Open-Source Vision-Language-Action Model,” arXiv:2406.09246: arxiv.org/abs/2406.09246
- Black et al., “π0: A Vision-Language-Action Flow Model for General Robot Control,” arXiv:2410.24164: arxiv.org/abs/2410.24164
- NVIDIA, “GR00T N1: An Open Foundation Model for Generalist Humanoid Robots,” arXiv:2503.14734: arxiv.org/abs/2503.14734
- Figure, “Helix: A Vision-Language-Action Model for Generalist Humanoid Control” (Feb 20, 2025): figure.ai/news/helix
- Google DeepMind, “Gemini Robotics 2 brings whole-body intelligence to robots” (July 30, 2026): deepmind.google
- Boston Dynamics, “Robot Hands for Modern AI and Real Work” (whole-body control trained with RL in simulation): bostondynamics.com
Last updated: October 4, 2026. To report an error, see our corrections page. Articles are drafted with AI assistance and reviewed and edited by an editor; see our editorial policy.
