Every humanoid company now says it’s building a “foundation model.” NVIDIA has GR00T, Google DeepMind has Gemini Robotics, Figure has Helix and 1X has Redwood. Only a few of them will tell you what the model was trained on, how often it fails, or whether you can download it.
The term is borrowed from language AI, where “foundation model” means one large model pre-trained on huge amounts of data and then adapted to many jobs. In robotics, a robot foundation model is supposed to do the same for physical skills: one pre-trained system that handles perception, understands instructions and produces actions, adaptable to many tasks and ideally to many robot bodies. In practice, the phrase is doing a lot of marketing work. Here’s how to tell a real one from a rebranded demo.
What earns the “foundation model” label
There’s no official definition, so we use a practical test. Call something a robot foundation model when most of these are true. It’s pre-trained on large, diverse robot data, video-and-action data, or both. It has a general interface, usually camera images plus language in and robot actions out. That’s the vision-language-action (VLA) pattern. It can be adapted to new tasks, by fine-tuning or prompting, with far less task-specific engineering than a classic robot software stack. And its makers claim, and preferably measure, that it transfers across tasks or across different robots.
By that test, not every VLA qualifies. A small VLA trained for one task in one lab isn’t a foundation model. The word implies broad pre-training and reuse.
How it relates to the other buzzwords
VLA is the most common architecture for robot foundation models, not a separate idea. LLMs and vision-language models are often the backbone, but on their own they can’t control motors. Imitation-learning and reinforcement-learning policies are narrower specialists. They can sit on top of a foundation model or be distilled from one. “Physical AI” is a vendor slogan that bundles models, robots and simulation together. It isn’t a model type at all.
Five that get called foundation models in 2026
NVIDIA’s Isaac GR00T N1.7 is the most open of the big names. It’s described as an open, commercially licensed VLA, now paired with a reference humanoid design. Google DeepMind’s Gemini Robotics models are frontier-scale, and Apptronik says it supplies training data through a partnership. Figure’s Helix runs on Figure’s own humanoids and isn’t available as open weights; see our Figure 03 report. 1X describes Redwood AI as a vision-language model for household chores on NEO. And Unitree has posted UnifoLM-WLA-1.0 as an open-weight release on Hugging Face (check the model card for licence terms).
Notice that “foundation model” covers everything from downloadable weights to a model you’ll only ever see in a company video. That range is the reason to ask questions.
Six questions to ask before you trust the claim
- What’s the licence, and does it allow commercial use?
- Which robot bodies does it support, and which were actually in the training data?
- How many hours of real-robot data versus simulation went in?
- How was it evaluated, and was a teleoperator involved in any demo?
- Who owns the safety layer? A model isn’t a certified controller.
- Does it run on the robot or in the cloud, and what happens when the network drops?
The mistakes we see most
The most common one is treating a single, cherry-picked video as a benchmark. Close behind are assuming success in simulation means success on a real robot, and missing the teleoperator behind an “autonomous” clip (our teleoperation guide explains how to spot one). Two more are costlier. One is deploying open weights without a safety PLC or certified stop path. The other is mistaking a partner logo wall for a production service agreement.
The open question for the next year is evaluation. Language models got leaderboards. Robot foundation models still mostly grade their own homework. The first widely accepted, independently run benchmark across physical robots will do more to sort real foundation models from rebrands than any launch video.
Frequently asked questions
What is a robot foundation model?
It’s a large pre-trained AI model meant to give robots general skills, usually by mapping camera images and language to actions, that can be adapted to many tasks and robots instead of hard-coding one skill at a time.
Is every VLA a foundation model?
No. A small, task-specific VLA isn’t a foundation model. The term implies broad pre-training and intended reuse across tasks.
Are open weights enough to put a robot to work?
No. You still need to integrate the model with your robot, add safety systems, run your own evaluation, and usually collect task data.
Sources
- NVIDIA Isaac GR00T / N1.7 public materials
- Google DeepMind Gemini Robotics posts
- Figure Helix and product posts; 1X Redwood; Unitree UnifoLM Hugging Face model card
Last updated: October 7, 2026. To report an error, see our corrections page. Articles are drafted with AI assistance and reviewed and edited by an editor; see our editorial policy.
