What Is RT-2? Google DeepMind’s Robotics Transformer Explained (2026)

RT-2 is the 2023 Google DeepMind model that coined “vision-language-action.” How it turns PaLI-X and PaLM-E into robot actions, the 62% vs 32% result, its limits, and why you cannot download it.
PAL Robotics TIAGo mobile manipulator robot at a RoboCup exhibit stand, with a REEM-C humanoid robot behind it PAL Robotics TIAGo mobile manipulator robot at a RoboCup exhibit stand, with a REEM-C humanoid robot behind it
PAL Robotics TIAGo at RoboCup 2016 in Leipzig: the same class of wheeled, arm-equipped mobile manipulator as the Everyday Robots units RT-2 was trained and tested on. Not the RT-2 robot and not a Google DeepMind photo. Photo: ubahnverleih / Wikimedia Commons, CC0.

Asked to “move the apple to the cup with the same color,” a Google DeepMind robot in 2023 did it, even though nothing like that instruction appeared in its robot training data. The model running it had picked up colors, numbers and symbols from the web, and it turned out that knowledge could steer a robot. In one demo it even wrote out a short plan and chose a rock as an improvised hammer.

That model was RT-2 (Robotic Transformer 2), which Google DeepMind announced on July 28, 2023. It’s the paper that made “vision-language-action model” a mainstream robotics term. RT-2 took large vision-language models trained on web data and fine-tuned them to output robot actions, so a single network could read a camera image and an instruction and command a robot arm directly. Almost every VLA headline since, from OpenVLA to Gemini Robotics 2, builds on that recipe. You can’t download RT-2, and it’s three years old. It still sets the questions worth asking about any robot that claims to “run on a VLA.”

RT-2 by the numbers

Item What is published
Developer Google DeepMind (robotics team, building on Google’s RT-1 work)
Announced July 28, 2023 (arXiv 2307.15818; presented at CoRL 2023)
Model type Vision-language-action (VLA) model. The paper introduced the term.
Backbones Two versions: RT-2-PaLI-X (5B and 55B parameters) and RT-2-PaLM-E (12B)
Robot training data The RT-1 dataset: demonstrations collected with 13 robots over 17 months in an office kitchen, mixed with web-scale vision-language data
Robot used A 7-degree-of-freedom mobile manipulator (the Everyday Robots platform used in the RT-series papers)
Output Robot actions written as 8 integer “text” tokens (256 bins per continuous dimension)
Control speed 55B model: 1–3 Hz. 5B model: about 5 Hz. Both run on a multi-TPU cloud service, not on the robot.
Evaluation About 6,000 real-robot evaluation trials
Weights / API Not released. No public RT-2 checkpoint or API.

Why Google started from a web model

RT-2’s predecessor, RT-1 (December 2022), was a much smaller transformer, with roughly 35 million parameters, trained only on robot data. That data was about 130,000 demonstrations of more than 700 tasks, collected by a fleet of 13 robots over 17 months. RT-1 handled tasks it had seen well, but struggled with new objects, backgrounds and instructions. Google open-sourced the RT-1 code.

Advertisement

RT-2 asked a simple question: if a model has already learned from web images and captions what colours, numbers, logos and “the smallest object” are, can that knowledge carry over to a robot? To test it, the team reused the RT-1 robot data but started from large pretrained vision-language models instead of training from scratch.

The trick: writing robot actions as text

The key trick is in how RT-2 represents actions. Each robot command has eight parts: six numbers for the change in end-effector position and rotation, one for how far the gripper opens, and one flag that ends the episode. RT-2 splits each continuous value into 256 evenly spaced bins, so a whole action becomes a string of eight integers. Those integers are mapped to tokens the language model already knows. For the model, “move the arm” looks the same as writing a short sentence.

The team then co-fine-tuned the model on two kinds of data at once:

  • the original web data, such as visual question answering and captioning, so the model keeps its general knowledge;
  • robot trajectories, where the answer to “what should the robot do to pick up the apple?” is an action string.

When RT-2 is asked for a robot action, its output is limited to valid action tokens. Those tokens are turned back into motor commands and sent to the robot in a closed loop. Because a 55-billion-parameter model cannot run on a robot’s onboard computer, DeepMind served it from TPUs in the cloud and queried it over the network.

Twice as good on things it had never seen

Test Published result What it means
Seen tasks (RT-1 suite, 200+ instructions) RT-2 roughly matched RT-1 Adding web knowledge did not hurt basic skills
Unseen objects, backgrounds and environments Average success 62% for RT-2 vs 32% for RT-1 About 2× better generalization from web pretraining
“Emergent” instructions (symbols, reasoning, recognising people) 2× to 3× RT-1’s success; the best PaLI-X model averaged more than 3× RT-1 The robot can follow commands that never appeared in its robot data, such as “move the apple to the cup with the same color” or “move X near the sum of two plus one”
Chain-of-thought variant Qualitative demos The model wrote a short plan before acting, for example choosing a rock as an improvised hammer
PaLI-X vs PaLM-E Similar on average; PaLM-E better on harder generalization cases and math, PaLI-X better on easier cases The choice of backbone shapes what transfers

The headline number is 62% versus 32% on unseen objects, backgrounds and environments. That’s the clearest evidence in the paper that web knowledge transfers to robot control. Keep in mind that these results come from DeepMind’s own lab robots and evaluation protocol. They aren’t proof that RT-2 would work in a customer’s warehouse or home.

What RT-2 couldn’t do, in its authors’ own words

The most important limit is one the paper states plainly: web pretraining didn’t give the robot any new physical skills. RT-2 learned to apply the skills in its robot data to new situations, like picking up an object it hadn’t seen. It couldn’t invent a motion it had never been shown. Knowing what a hammer is doesn’t teach you how to swing one.

It was also slow and tied to the cloud. Control at 1 to 3 Hz from a cloud TPU service is far below the 30 to 200 Hz loops used for dexterous or whole-body humanoid control. That’s why later systems added fast low-level controllers; our VLA guide covers these “System 1 / System 2” designs. The authors also noted that only a handful of vision-language models were available to build RT-2-style systems, and they called for more open ones.

And RT-2 itself stayed closed. DeepMind released no weights, code or API, so outside teams couldn’t reproduce the results directly. That gap is a big part of why what came next looks the way it does.

What came after: RT-X, OpenVLA, π0 and Gemini Robotics

Model / project Date Relationship to RT-2 Open?
Open X-Embodiment / RT-X October 2023 Pooled more than 1 million real robot trajectories from 22 robot types, contributed by 21 institutions. RT-2-X, trained on this mix, scored about 3× higher than RT-2 on emergent-skill tests. Dataset and RT-1-X checkpoint released; RT-2-X not released
OpenVLA June 2024 A 7B open VLA trained on 970,000 Open X-Embodiment episodes. Its paper reports 16.5 points higher absolute task success than the 55B RT-2-X across 29 tasks, with 7× fewer parameters. Yes: weights and code
Physical Intelligence π0 October 2024 onward Keeps the VLA idea but outputs continuous actions through a flow-matching “action expert” instead of RT-2’s discrete tokens. See our π0 explainer. Open weights (openpi)
Gemini Robotics / Gemini Robotics 2 March 2025 / July 30, 2026 DeepMind’s successors, built on Gemini. Gemini Robotics 2 targets full humanoids and two-arm robots. See Gemini Robotics 2 explained. Only the ER 2 reasoning model is publicly available; VLAs are partner-only

Can you use RT-2 in 2026?

No, not directly. Google DeepMind never published RT-2 weights or an RT-2 API, and its current robotics work is branded Gemini Robotics. To experiment with an RT-2-style model today, the practical routes are open VLAs such as OpenVLA or π0 (openpi), NVIDIA’s Isaac GR00T models for humanoids, or the publicly available Gemini Robotics ER 2 reasoning model, which plans tasks but does not output motor commands itself.

Five RT-2 questions to ask any “VLA-powered” robot

RT-2’s real legacy is a checklist. Ask these of any vendor that says its robot runs on a VLA:

  1. How are actions represented? Discrete tokens like RT-2, or continuous chunks from a diffusion or flow head?
  2. What is the control rate, and where does the model run? On the robot or in the cloud, and what happens when the network drops?
  3. What robot data is behind it? Web knowledge helps with new objects and instructions, but the paper shows physical skills still come from robot demonstrations.
  4. How was generalization tested? On unseen objects, backgrounds and rooms, or only on scenes from the training data?
  5. Is it autonomous? Check whether the demo was teleoperated. See teleoperation vs autonomy.

The open question RT-2 raised three years ago still hasn’t been fully answered: how much of a robot’s skill can come from the web, and how much still has to come from expensive robot demonstrations? Watch the data sections of the next big VLA papers. That’s where the answer will show up first.

Frequently asked questions

What is RT-2 in robotics?

RT-2 (Robotic Transformer 2) is a vision-language-action model that Google DeepMind announced in July 2023. It fine-tunes large vision-language models (PaLI-X or PaLM-E) on robot demonstrations, so one model maps a camera image and a text instruction directly to robot actions.

How is RT-2 different from RT-1?

RT-1 was a roughly 35-million-parameter transformer trained only on robot data. RT-2 starts from vision-language models with billions of parameters that were pretrained on web data. In DeepMind’s tests, that raised average success on unseen scenarios from 32% to 62%.

Is RT-2 open source?

No. Google DeepMind did not release RT-2 weights, code or an API. The closest open alternatives are OpenVLA, which its paper reports outperformed the 55B RT-2-X, and Physical Intelligence’s open-weight π0 models.

Does RT-2 control humanoid robots?

Not in the published work. RT-2 was trained and tested on a single-arm mobile manipulator. DeepMind’s later Gemini Robotics 2 family is the line aimed at full humanoids and two-arm robots.

Sources

  • Google DeepMind, “RT-2: New model translates vision and language into action” (July 28, 2023) [company]
  • Brohan, Zitkovich et al., “RT-2: Vision-Language-Action Models Transfer Web Knowledge to Robotic Control,” arXiv 2307.15818 / CoRL 2023 (PMLR v229) [paper]
  • Brohan et al., “RT-1: Robotics Transformer for Real-World Control at Scale” (2022) [paper]
  • Open X-Embodiment Collaboration, “Open X-Embodiment: Robotic Learning Datasets and RT-X Models,” arXiv 2310.08864 [paper]
  • Kim, Pertsch, Karamcheti et al., “OpenVLA: An Open-Source Vision-Language-Action Model,” CoRL 2024 (PMLR v270) [paper]
  • Google DeepMind Gemini Robotics 2 announcement and model pages (July 30, 2026) [company], as summarised in our Gemini Robotics 2 explainer

Related: What is a VLA model? · What is a robot foundation model? · Gemini Robotics 2 explained · What is Diffusion Policy?

Last updated: October 7, 2026. To report an error, see our corrections page. Articles are drafted with AI assistance and reviewed and edited by an editor; see our editorial policy.

Keep Up to Date with the Most Important News

By pressing the Subscribe button, you confirm that you have read and are agreeing to our Privacy Policy and Terms of Use
Advertisement