What Is Diffusion Policy for Robot Learning? (2026 Explainer)

Diffusion Policy generates robot motions the way image models generate pictures, by denoising. How it works, what the Columbia/TRI/MIT paper and TRI’s Large Behavior Models showed, how it relates to VLAs, and its limits.
Toyota Research Institute bimanual robot with two white robot arms whisking eggs in a bowl on a wooden table while a researcher behind it holds teleoperation controllers in a robotics lab Toyota Research Institute bimanual robot with two white robot arms whisking eggs in a bowl on a wooden table while a researcher behind it holds teleoperation controllers in a robotics lab
A Toyota Research Institute researcher teaches a two-arm robot to whisk eggs by teleoperation — the kind of demonstration Diffusion Policy learns from. Photo: Toyota Research Institute (Toyota USA Newsroom).

In September 2023, Toyota Research Institute said it had taught robots more than 60 dexterous skills, including pouring, using tools and handling floppy objects, “without writing a single line of new code.” Each new skill took only new human demonstrations. Two years later, a descendant of the same technique was controlling Boston Dynamics’ electric Atlas, walking and manipulating, end to end.

The technique is called Diffusion Policy. It borrows the “denoising” trick that image generators such as Stable Diffusion use, but instead of turning random noise into a picture, it turns random noise into a short sequence of robot motions, guided by what the robot’s cameras see. Researchers at Columbia University, TRI and MIT introduced it in 2023. It’s since become one of the most widely used methods in robot learning, and its action-generation idea now sits inside many of the bigger robot AI models you read about. If you want to understand how today’s robots learn from demonstrations, this is the paper to know.

The paper behind it

Item What is published
Authors / labs Cheng Chi, Siyuan Feng, Yilun Du, Zhenjia Xu, Eric Cousineau, Benjamin Burchfiel and Shuran Song (Russ Tedrake joined the extended version). Columbia University, Toyota Research Institute, MIT
First published March 2023 (arXiv 2303.04137); Robotics: Science and Systems (RSS) 2023; extended version March 2024
What it is A policy representation: the robot’s behaviour is modelled as a conditional denoising diffusion process over actions
Learning type Imitation learning (behaviour cloning) from teleoperated demonstrations
Headline benchmark result Average 46.9% improvement over the best prior methods (12 tasks in the conference paper, 15 tasks from 4 benchmarks in the extended version)
Inference speed About 0.1 s per inference on an NVIDIA RTX 3080, using 10 denoising steps (DDIM) at test time vs 100 in training
Robots in the paper UR5 and Franka arms, plus bimanual setups for egg beating, mat unrolling and shirt folding
Code Open source (MIT License), with data and training details

How Diffusion Policy turns noise into motion

A conventional learned policy looks at the current camera image and outputs one action, such as “move the gripper 2 cm left.” Diffusion Policy does three things differently:

Advertisement

  1. It starts from noise and refines. The model begins with a random sequence of actions. Over several denoising steps, guided by the camera images and robot state, it pushes that sequence towards something that looks like the human demonstrations.
  2. It predicts action chunks, not single steps. At each step it looks at the last few observations, predicts a block of future actions, executes the first part of that block, then replans. Control engineers call this receding-horizon control. It keeps motion smooth while letting the robot react to changes.
  3. It can represent several valid answers. Human demonstrations are often “multimodal”: one person goes around an obstacle on the left, another on the right. A simple regression policy averages the two and drives into the obstacle. A diffusion model can commit to one valid mode. The authors list this as a main advantage, together with handling high-dimensional action spaces and stable training.

The paper offers two network backbones: a convolutional (CNN) version that the authors found works well out of the box on most tasks, and a time-series transformer version for tasks where actions change quickly.

From pushing a T-block to flipping mugs

Task Setup Published result
Push-T (push a T-shaped block into a target) UR5 arm; policy runs at 10 Hz, commands interpolated to 125 Hz 95% success, close to human demonstrators
Mug flipping (pick a mug, place it lip-down, handle left) 6-DoF arm 90% success; the LSTM-GMM baseline scored 0%
Sauce pouring and spreading on pizza dough Liquid / periodic motion Coverage close to human demonstrations; baseline largely failed
Bimanual egg beater, mat unrolling, shirt folding Two-arm teleoperated setups (extended paper) Added to show the method scales to two-arm, deformable-object tasks

That mug result is the one that stands out: 90% for Diffusion Policy against 0% for the LSTM-GMM baseline. Remember, though, that these are lab results on fixed setups, measured by the authors. They show the method is data-efficient and reliable for single skills. They don’t show a general-purpose robot.

How Toyota scaled it into “Large Behavior Models”

TRI made Diffusion Policy the basis of its robot-learning programme:

  • September 19, 2023: TRI said it had taught robots more than 60 dexterous skills, including pouring, using tools and handling deformable objects, “without writing a single line of new code,” only new demonstrations. It set targets of hundreds of skills by the end of 2023 and 1,000 by the end of 2024. In the sources we checked, TRI has not published a confirmed count against the 1,000 target.
  • July 2025: TRI’s paper “A Careful Examination of Large Behavior Models for Multitask Dexterous Manipulation” scaled the approach into multitask diffusion-transformer models trained on almost 1,700 hours of robot data. It ran 1,800 real-world and more than 47,000 simulated evaluation rollouts. Fine-tuned LBMs learned new tasks with 3–5× less data than single-task policies trained from scratch. TRI also reported mixed results for the pretrained models used without fine-tuning.
  • August 2025: Boston Dynamics and TRI showed an LBM controlling electric Atlas end to end, for both walking and manipulation. They described it as a 450-million-parameter diffusion transformer trained with a flow-matching objective, predicting 1.6-second action chunks at 30 Hz. See our Atlas commercialization coverage for the product side.

Diffusion Policy isn’t a VLA, but many VLAs borrow from it

The two terms are often confused. A vision-language-action (VLA) model is a class of model: a vision-language backbone that outputs robot actions. Diffusion Policy is a method for producing actions. Many 2024–2026 VLAs use a diffusion or flow-matching “action head” descended from this work.

System How it outputs actions Link to Diffusion Policy
Original Diffusion Policy (2023) Denoising diffusion over action chunks; no language backbone required The original method
Google DeepMind RT-2 (2023) Discrete action tokens from a vision-language model A different approach; see our RT-2 explainer
Physical Intelligence π0 Flow-matching “action expert” on a vision-language backbone Closely related generative family (flow matching); see π0 explained
NVIDIA GR00T N1 Fast diffusion-transformer controller under a slower vision-language model Diffusion-style action generation for humanoids; see GR00T explained
TRI LBMs / Atlas LBM Multitask diffusion transformers with language prompts Direct scale-up of Diffusion Policy

The comparison works in both directions. The OpenVLA paper (2024) reports that a fine-tuned 7B VLA beat from-scratch Diffusion Policy by 20.4% in multi-task settings with several objects and language instructions. Diffusion Policy remained strong on narrow, single-task dexterous skills. Which approach is better depends on the job.

Where Diffusion Policy still falls short

It’s only as good as its demonstrations. The authors note that it inherits the limits of behavior cloning: with too few or poor demonstrations, performance drops, and by default it doesn’t learn from trial and error. It’s also computationally heavier than a simple policy. Several denoising steps cost more than a single forward pass, and the authors say action-chunk prediction only partly offsets this and “may not suffice for tasks requiring high rate control.”

The 2023 method also trains one policy per skill. Multitask, language-steerable versions like TRI’s LBMs are newer, and by TRI’s own account they still give mixed results without fine-tuning. Then there’s evaluation. TRI’s 2025 work stresses that robot policies have to be tested on real hardware with blind, statistically controlled trials, because offline metrics aren’t enough. That’s a good reason to treat any single demo video with care (see teleoperation vs autonomy).

Five questions for a vendor that says “we use diffusion policies”

  1. How many demonstrations did each skill need, and who collected them?
  2. Is it one policy per task or one multitask model? Does it take language instructions?
  3. What is the control rate, and how is latency handled (action chunks, onboard GPU)?
  4. What success rate was measured, over how many blind trials, under what changes (lighting, object positions, new objects)?
  5. What happens on failure: recovery behaviour, operator takeover or a safety stop? See humanoid robot safety systems.

One number from 2023 is still worth chasing. TRI set a target of 1,000 skills by the end of 2024 and, in the sources we checked, hasn’t published a count against it. Whether that figure appears, and how many of those skills transfer to a humanoid like Atlas, is the best test of whether learning from demonstrations can scale.

Frequently asked questions

What is Diffusion Policy in robotics?

Diffusion Policy is a robot-learning method that represents a robot’s behaviour as a denoising diffusion process. Starting from random noise, the model refines a short sequence of future actions based on camera images and robot state. It was introduced by Columbia University, Toyota Research Institute and MIT in 2023.

Is Diffusion Policy the same as Stable Diffusion?

No. Both use denoising diffusion, but Stable Diffusion generates images from text, while Diffusion Policy generates robot action sequences from sensor observations. It does not create pictures.

Why is Diffusion Policy good for robots?

The authors highlight three strengths: it handles demonstrations where several different motions are valid, it copes with high-dimensional actions such as two-arm control, and it trains stably. Predicting action chunks with receding-horizon control also makes motion smoother.

Do humanoid robots use Diffusion Policy?

Descendants of it, yes. Boston Dynamics and Toyota Research Institute showed a 450-million-parameter diffusion-transformer Large Behavior Model controlling Atlas in August 2025, and NVIDIA’s GR00T N1 uses a diffusion-transformer action module. Vendors rarely publish which exact method runs in shipping products.

Sources

  • Chi, Xu, Feng, Cousineau, Du, Burchfiel, Tedrake, Song, “Diffusion Policy: Visuomotor Policy Learning via Action Diffusion,” arXiv 2303.04137 (RSS 2023; extended version v5, March 2024) [paper]
  • Diffusion Policy project page and code (diffusion-policy.cs.columbia.edu; MIT License) [academic]
  • Toyota Research Institute, “Toyota Research Institute Unveils Breakthrough in Teaching Robots New Behaviors” (Sept 19, 2023) and TRI Medium technical post [company]
  • TRI LBM Team, “A Careful Examination of Large Behavior Models for Multitask Dexterous Manipulation,” arXiv 2507.05331 (July 2025) [paper]
  • Boston Dynamics, “Large Behavior Models and Atlas Find New Footing” (Aug 2025) and TRI/Toyota newsroom release [company]
  • Kim et al., “OpenVLA: An Open-Source Vision-Language-Action Model,” CoRL 2024 [paper]

Related: What is RT-2? · What is a VLA model? · π0 explained · NVIDIA GR00T explained

Last updated: October 7, 2026. To report an error, see our corrections page. Articles are drafted with AI assistance and reviewed and edited by an editor; see our editorial policy.

Keep Up to Date with the Most Important News

By pressing the Subscribe button, you confirm that you have read and are agreeing to our Privacy Policy and Terms of Use
Advertisement