Unitree Open-Sourced Its 6B Humanoid Model. Here’s What You Can Actually Run

Unitree has released the weights and fine-tuning code for UnifoLM-WLA-1.0, a 6-billion-parameter humanoid model it says handles 64 tasks. What’s in the release, what the docs say you can run on one GPU, and what Unitree hasn’t published.
Silver and black Unitree humanoid robot standing in a white kitchen, holding a mug above an open dishwasher drawer Silver and black Unitree humanoid robot standing in a white kitchen, holding a mug above an open dishwasher drawer
A Unitree humanoid loading a dishwasher, in a frame from the demo video on Unitree's UnifoLM-WLA-1.0 project page. Unitree labels the video's whole-body clips autonomous and not sped up. Image: Unitree Robotics.

Anyone can now download UnifoLM-WLA-1.0, the 6-billion-parameter model Unitree says drives its G1 humanoid through 64 tasks. It’s under the Apache 2.0 license, the checkpoint is a 12.5 GB file on Hugging Face, and the company has published code to fine-tune it on a single 24 GB graphics card. What Unitree hasn’t published, as far as we can find, is how often the model succeeds.

That gap is the story. Open weights from a company that has shipped thousands of humanoids are useful to researchers and robot startups, but a demo reel isn’t a benchmark. Here’s what’s in the release, what the documentation says you can run, and what to treat as a company claim.

What Unitree actually released

The release came in three steps, according to the news log in the project’s GitHub README. On September 11 Unitree published the weights of two embodied-reasoning models, UnifoLM-ER-1 and UnifoLM-ER-Flow. On September 20 it released the model modules and the code for training an action expert, the part that turns the model’s output into motion. On September 28 it posted UnifoLM-WLA-1.0-Base and code for full and LoRA fine-tuning. Every box in the README’s open-source checklist is ticked.

Advertisement

Hugging Face lists the Base model’s license as Apache 2.0, which permits commercial use, though the README says third-party components keep their own licenses. The model card itself is empty apart from the license line, so the real documentation lives on GitHub and on Unitree’s project page. Unitree has also put task datasets for the G1 and R1 on Hugging Face, including whole-body and gripper collections. We couldn’t find where Unitree says how much of the roughly 2,500 hours of robot data used to train the model is in those public sets.

One model for the hands and the feet

Most robot models handle either a tabletop arm or a whole body. Unitree says UnifoLM-WLA-1.0 does both with one set of weights: 64 tasks, of which 54 are tabletop manipulation and 10 are whole-body, using two-finger grippers and several five-finger hands. The company says the model was trained on about 2,500 hours of real-robot data, including its own open datasets and one called BitRobot-HIW-500, plus more than 5 million embodied-reasoning samples for the vision-language backbone.

The design follows the pattern we describe in our explainers on vision-language-action models and robot foundation models. A language-and-vision model reads the cameras and the instruction. A separate “action expert,” a flow-based decoder Unitree describes as MMDiT (a diffusion-transformer design), outputs the joint movements. Unitree also trains the backbone to predict which parts of the scene will change next, using optical flow, so the model learns what its own actions do to the world. The README credits the open-source starVLA and Qwen-Image projects as foundations.

For scale, our piece on humanoid training data notes that Physical Intelligence’s pi0 used more than 10,000 hours of robot data. A 2,500-hour model is smaller than that, and Unitree hasn’t published a head-to-head.

What the documentation says you can run

The training guide is unusually concrete. Fine-tuning keeps the vision-language backbone frozen and trains only the action head and the robot-state projector, and the default configuration targets a single 24 GB GPU. LoRA fine-tuning is supported for either part. A script replays a recorded Unitree episode through a checkpoint and plots predicted against recorded actions for every joint, which is how you’d check the model before putting it near hardware.

There’s also a model server that exposes a checkpoint over a websocket for the G1’s whole-body-control client. And here’s the catch the guide states itself: the server’s action definitions cover only the two-finger-gripper data type. To serve a checkpoint trained on Unitree’s whole-body dataset, the guide says, you would have to add the dexterous-hand and base-motion fields and extend the server’s protocol, which currently can’t carry them. The tabletop path looks ready to try. For whole-body control, the public documentation shows a gap that would take engineering work to close, though Unitree may use its own tools internally.

What the demos do and don’t show

Unitree’s demo video, hosted on the project page, labels its tabletop clips “2x Speed, Autonomous Execution” and its whole-body clips “Real Footage Throughout, No Speed-Up, Autonomous Execution.” Those labels are the company’s, and they answer the question we ask of every humanoid video, which is whether a person was steering it. Our guide to teleoperation versus autonomy explains why the label matters. Each clip is shown beside a live panel of camera views and a log of the policy running.

What’s missing is a number. We found no per-task success rate, no count of failed attempts, and no comparison on the same tasks with another open model on the project page or in the README. Unitree does publish benchmark scores for its embodied-reasoning model, UnifoLM-ER-1, which it says leads open-source models on seven of 16 perception benchmarks. That measures how well the model understands a scene, not how often a robot completes a task. The real-robot tests were run on Unitree’s G1, according to the project page, which doesn’t describe the test setups.

What to watch

Open weights change the question from “what does Unitree say?” to “what can anyone reproduce?” The first thing to look for is a lab or startup outside Unitree reporting success rates on a G1 it set up itself, and whether the whole-body server gap closes in a later release. Unitree sells the G1 for $13,500, as covered in our price guide, so independent tests are within reach of a well-funded lab. Until one appears, UnifoLM-WLA-1.0 is a real, downloadable model with no published track record.

Frequently asked questions

Is UnifoLM-WLA-1.0 open source?

Unitree released the UnifoLM-WLA-1.0-Base weights and fine-tuning code under the Apache 2.0 license, and Hugging Face lists the same license for the model. The README adds that third-party components remain under their own licenses. We couldn’t find a statement of how much of the training data is public.

How big is UnifoLM-WLA-1.0?

Unitree describes it as a 6-billion-parameter model, and the checkpoint file on Hugging Face is about 12.5 GB. The company says it was trained on roughly 2,500 hours of real-robot data.

What hardware do you need to fine-tune it?

Unitree’s guide says the default fine-tuning configuration, which freezes the vision-language backbone and trains the action head, targets a single 24 GB GPU, and it suggests lowering the batch size on less memory. LoRA fine-tuning is also supported.

Which robots does it work with?

Unitree’s real-robot tests were on the G1, and the model supports two-finger grippers and several five-finger hands. The released model server currently covers only the two-finger-gripper action format, according to Unitree’s own documentation.

Sources

All accessed October 10, 2026.

Related: NVIDIA GR00T explained · Gemini Robotics 2 explained

Last updated: October 10, 2026. To report an error, see our corrections page. Articles are drafted with AI assistance and reviewed and edited by an editor; see our editorial policy.

Keep Up to Date with the Most Important News

By pressing the Subscribe button, you confirm that you have read and are agreeing to our Privacy Policy and Terms of Use
Advertisement