Anyone can now download UnifoLM-WLA-1.0, the 6-billion-parameter model Unitree says drives its G1 humanoid through 64 tasks. It’s under the Apache 2.0 license, the checkpoint is a 12.5 GB file on Hugging Face, and the company has published code to fine-tune it on a single 24 GB graphics card. What Unitree hasn’t published, as far as we can find, is how often the model succeeds.
That gap is the story. Open weights from a company that has shipped thousands of humanoids are useful to researchers and robot startups, but a demo reel isn’t a benchmark. Here’s what’s in the release, what the documentation says you can run, and what to treat as a company claim.
What Unitree actually released
The release came in three steps, according to the news log in the project’s GitHub README. On September 11 Unitree published the weights of two embodied-reasoning models, UnifoLM-ER-1 and UnifoLM-ER-Flow. On September 20 it released the model modules and the code for training an action expert, the part that turns the model’s output into motion. On September 28 it posted UnifoLM-WLA-1.0-Base and code for full and LoRA fine-tuning. Every box in the README’s open-source checklist is ticked.
Hugging Face lists the Base model’s license as Apache 2.0, which permits commercial use, though the README says third-party components keep their own licenses. The model card itself is empty apart from the license line, so the real documentation lives on GitHub and on Unitree’s project page. Unitree has also put task datasets for the G1 and R1 on Hugging Face, including whole-body and gripper collections. We couldn’t find where Unitree says how much of the roughly 2,500 hours of robot data used to train the model is in those public sets.
One model for the hands and the feet
Most robot models handle either a tabletop arm or a whole body. Unitree says UnifoLM-WLA-1.0 does both with one set of weights: 64 tasks, of which 54 are tabletop manipulation and 10 are whole-body, using two-finger grippers and several five-finger hands. The company says the model was trained on about 2,500 hours of real-robot data, including its own open datasets and one called BitRobot-HIW-500, plus more than 5 million embodied-reasoning samples for the vision-language backbone.
The design follows the pattern we describe in our explainers on vision-language-action models and robot foundation models. A language-and-vision model reads the cameras and the instruction. A separate “action expert,” a flow-based decoder Unitree describes as MMDiT (a diffusion-transformer design), outputs the joint movements. Unitree also trains the backbone to predict which parts of the scene will change next, using optical flow, so the model learns what its own actions do to the world. The README credits the open-source starVLA and Qwen-Image projects as foundations.
For scale, our piece on humanoid training data notes that Physical Intelligence’s pi0 used more than 10,000 hours of robot data. A 2,500-hour model is smaller than that, and Unitree hasn’t published a head-to-head.
What the documentation says you can run
The training guide is unusually concrete. Fine-tuning keeps the vision-language backbone frozen and trains only the action head and the robot-state projector, and the default configuration targets a single 24 GB GPU. LoRA fine-tuning is supported for either part. A script replays a recorded Unitree episode through a checkpoint and plots predicted against recorded actions for every joint, which is how you’d check the model before putting it near hardware.
There’s also a model server that exposes a checkpoint over a websocket for the G1’s whole-body-control client. And here’s the catch the guide states itself: the server’s action definitions cover only the two-finger-gripper data type. To serve a checkpoint trained on Unitree’s whole-body dataset, the guide says, you would have to add the dexterous-hand and base-motion fields and extend the server’s protocol, which currently can’t carry them. The tabletop path looks ready to try. For whole-body control, the public documentation shows a gap that would take engineering work to close, though Unitree may use its own tools internally.
What the demos do and don’t show
Unitree’s demo video, hosted on the project page, labels its tabletop clips “2x Speed, Autonomous Execution” and its whole-body clips “Real Footage Throughout, No Speed-Up, Autonomous Execution.” Those labels are the company’s, and they answer the question we ask of every humanoid video, which is whether a person was steering it. Our guide to teleoperation versus autonomy explains why the label matters. Each clip is shown beside a live panel of camera views and a log of the policy running.
What’s missing is a number. We found no per-task success rate, no count of failed attempts, and no comparison on the same tasks with another open model on the project page or in the README. Unitree does publish benchmark scores for its embodied-reasoning model, UnifoLM-ER-1, which it says leads open-source models on seven of 16 perception benchmarks. That measures how well the model understands a scene, not how often a robot completes a task. The real-robot tests were run on Unitree’s G1, according to the project page, which doesn’t describe the test setups.
What to watch
Open weights change the question from “what does Unitree say?” to “what can anyone reproduce?” The first thing to look for is a lab or startup outside Unitree reporting success rates on a G1 it set up itself, and whether the whole-body server gap closes in a later release. Unitree sells the G1 for $13,500, as covered in our price guide, so independent tests are within reach of a well-funded lab. Until one appears, UnifoLM-WLA-1.0 is a real, downloadable model with no published track record.
Frequently asked questions
Is UnifoLM-WLA-1.0 open source?
Unitree released the UnifoLM-WLA-1.0-Base weights and fine-tuning code under the Apache 2.0 license, and Hugging Face lists the same license for the model. The README adds that third-party components remain under their own licenses. We couldn’t find a statement of how much of the training data is public.
How big is UnifoLM-WLA-1.0?
Unitree describes it as a 6-billion-parameter model, and the checkpoint file on Hugging Face is about 12.5 GB. The company says it was trained on roughly 2,500 hours of real-robot data.
What hardware do you need to fine-tune it?
Unitree’s guide says the default fine-tuning configuration, which freezes the vision-language backbone and trains the action head, targets a single 24 GB GPU, and it suggests lowering the batch size on less memory. LoRA fine-tuning is also supported.
Which robots does it work with?
Unitree’s real-robot tests were on the G1, and the model supports two-finger grippers and several five-finger hands. The released model server currently covers only the two-finger-gripper action format, according to Unitree’s own documentation.
Sources
All accessed October 10, 2026.
- Unitree, UnifoLM-WLA-1.0 repository and README (news log, open-source plan, license): github.com/unitreerobotics/unifolm-wla
- Unitree, “Train and Evaluate an Action Expert” (installation, fine-tuning, model server notes): docs/train_action_expert_en.md
- Unitree, UnifoLM-WLA-1.0 project page (parameters, data, tasks, demo video): unigen-x.github.io
- Hugging Face, UnifoLM-WLA-1.0-Base model page and file listing (license, checkpoint size, upload date Sep 28, 2026): huggingface.co
- Unitree US store, G1 price: shop.unitree.com
Related: NVIDIA GR00T explained · Gemini Robotics 2 explained
Last updated: October 10, 2026. To report an error, see our corrections page. Articles are drafted with AI assistance and reviewed and edited by an editor; see our editorial policy.
