Sequoia Capital has led a $60 million Series B in Mecka, a startup founded in 2024 that pays people to record themselves doing ordinary tasks so robot makers can train on the footage. Mecka announced the round on October 7, 2026, with NVIDIA, Microsoft’s venture fund M12, Qualcomm Ventures and Samsung joining as new investors. It is the clearest sign yet that investors see humanoid robot training data as a business in its own right, separate from the robots.
The reason is simple to state and hard to fix. Language models learned from an internet’s worth of text. There is no equivalent library of robots folding towels, sorting parts or tidying living rooms, and most of what exists took a person, a robot and a lab to record. Mecka, Figure, NVIDIA, Tesla and Apptronik are each attacking that shortage from a different angle. Here’s what they’re doing, what the numbers actually show and what remains unproven.
What Mecka is selling
According to TechCrunch, Mecka pays contributors to capture everyday tasks with body-worn sensors and smartphones, then turns the recordings into training-ready data. In its funding announcement, the company says it passed a $100 million annual revenue run-rate in June 2026 and expects to reach $300 million by the end of the year. It says its customers include “several of the top frontier robotics labs and multiple Mag 7 companies,” but it doesn’t name any of them, and none of the revenue figures has been independently verified.
Mecka didn’t disclose a valuation. TechCrunch reported in September that the deal was nearing a $500 million valuation, and its October 7 story said rival XDOF was in talks at $1.2 billion. Both figures come from unnamed sources. Mecka’s angel investors include Milan Kovac, the former head of Tesla’s Optimus program, which says something about where people who have built humanoids think the constraint sits.
Why humanoid robot training data is so scarce
NVIDIA’s researchers put the problem bluntly in the GR00T N1 paper: “no Internet of humanoid robot datasets exist for large-scale pre-training.” The data available for any single humanoid, they wrote, “would be orders of magnitude too small.” And it’s slow to make. “Robot data scales linearly with human labor,” the paper says, because someone usually has to teleoperate the robot to produce each example.
The largest public pools show the scale. The Open X-Embodiment collaboration, published in 2023, combined 60 existing datasets into more than one million robot trajectories from 22 robot types, contributed by 21 institutions. That was a landmark for robot learning, and it is still small and mostly made up of robot arms rather than humanoids. Today’s best-known robot models are trained on thousands of hours of robot time, not millions:
- Physical Intelligence’s π0 paper describes pretraining on more than 10,000 hours of robot data.
- Toyota Research Institute’s Large Behavior Models were trained on almost 1,700 hours.
- Unitree’s UnifoLM-WLA-1.0 model card cites about 2,500 hours of real-robot data.
For background on how these models turn data into motion, see our explainers on robot foundation models and vision-language-action models.
The data pyramid: web video, simulation and robot time
The most useful way to picture the industry’s answer comes from that same NVIDIA paper, which describes a “data pyramid.” The wide base is web data and video of people doing things. The middle layer is synthetic data: simulated robots and AI-generated video. The narrow top is data recorded on real robot hardware. Moving up, there’s less of it, but it matches the robot’s body more closely.
NVIDIA’s own numbers show why the lower layers matter. Its team took 88 hours of in-house teleoperation data and used video-generation models to produce about 827 hours of synthetic “neural trajectories,” roughly a tenfold expansion. It also generated 780,000 simulated trajectories, which it described as equal to 6,500 hours of human demonstrations, in 11 hours of compute. The catch, which the paper acknowledges, is that synthetic data still struggles to stay physically accurate. Our GR00T explainer covers how NVIDIA’s models use each layer.
Human video is the fastest-growing layer
Figure AI has made the biggest public bet on the base of the pyramid. On August 25 it launched Index, an app that pays “Creators” to film household and workplace tasks. Figure said the app had passed 264,000 downloads in 108 countries, with more than 44,000 weekly active users, more than 16 million uploaded videos and $15 million paid to contributors. It also said it is committed to spending more than $1 billion over 12 months on data and compute. Figure wrote that it had tried buying data first, but vendors “couldn’t hit the throughput, diversity, or quality bar” it needed.
Three weeks later Figure published the first results. Its Helix 2.5 model was pretrained on Index and then sent into 30 Bay Area homes where no data had been collected, to tidy living rooms, fold towels and make beds. Holding everything else fixed, Figure says Index pretraining raised zero-shot success from 9% to 56% of trials. It also says doubling the Index data improved robot action prediction predictably enough to forecast its largest training run in advance, which it calls a human-to-robot scaling law. Index, the company says, now adds about 35 minutes of human video every second.
Those are Figure’s own evaluations, run by its own graders, and they haven’t been replicated outside the company. Even on Figure’s numbers, the robot failed nearly half of the trials, and success at three household tasks says little about a fourth. But the results explain why investors are paying for human recordings, from Mecka and others: they suggest the cheap layer of the pyramid can do much of the heavy lifting. Our Figure 03 guide tracks what the company has and hasn’t disclosed about the robot itself.
Robot fleets are turning into data factories
The top of the pyramid still needs real robots, and several companies now describe their early fleets mainly as data collectors:
- Tesla. Its Q2 2026 shareholder update said, “The initial Optimus builds will be used in our Optimus Academy for training data collection and further functionality development.” More in our Optimus status report.
- Apptronik. Its Robot Park in Austin is a data-collection facility for Apollo 2, and Apptronik says the data feeds Google DeepMind’s Gemini Robotics models. See our Apollo 2 status piece.
- Figure. It said on April 29 that it had delivered more than 350 Figure 03 robots from its BotQ plant, without saying how many are with outside customers.
Much of that robot data still comes from people steering the machines remotely. That’s normal, and it isn’t the same as autonomy, which is why we treat teleoperated footage carefully (see our guide to teleoperation versus autonomy).
What the money says, and what it doesn’t
The funding is real. Mecka’s round is confirmed by the company and its lead investor. Bigger numbers are still rumors: The Information reported on October 7 that NVIDIA discussed investing another $1 billion in Figure at about a $38 billion pre-money valuation, a report neither company has confirmed.
What nobody has shown yet is how much data a humanoid needs before it can be trusted with paid work outside a demo. Figure’s scaling law measures prediction error, not reliability on a factory floor. Mecka’s revenue claim doesn’t say who is buying or what they’ve trained. And the companies collecting footage in people’s homes haven’t said much publicly about how those recordings are reviewed and stored beyond Figure’s description of its filtering and fraud checks.
The next test comes soon. Tesla reports third-quarter results on October 21, and investors will want to know whether the Optimus Academy has moved beyond collecting data. Until a robot maker shows a model trained mostly on human video doing a paying customer’s job, week after week, the data race is a bet on a scaling curve, not proof that it holds.
Frequently asked questions
What is humanoid robot training data?
It’s the recorded examples that robot AI models learn from: camera video, joint positions, forces and the actions taken. It comes from three main places: video of humans doing tasks, simulated or AI-generated robot data, and data recorded on real robots, often while a person teleoperates them.
What is the robot “data pyramid”?
It’s a framing from NVIDIA’s GR00T N1 paper. Web data and human videos form the large base, synthetic data from simulation and video-generation models forms the middle, and a smaller amount of real-robot data sits at the top. Higher layers are scarcer but match the robot more closely.
Why are companies paying people to film household chores?
Because human video is far cheaper to collect at scale than robot time. Figure says its Index app has paid contributors $15 million, and that pretraining on that footage raised its Helix 2.5 model’s success rate in unseen homes from 9% to 56% in company-run tests. Mecka sells similar recordings to robotics labs.
How much data do robot foundation models use?
Published figures are in the thousands of hours of robot data: more than 10,000 hours for Physical Intelligence’s π0, almost 1,700 hours for Toyota Research Institute’s Large Behavior Models and about 2,500 hours for Unitree’s UnifoLM-WLA-1.0. Figure hasn’t published the total size of Index in hours.
Sources
All accessed October 8, 2026.
- Mecka, “Mecka raises $60M Series B led by Sequoia Capital” (Oct 7, 2026): mecka.ai
- Sequoia Capital, “Partnering with Mecka”: sequoiacap.com
- TechCrunch, “Robot data startup Mecka AI nabs $60M from Sequoia” (Oct 7, 2026): techcrunch.com; earlier valuation report (Sep 11, 2026): techcrunch.com
- Figure AI, “Introducing Index” (Aug 25, 2026): figure.ai
- Figure AI, “Helix 2.5: Zero-Shot 30-Home Generalization” (Sep 17, 2026): figure.ai
- Figure AI, “Ramping Figure 03 production” (Apr 29, 2026): figure.ai
- NVIDIA, “GR00T N1: An Open Foundation Model for Generalist Humanoid Robots” (arXiv 2503.14734): arxiv.org
- Open X-Embodiment Collaboration, “Open X-Embodiment: Robotic Learning Datasets and RT-X Models” (arXiv 2310.08864): arxiv.org
- Physical Intelligence, “π0: A Vision-Language-Action Flow Model for General Robot Control” (arXiv 2410.24164): arxiv.org
- Toyota Research Institute, “A Careful Examination of Large Behavior Models for Multitask Dexterous Manipulation” (arXiv 2507.05331): arxiv.org
- Unitree, UnifoLM-WLA-1.0 model card: huggingface.co
- Tesla, Q2 2026 shareholder update (Jul 22, 2026): ir.tesla.com
- Apptronik, “Welcome to Robot Park”: apptronik.com
- The Information report on NVIDIA and Figure (Oct 7, 2026; paywalled), as summarized by Börsvärlden/TIF: borsvarlden.com
Related: What is a robot foundation model? · NVIDIA GR00T explained · π0 explained · What is Diffusion Policy?
Last updated: October 8, 2026. To report an error, see our corrections page. Articles are drafted with AI assistance and reviewed and edited by an editor; see our editorial policy.
