| Myth | Reality |
|---|---|
| A few hundred demonstrations are enough to train a reliable robot skill | Published scaling studies show success rates continuing to climb well past a thousand demonstrations per task before plateauing |
| Any human demonstration is equally useful training data | Data diversity across objects, backgrounds, and operator style matters more than raw episode count once a baseline volume is reached |
| Teleoperation always requires an actual robot to be present | Newer “robot-free” methods capture human hand and body motion directly, without a physical robot in the loop, then retarget it afterward |
| Data collection is a one-time cost before training starts | Production pipelines run continuous data collection and re-collection cycles to patch specific failure modes discovered after deployment |
Why the Data Pipeline Is the Real Bottleneck
Every headline about a new robot foundation model eventually traces back to the same unglamorous input: hours of human demonstration data. Vision-language-action policies, imitation-learning controllers, and diffusion-based manipulation models are all trained primarily on recorded examples of a human operator performing a task, not on hand-written rules or reward functions. That makes the process of collecting, labeling, and curating demonstration data one of the most consequential and most under-discussed parts of humanoid robotics, and the actual economics of that process, hours needed, cost per hour, diversity requirements, are rarely covered with any numbers attached.
How Teleoperation Data Collection Actually Works
VR and Motion-Controller Rigs
The most common current setup uses consumer or near-consumer VR hardware, headsets such as Meta Quest 3 or Apple Vision Pro, paired with hand-tracking controllers, to let an operator control a robot’s end effectors in real time while the system logs synchronized camera frames, joint states, and the operator’s commanded motion as one training episode. This approach reached broad adoption because it repurposes off-the-shelf consumer hardware with mature six-degree-of-freedom tracking, rather than requiring custom-built teleoperation hardware for every lab.
Exoskeletons and Full-Body Motion Capture
For humanoid tasks that require coordinated whole-body motion rather than just arm and hand control, teams use wearable exoskeleton rigs or full-body motion-capture suits that sample kinematic data at sampling rates above 100 Hz, capturing the operator’s entire body trajectory, not just hand position, so the resulting robot policy can learn coordinated stepping, reaching, and balance adjustments together rather than as separate subsystems.
Robot-Free Egocentric Capture
A newer and increasingly important category skips the physical robot entirely during data collection. Systems in this category have an operator wear a lightweight headset and a handheld or glove-mounted gripper analog, capturing sparse human keypoint trajectories and wrist-view camera footage as the person performs the task directly with their own hands, then retarget that recorded motion onto a robot’s kinematics afterward in software. This class of method, sometimes described under names like UMI-style capture, exists specifically because physical robot teleoperation is constrained by hardware availability, operator expertise, and the sheer throughput limit of only being able to record as fast as a robot can physically move; a human demonstrating with their bare hands can record many more episodes per hour, in more locations, without monopolizing scarce robot hardware.
| Collection method | Hardware required | Throughput characteristic | Best suited for |
|---|---|---|---|
| VR controller teleoperation | Consumer VR headset, hand controllers, physical robot present | Moderate; limited by robot execution speed | Arm and gripper manipulation tasks |
| Exoskeleton / full-body mocap | Wearable exoskeleton or motion-capture suit, physical robot present | Lower; rigging and calibration overhead per session | Whole-body humanoid loco-manipulation |
| Robot-free egocentric capture | Lightweight headset, wrist camera, handheld gripper analog, no robot required | High; many episodes per hour, portable to any location | Scaling data volume and environment diversity cheaply |
How Much Data Is Actually Needed
Published data-scaling studies in imitation learning for robotic manipulation show a pattern that will be familiar from other machine-learning domains: task success rate improves steadily with more demonstrations, with diminishing but still meaningful returns well beyond what early robotics research assumed was sufficient. Where early academic demonstrations often trained narrow single-task policies on a few hundred episodes, current foundation-model-style policies intended to generalize across many objects and instructions are trained on datasets combining hundreds of thousands of episodes across many tasks and embodiments, with OpenVLA’s public training set alone comprising roughly 970,000 real-world demonstration episodes drawn from a broad open robotics dataset collaboration.
The practical planning number most teams work with for a single new task on a single robot embodiment, rather than a broad foundation model, tends to fall in the range of several hundred to a few thousand demonstration episodes to reach a reliable, deployable success rate, with the exact number depending heavily on how visually and physically varied the task’s object set and environment are. A simple, constrained pick-and-place task with a fixed object needs meaningfully less data than an open-vocabulary task expected to generalize to novel objects never seen during collection.
Diminishing returns in demonstration scaling
Task success rate climbs steeply from near-zero up through the first few hundred demonstrations, continues rising more gradually through the low thousands, and then flattens into a long tail where additional episodes mostly help with rare edge cases and out-of-distribution objects rather than raising the average success rate.
The Cost Side Nobody Puts a Number On
Collecting a single demonstration episode is not free, and the true cost includes far more than the operator’s hourly wage. A teleoperation session requires the physical robot to be reserved and often physically reset between episodes, an operator trained on the specific interface, a supervisor or engineer to flag and discard corrupted or unsafe episodes, and downstream labeling or annotation work to tag episodes with the language instruction, success or failure outcome, and any segmentation needed for the model’s training pipeline. Robot-free egocentric methods reduce several of these costs simultaneously, since no physical robot needs to be reserved or reset, and an operator can record many short episodes back to back in ordinary environments rather than a calibrated lab cell, which is a major part of why the field has been moving toward these methods as a way to scale data volume without scaling robot fleet size in lockstep.
| Cost driver | Physical-robot teleoperation | Robot-free egocentric capture |
|---|---|---|
| Hardware reservation | Requires dedicated robot and workcell time | Requires only headset and handheld gripper analog |
| Operator throughput per hour | Bounded by robot execution speed and reset time | Bounded mainly by human task-performance speed |
| Environment diversity achievable | Limited to available robot workcells | Can be recorded in varied real-world locations |
| Post-processing requirement | Lower; actions already in robot action space | Higher; requires retargeting human motion onto robot kinematics |
Data Quality and Diversity Requirements
Volume alone does not produce a reliable policy. Data-scaling research consistently finds that diversity across object instances, background clutter, lighting, and even operator demonstration style matters as much as sheer episode count, because a policy trained on visually narrow data learns to rely on incidental cues, a specific tabletop color, a specific camera angle, rather than the task-relevant structure. Teams building serious demonstration datasets deliberately vary the object set, camera placement, and environment across collection sessions specifically to prevent the model from overfitting to collection-studio artifacts that will not be present at deployment.
Labeling considerations add another quality dimension: episodes need accurate language instruction pairing, since the instruction text is itself part of the training signal for vision-language-action models, and episodes containing operator errors, collisions, or ambiguous outcomes need to be flagged and either excluded or explicitly labeled as failures, since silently including failed demonstrations as if they were successes actively teaches the model the wrong behavior.
Common mistake
Teams under time pressure sometimes pad out a demonstration dataset with many near-duplicate episodes, the same object, same background, same operator, recorded dozens of times, to hit an episode-count target quickly. This inflates the raw number without improving generalization, since the model has still only seen one real scenario repeated many times. A smaller dataset with genuine variation in objects, lighting, and camera angle consistently outperforms a larger but repetitive one on held-out evaluation.
What worked
Mixing robot-free egocentric human demonstrations with a smaller core set of physical-robot teleoperation episodes let teams scale total data volume and environment diversity quickly through the cheap egocentric channel, while still anchoring the policy in accurate robot-specific action data through the smaller, higher-fidelity physical-robot set. Neither channel alone matched the result of combining both.
Companies Running Large-Scale Teleoperation Data Programs
Several organizations have built dedicated infrastructure and, in some cases, paid human operator programs specifically to generate demonstration data at scale, treating it as core infrastructure rather than a one-off research exercise. This includes robotics labs building open, cross-embodiment demonstration datasets through multi-institution collaborations, humanoid robot companies running in-house teleoperation studios staffed by trained human demonstrators, and foundation-model developers such as Physical Intelligence collecting cross-embodiment data specifically to train broadly generalizing action-expert policies like the pi-0 family. The common thread across all of these programs is that demonstration data collection has shifted from an academic afterthought into a standing operational function, with its own staffing, tooling, and quality-control processes, much like data labeling became a standing function in the earlier scaling of large language models.
Frequently Overlooked Details in Data Collection Programs
- Episode metadata matters as much as the trajectoryAccurate language instruction labels, success or failure tags, and camera calibration metadata are as critical to training value as the raw motion data itself.
- Operator style variance is a feature, not noiseCollecting demonstrations from multiple operators with different movement styles helps a policy generalize better than data from one highly consistent operator.
- Failure episodes have training valueExplicitly labeled failed or corrected episodes can teach a policy what not to do, but only if they are flagged rather than silently mixed in as if successful.
- Retargeting error is an underappreciated noise sourceRobot-free human motion capture must be mathematically retargeted onto a robot’s different kinematics, and retargeting inaccuracies introduce systematic errors that are easy to miss during review.
- Camera viewpoint must match deploymentData collected from a camera angle that will differ from the deployed robot’s actual camera placement quietly degrades real-world performance.
- Re-collection after deployment is a standing costProduction teams budget for ongoing targeted data collection to patch specific failure modes discovered once a policy is already deployed, not just a single upfront collection phase.
Glossary
- Demonstration episode
- One complete recorded instance of a human operator, directly or through teleoperation, performing a task, including the synchronized camera frames, instruction, and resulting action sequence.
- Teleoperation
- Real-time human control of a physical robot’s movements through an interface such as VR controllers or a motion-capture rig, used to generate training demonstrations.
- Retargeting
- The process of mathematically converting recorded human motion or hand poses into the equivalent motion for a robot’s different physical kinematics, used in robot-free data collection.
- Cross-embodiment dataset
- A demonstration dataset combining episodes collected across multiple different robot bodies, used to train policies that transfer skills across hardware more efficiently.
- Data scaling law (imitation learning)
- The observed relationship between the number of demonstration episodes used in training and the resulting task success rate, typically showing steep early gains followed by diminishing returns.
Key Takeaways
- Robot policies are trained primarily on recorded human demonstration episodes, not hand-written rules, making data collection a core part of the robotics pipeline.
- VR controller teleoperation, exoskeleton or full-body motion capture, and robot-free egocentric capture are the three main current data-collection approaches.
- Public datasets like OpenVLA’s roughly 970,000 real-world demonstration episodes illustrate the scale foundation-model-style policies now train on.
- A single new task on one robot embodiment typically needs several hundred to a few thousand demonstrations to reach a reliable success rate, depending on object and environment variety.
- Robot-free egocentric capture methods scale data volume and environment diversity faster and more cheaply than physical-robot teleoperation, since no robot hardware needs to be reserved.
- Data diversity across objects, lighting, and operator style matters as much as raw episode count; padding a dataset with near-duplicate episodes does not improve generalization.
- Several organizations, including cross-institution open-dataset collaborations and dedicated foundation-model companies, now run standing teleoperation data programs as core infrastructure.
FAQs
What is teleoperation data collection in robotics?
It is the process of recording a human operator controlling a robot, or performing a task with their own hands using motion-capture hardware, so that the synchronized camera footage, instruction, and resulting motion can be used as training data for a robot learning policy.
How much demonstration data does a robot policy need?
A single new task on one robot typically needs several hundred to a few thousand demonstration episodes for reliable performance, while broad foundation-model-style policies intended to generalize across many tasks train on datasets with hundreds of thousands of episodes.
What hardware is used to collect robot teleoperation data?
Common hardware includes consumer VR headsets and hand controllers for arm and gripper tasks, wearable exoskeletons or motion-capture suits for whole-body humanoid tasks, and lightweight headset-plus-gripper rigs for robot-free egocentric human demonstration capture.
What is robot-free data collection?
It is a method where a human performs a task directly with their own hands while wearing tracking hardware, without any physical robot present, and the recorded human motion is mathematically retargeted onto a robot’s kinematics afterward, allowing much higher data-collection throughput than physical teleoperation.
Does more demonstration data always improve a robot policy?
Success rate generally improves with more data but with diminishing returns, and data diversity across objects, environments, and operator style matters as much as raw episode count; a large but repetitive dataset underperforms a smaller, more varied one.
Why do companies keep collecting data after a robot policy is already deployed?
Deployment reveals specific real-world failure modes that were not represented in the original training data, so teams run targeted re-collection to gather demonstrations that directly address those observed failures rather than relying solely on the original dataset.
What does data labeling involve for robot demonstrations?
It involves pairing each episode with an accurate language instruction, tagging whether the episode was a success or failure, and verifying camera calibration and synchronization metadata, since inaccurate labels actively teach the model incorrect associations.
Which organizations run large-scale robot demonstration data programs?
Cross-institution open-dataset collaborations, in-house teleoperation studios run by humanoid robot companies, and foundation-model developers such as Physical Intelligence, which trains cross-embodiment action-expert models, all run standing data-collection operations rather than one-time collection exercises.
The models trained on this data are examined in depth in our piece on vision-language-action models. Why hand dexterity remains hard to teach even with abundant demonstration data is covered in dexterous manipulation. How simulation-trained policies still need real demonstration data to close remaining gaps is discussed in sim-to-real transfer. For the underlying foundation-model landscape these datasets feed, see robot foundation models explained. Readers interested in how this data pipeline scales across deployed robot fleets can also see robot fleet learning.
- arXiv, “Data Scaling Laws in Imitation Learning for Robotic Manipulation”
- arXiv, “HumanoidUMI: Bridging Robot-Free Demonstrations and Humanoid Whole-Body Manipulation”
- arXiv, “BifrostUMI: Bridging Robot-Free Demonstrations and Humanoid Whole-Body Manipulation”
- arXiv, “EgoHumanoid: Unlocking In-the-Wild Loco-Manipulation with Robot-Free Egocentric Demonstration”
- arXiv, “UMIGen: A Unified Framework for Egocentric Point Cloud Generation and Cross-Embodiment Robotic Imitation Learning”
- EVS Int, “Embodied AI Data Collection: Teleoperation Guide (2026)”
