Robots learn from what people do.
We turn it into training data at scale.
Terra is Labelbox’s data engine for robotics: egocentric and teleoperated demonstrations captured on purpose-built rigs, annotated and validated by an AI diversity engine, and delivered from pre-training through evals.
- 1PB+produced as of January 2026
- 7capture rig configurations
- Up to 4synced camera streams per episode
- Pre-training → evalsone pipeline, one episode format
Physical AI is blocked by data, not models
Language models had the internet. Robots have nothing of the kind: every demonstration has to be recorded on purpose, in the real world, by a person. Whoever solves that at scale, with the variety the world actually has, sets the pace for the field.
- 01
Scale
One task needs thousands of demonstrations to generalize. No lab can record that in-house.
- 02
Diversity
Ten thousand episodes in one kitchen is one episode. Coverage across environments, objects, operators and lighting is the whole point.
- 03
Ground truth
Raw video is not training data. Actions, poses, contacts and success labels have to be right, or the model learns the noise.
One pipeline, pre-training to evals
Every rig writes the same episode format, so the data that pre-trains a model is the same shape as the demonstrations that post-train it and the episodes that evaluate it.
Pre-training
Diverse egocentric and third-person video across environments, objects and tasks, with auto-labels for scene, object and action segments.
Post-training
Expert teleoperation and instrumented-gripper demonstrations with synchronized action streams and success labels, for imitation learning and RL fine-tuning.
Evals
Standardized manipulation and navigation tasks with scoring rubrics, run against your policy and reported per environment.
Seven ways to see a task
Every rig is hardware-synchronized, ships to the operator pre-configured, and writes to the same episode format. Pick the viewpoints your policy needs and mix them in one program.
| Rig | Cameras | Modality | Best for |
|---|---|---|---|
| Ego mono | 1 head-mounted | RGB | Scene understanding, pre-training |
| Ego stereo | 2 head-mounted, synced | RGB + depth | 3D scene, navigation |
| Ego trio | 3 head + peripheral | RGB wide FOV | Full-context manipulation |
| Ego + wrist | Head + 1–2 wrist | RGB close-up | Grasping, fine manipulation |
| Trio + gripper | 3 head + instrumented gripper | RGB + gripper state | Action-labelled demonstrations |
| Overhead | 1–2 fixed above the workspace | RGB | Bird’s-eye view, bimanual tracking |
| Teleoperation | Robot-mounted + operator | Joint states, actions, video | Post-training, imitation learning |
An engine that knows what it has not seen yet
Every episode is ingested, auto-categorized by environment, object, task and operator, validated, and counted against a target distribution. Where coverage is thin, collection is steered there next.
- 1Ingest · Raw streams land from operators worldwide
- 2Categorize · Environment, objects, task and operator style auto-tagged
- 3Validate · Automated checks, human review on edge cases
- 4Steer · Gaps found, collection redirected to fill them
Five episodes from the corpus
Real footage from Terra rigs across tasks, sensors and environments. Bimanual and gripper-instrumented episodes first; egocentric captures follow.
Stacking plates
Packing garments
Washing dishes
Groceries
Folding clothes
Dwarkesh visits the San Francisco robotics lab
Where rigs are built, operators are trained and episodes are validated before they ship. A first-hand look at how Terra data gets made.
Questions robotics teams ask us
Does Labelbox build robots?
No. Terra produces the data robots learn from and the evals that measure them. We are embodiment-agnostic and work with your hardware or ours.
What do you deliver?
Synchronized multi-camera video, action and state streams, annotations and success labels, in the formats your training stack expects.
Who records the data?
Collection runs through the Alignerr network and specialist third-party partners, on rigs we supply, in homes, workplaces and our San Francisco robotics lab. Labelbox specializes in what happens next: turning that raw footage into validated, training-ready datasets with our technology.
Can we commission a task?
Yes. Most programs start with a task list and a target distribution; the diversity engine steers collection until coverage is met.
How is quality checked?
Every episode is auto-categorized and validated, with human review on edge cases. Episodes that fail are re-recorded, not shipped.
Tell us what your robots need to learn
Share the tasks and environments you need covered. We will scope the program and show you the first episodes.