Robot manipulation data · India
We record the work robots need to learn from.
Four Vectors collects robot manipulation data in working Indian shops, kitchens and workshops. We ship the kit to the room, coach the recording while it's happening, and run twenty-five checks on every episode before it leaves us.
Almost all of it comes from a handful of buildings.
Manipulation data gets collected wherever someone can stand next to the rig and keep an eye on it. Which in practice means a lab.
| CORPUS | VOLUME | REAL DENOMINATOR | |
|---|---|---|---|
| 01 | AgiBot World | 1,001,552 trajectories | One 4,000 m² hall |
| 02 | DROID | 76,000 trajectories · 564 scenes | 52 buildings, mostly labs |
| 03 | BridgeData V2 | 60,096 trajectories | 24 environments, seven toy kitchens |
The trouble starts when you try to record anywhere else.EVIDENCE→
Ego-Exo4D was run by expert teams across thirteen cities, and published its own tracking numbers. About one recording in twenty-four came back partly or wholly untracked: 95.9% clean, 3.5% partial failure, 0.6% failed outright. That's with trained people and scenarios they controlled. It's the good case.
What we do instead.
Coach while it's recording. Measure everything once it's back.
Your tasks become a grid with a target per cell
A contributor books a kit in the app
In-person pickup, with a guided calibration check
Live tracking health, drift, framing by voice
Every episode runs the register
Per episode, with a floor under it
Why the rooms are worth it.
The evidence on which variety improves a policy isn't the answer most people assume.
| AXIS | EFFECT | TRIALS | VERDICT |
|---|---|---|---|
| Spatial arrangement | 10% → 70% | n=20, real | Critical |
| Camera pose | 17% → 90% | n=30, sim | Critical |
| Spatial arrangement | 10% → 57% | n=30, sim | Same effect, in sim |
| Object texture | 93% | n=30, sim | Negligible |
A separate study across more than 15,000 real robot trials found the same thing from the other direction: environment and object variety mattered far more than the sheer number of demonstrations, and the returns flattened out at around fifty demonstrations per environment.
The simulation rows are single-seed at thirty trials, so we've rounded them; a second decimal would be noise. They're here because they agree with the hardware result, not because they'd carry the argument on their own.
Which is why we don't build a program around exotic objects.WHAT'S SCARCE→
A counter that's too low. Work done sitting on the floor. A galley kitchen where the wall decides how far you can reach. Camera angles set by where there's room to stand rather than by a rig.
What Matters in Learning from Large-Scale Datasets · Data Scaling Laws in Imitation Learning
What we don't claim.
Egocentric human video is already globalCORRECTION→
It more or less is. Ego4D covers 3,670 hours from 74 locations across nine countries, India included: IIIT Hyderabad alone recorded over 130 participants at 25 sites. So we're not going to tell you nobody has filmed India. What's actually missing is manipulation data with calibrated poses, gripper state, action timing you can trust, and a check run on every episode.
The biggest action datasets aren't AmericanCORRECTION→
AgiBot World and ARIO are both Chinese, and both bigger by trajectory count than anything out of the Bay Area. We used to build the pitch around geography. It didn't survive contact with the numbers, so we dropped it. What we think actually matters is whether the room was there before the camera was.
Working this way is harder, not easierCORRECTION→
A lab can fix a bad take while the person is still standing there. We can't. Whatever goes wrong, we find out after the kit comes back, which means the checks have to be good enough to carry the whole thing. If you don't believe they are, not much else here matters, and that's why the register is the page we've written out in full.