Skip to content
Four Vectors

01 · Capture & embodiment

What we can put in your hands.

Modality, mounting and format follow your robot and your loader, not whichever rig we happen to own.

Point-cloud reconstruction of a robotic hand observed by a machine-vision camera, with gripper frame and action-space geometry
AWHAT WE RECORD

What we record.

What each setup records, and whether we can start with it today.

MODALITYSTREAMSAVAILABILITY
RGB egocentric + IMUWide-FOV stereo RGB · 9-DoF IMUIn stock kits
RGBD egocentricRGB · ToF depth · IMUIn stock kits
UMI-style handheld gripperWrist RGB · EE pose · apertureIn stock kits
Synchronised multiview RGBn × RGB, hardware triggerBuilt to spec
Multiview RGBD + tactileRGB · depth · fingertip arraysBuilt to spec
Full-body motion captureBody pose · hand keypointsBuilt to spec
What "built to spec" means
NOTE

In stock kits ships from the hardware below. Built to spec means we build the rig for your program: real and quotable, longer lead time. Multiview needs fixed cameras rigged per site; full-body needs a suit, since head-mounted kits give hand keypoints rather than body pose; and fingertip sensors change how a task feels to do, which is worth knowing before you specify tactile. Annotation schema conforms to yours.

BKITS WE RUN

Kits we run.

The rigs we have calibrated and out in the field.

DEVICEFORMKEY SPECIFICATIONSYNC
Trossen TRumiHandheld, uni- or bimanual
GoPro HERO13 · 177° FOV
Zarr / MCAP
Visual-inertial, CNC frame
GI Lab Ego1GSHead-mounted stereo
Global shutter · 2 × 1080p30
157° × 84° per camera
9-DoF IMU 400 Hz · H.265 on device
Stereo mic 16 kHz · on-device hand detection
Hardware-synced to one MCAP
Meta Quest 3Teleop
Teleop, at a customer robot cell
Hand tracking 30 Hz · 60 Hz Fast Motion
Passthrough 1280×1280 @ 60 Hz
Host-side timestamping
Why the Ego1GS
NOTE

Global shutter on a head-mounted kit is most of the reason we run this one. Rolling shutter left unmodelled makes trajectory error worse by roughly an order of magnitude, and on something a person wears you have no control over how much the head moves.

It also has a speaker, which is how the coaching in step 04 actually reaches the person: spoken, while they're working, in the app's language. Hand detection runs on the device and is off by default; we turn it on per program.

Why the Quest 3 row is different
NOTE

That row is teleoperation at your robot cell, for customers who already have arms. It isn't part of the in-the-wild model (there is no robot in a tailor's workshop) and because it's host-timestamped rather than hardware-triggered, its offsets are measured differently. Separate programme, separate spec sheet.

TRumi · GI Lab Ego1GS · Open Teach · UMI

CINTEGRATION SPEC

Five answers set the program.

Five things we need settled before we start recording, so the data loads in your stack without a re-export.

01Your embodiment
MATCH

Jaw width, stroke, finger geometry, joint ranges, and where the cameras sit, plus the transform into your end-effector frame. UMI-style capture transfers well when the aperture range matches yours and noticeably worse when it doesn't. Mounting is the parameter that moves results most, so we'd rather match yours than hand you something to domain-adapt.

02Your action space
MATCH

Absolute or delta, which frame, which units, what rate you train at. We ship in your convention and put the measured observation-to-action offset next to it.

03Your loader
MATCH

LeRobot, RLDS, MCAP, HDF5 or Zarr, agreed before collection starts. Choose after delivery and it's a re-export, which takes about a month.

04The distribution you're missing
MATCH

Tasks, environments, the edge cases you can't get to. This is what becomes the schema, and the schema is what we fill cell by cell.

05Whether you want the failures
MATCH

Most people say no. A fair number of them wish later that they hadn't. We can keep failed attempts as their own stratum instead of dropping them at capture, but once they're discarded they're gone.

DDELIVERY SPEC

What ships with every episode.

Five records travel with every episode.

ARTIFACTCONTENTS
Recording

Every stream at the agreed rate, with codec, bitrate and any re-encode written down.

Calibration record

Depot intrinsics and extrinsics, the online trace, and the board residual at the start and at the end.

Timing record

Offsets for each pair of sensors with the method named, shutter type and line readout per sensor, and the observation-to-action distribution.

Verification record

Every check, its measured value and its threshold. Signed at ingest.

Rights record

The consent artifacts, and the chain back to the contributor.