The handheld capture rig: a 360° camera above a gripper with yellow jaws, on a handle

Handheld capture system & data pipeline

Ground truth
you can carry.

A handheld 360° rig and the pipeline behind it. Neural VIO holds millimetre-level pose for six hours at a stretch, every stream lands on one clock within a millisecond, and what comes out is a metric 3D world with actions, calibration and open-set labels already attached.

loading capture…
CH.01Tracking

Neural VIO that stays
in one piece.

Our own feedforward VIO tracks the rig from the first frame to the last — continuously, with no reset and no external tracker. It holds global metric scale for the whole session: a metre in the first minute is a metre in the sixth hour, in one room or across a whole building.

It keeps that accuracy through the conditions that end a classical run — people crossing the frame, the gripper filling the view, a doorway blowing out the exposure, a light switched off mid-take. No markers, no motion capture, nothing installed in the room.

Pose resolution
mm-level
Continuous session
6 h+
Scale
global metric
recording_20260523_230639
CH.02Time base

One clock.
Every stream.

Every sensor on the rig lands on a single time base — both camera bodies, the IMU, the gripper encoder, the microphones. Our synchronisation works across separate devices that never share a cable and across sensors running at completely different rates.

Theoretical alignment error is under a millisecond: a twentieth of a camera frame at 50 fps, a fifth of an IMU sample at 200 Hz. It holds for the length of the session rather than only at the start, and every session is checked before it leaves the pipeline.

Alignment
< 1 ms
At 50 fps
1/20 frame
Scope
cross-device
five streams · common time base alignment Δt 0.40 ms
CH.03Modalities

Everything a policy needs,
from one pass of the room.

Metric 3D environment

Dense coloured point cloud in real metres, with a pose for every frame. Global metric scale, held across the whole session.

Annotated fisheye frame from a capture, with detection boxes

Dual 200° fisheye

4K per lens at 50 fps, back to back — the full sphere, including whatever is behind the operator.

IMU at 200 Hz

Accelerometer and gyroscope, calibrated per device and shipped with the extrinsic to the camera it sits behind.

7-DoF end-effector pose

Position and orientation of the gripper at 50 Hz in the same metric frame as the map, with the jaw angle alongside — a demonstration you can replay.

Audio, 48 kHz

Stereo sound for contact events and speech, on the same time base as every other stream.

session#245
rigrig 2 · 2 clips
started2026-06-17T19:11:50
statesynced · cooked
operatoron record

Session metadata

Who shot it, on which rig, when, where it got to in the pipeline — carried with the clip instead of living in a spreadsheet.

Intrinsics & extrinsics

Every capture ships with its own calibration — per lens, per body, per rig. Nothing is assumed shared.

modelfisheye · 200°
fx, fy206.30, 204.88
cx, cy256.0, 256.0
Ticper device, factory
biasgyro + accel, logged
CH.04Labels

Labels that survive
contact with the real world.

In-the-wild footage doesn't fit a fixed vocabulary. Every clip is decomposed by a vision-language model into a task, ordered minitasks, and per-moment affordances grounded back onto the frame — then a reviewer confirms it. Below is one real capture with its real annotation, playing.

pass 2
cam R · time-aligned
next action

0.0 / 31.0 s
Progress

How far through the task this moment is — not just what is happening in it.

Affordance

Where to act and which way the action goes, tracked live against the clip.

Open-set objects

Free-form targets written by the model, not picked from a closed class list.

CH.05Loop

Eval is not the last step.
It is the instrument.

There is no closed form between a data distribution and how a policy behaves on a real robot. It can only be measured. So the measurement sits inside the line rather than at the end of it: a batch trains the moment it is complete, only the parameters that changed go back out to the rigs, and the operator sees what the last hour did to the policy before the next hour starts.

The gap between "the model fails at this" and "we are out collecting exactly that" closes in hours instead of weeks. Nobody collects blind, and nothing is judged three weeks late by a training run that diverged.

Loop time
hours
Per operator-day
6 h
Ships with
a checkpoint
collect · train · update · verify
Ten rounds to three

Engineering compresses how many rounds of exploration a capability takes. It cannot compress that to one — the last stretch has no analytic solution, only measurement. What it can do is make every round fast, cheap and repeatable.

Checked where it is captured

A batch that will not help is caught on site, at the moment it is shot, by the model it was meant to improve. The rework, the re-shoots and the discarded hours go away with it.

Every batch ships with proof

The dataset, the training recipe, the checkpoint trained on it and that checkpoint's real-robot eval report — so what arrives is a measured result, not hours of footage that should help.

three weeks on one clock · stage split illustrative
CH.06Rig
The gripper jaws open, seen from the front
Two rigs standing side by side
A hand holding the rig by its handle

Built to be
carried all day.

Weight
200 g
Runtime
unlimited
Storage
1 TB
Battery
Swappable — change it mid-shift and keep going
Handles
A 1 mm sheet up to a 20 kg box, two-handed
Coverage
About 90% of everyday manipulation work
Camera
2 × 360° body · 200° lenses, back to back
Video
4K per lens, 50 fps
Inertial
200 Hz, factory cam–IMU calibration on board
Action
Encoder-instrumented gripper, jaw angle logged
Audio
48 kHz stereo
Offload
Portable SSD → hub, mirrored to object storage
Fleet today
38 devices · 19 rigs

Access

Tell us what you're training.
We'll tell you what we can capture.

Datasets, rigs, or the whole pipeline running inside your own infrastructure — start with the task you're stuck on.