A warehouse that makes data.
Not storage. Production. Every bay runs a real task and records it from the operator's point of view.
A capture floor in development. What follows is the design, and the first bays already shooting against it.
Built to be filmed, not to hold stock.
A normal warehouse is designed around throughput. This one is designed around capture: fixed lighting, repeatable setups, and clean sightlines so the same task can be shot to spec without a crew rebuilding the scene each time.
Zones sit off one central aisle, so a crew moves between trades in minutes and footage shot in March matches footage shot in September.
Everything a pair of hands can do.
Each zone is a trade with enough distinct sub tasks to shoot for weeks rather than hours. The first are running. The rest are specified and waiting on a bay.
A robot never sees the room from the doorway.
Third person footage tells you that a task happened. It does not tell you what the hands did, in what order, with what grip, or where the operator was looking when they decided. Policy learning needs the second thing.
A head mounted camera sits roughly where a robot's own sensors sit. The frame moves when attention moves, so occlusion, reach and timing arrive in the shape the machine will meet them.
Worn, not mounted.
Three cameras move with the operator: one at the eye line, one on the chest for a stabilised wide, one at the wrist for grip and contact detail. A fourth is fixed to the bay for the room view. Only the floor cam belongs to the room, so the worn rig travels to any zone unchanged.
Depth, hand pose and timing come out of the same pass. Nothing is staged afterwards, and nothing is reconstructed from a single angle.
Footage is the raw material. The dataset is the product.
Captures ship as synchronised streams with a manifest, not as a folder of clips. You get every view on one timecode, plus the metadata needed to filter and resample without re-watching anything.
- Capture
- 4K 60 at the head, 1080 60 at chest, wrist and floor cam
- Sync
- One timecode across all four streams, frame accurate
- Metadata
- Trade, task, operator id, take number, lighting state
- Annotation
- Task boundaries and failure flags as standard, hand pose on request
- Delivery
- Your bucket or ours, checksummed, resumable
You write the spec. We go and get it.
- BriefYou describe the task, the environment, and the failure cases you care about.
- PilotWe shoot one short session and you sign off the framing before anyone commits to volume.
- CaptureThe zone runs to that spec. Multiple operators, multiple takes, deliberate variation.
- CheckEvery session is reviewed before it leaves the floor. The ones that fail are reshot.
- DeliverSynchronised streams and manifest, priced per usable hour.
Most of the work is refusing to ship bad footage.
A session fails review if the head camera drifts off the work, if the hands leave frame during the critical moment, if exposure shifts mid take, or if the operator narrates instead of working. Rejected takes are reshot, not padded out.
You are billed for usable hours. That number is smaller than the number of hours recorded, and the gap is the point.
The robot learns in the next bay over.
Footage goes straight from the bay to the cell beside it. That adjacency is the reason the floor is laid out this way. The arm attempts the same task, fails in a specific way, and the failure says which zone to shoot again and what to vary.
The floor runs as a loop rather than a pipeline.
No hand gets recorded without agreeing to it first.
Operators are hired to be filmed and sign a release covering the uses the footage is sold for. Nobody is filmed who has not signed one.
Sessions run inside our own space rather than in customer homes or businesses. That rules out bystanders who never agreed to anything, and keeps the premises under our control.
Contract capture now. A standing library after.
Today every hour is shot against a specific brief. That is the right way to start, because it forces the floor to be useful before it is large.
The end state is a library deep enough that most requests are answered from footage that already exists, and only the new task needs a shoot.
Every task worth doing is worth recording once.
Build the floor, run the tasks, capture them, and turn the output into the training set physical AI has been missing.
This browser could not start WebGL, so the 3D floor is hidden. The page still reads top to bottom.