Capture Floor
Building the floor
Think Layer · capture floor

A warehouse that makes data.

Not storage. Production. Every bay runs a real task and records it from the operator's point of view.

A capture floor in development. What follows is the design, and the first bays already shooting against it.

Scroll to walk the floor
The floor

Built to be filmed, not to hold stock.

A normal warehouse is designed around throughput. This one is designed around capture: fixed lighting, repeatable setups, and clean sightlines so the same task can be shot to spec without a crew rebuilding the scene each time.

Zones sit off one central aisle, so a crew moves between trades in minutes and footage shot in March matches footage shot in September.

Eight zones off one aisle

Everything a pair of hands can do.

Each zone is a trade with enough distinct sub tasks to shoot for weeks rather than hours. The first are running. The rest are specified and waiting on a bay.

01 Kitchen and food prep 02 Data and robotics control 03 Industrial applications 04 Construction assembly 05 Agriculture and hydroponics 06 Advanced manufacturing 07 Textiles and smart fabrics 08 Remote field robotics
Why first person

A robot never sees the room from the doorway.

Third person footage tells you that a task happened. It does not tell you what the hands did, in what order, with what grip, or where the operator was looking when they decided. Policy learning needs the second thing.

A head mounted camera sits roughly where a robot's own sensors sit. The frame moves when attention moves, so occlusion, reach and timing arrive in the shape the machine will meet them.

The rig

Worn, not mounted.

Three cameras move with the operator: one at the eye line, one on the chest for a stabilised wide, one at the wrist for grip and contact detail. A fourth is fixed to the bay for the room view. Only the floor cam belongs to the room, so the worn rig travels to any zone unchanged.

Depth, hand pose and timing come out of the same pass. Nothing is staged afterwards, and nothing is reconstructed from a single angle.

The deliverable

Footage is the raw material. The dataset is the product.

Captures ship as synchronised streams with a manifest, not as a folder of clips. You get every view on one timecode, plus the metadata needed to filter and resample without re-watching anything.

Capture
4K 60 at the head, 1080 60 at chest, wrist and floor cam
Sync
One timecode across all four streams, frame accurate
Metadata
Trade, task, operator id, take number, lighting state
Annotation
Task boundaries and failure flags as standard, hand pose on request
Delivery
Your bucket or ours, checksummed, resumable
How an engagement runs

You write the spec. We go and get it.

  1. BriefYou describe the task, the environment, and the failure cases you care about.
  2. PilotWe shoot one short session and you sign off the framing before anyone commits to volume.
  3. CaptureThe zone runs to that spec. Multiple operators, multiple takes, deliberate variation.
  4. CheckEvery session is reviewed before it leaves the floor. The ones that fail are reshot.
  5. DeliverSynchronised streams and manifest, priced per usable hour.
What gets thrown away

Most of the work is refusing to ship bad footage.

A session fails review if the head camera drifts off the work, if the hands leave frame during the critical moment, if exposure shifts mid take, or if the operator narrates instead of working. Rejected takes are reshot, not padded out.

You are billed for usable hours. That number is smaller than the number of hours recorded, and the gap is the point.

Training cells

The robot learns in the next bay over.

Footage goes straight from the bay to the cell beside it. That adjacency is the reason the floor is laid out this way. The arm attempts the same task, fails in a specific way, and the failure says which zone to shoot again and what to vary.

The floor runs as a loop rather than a pipeline.

Consent policy

No hand gets recorded without agreeing to it first.

Operators are hired to be filmed and sign a release covering the uses the footage is sold for. Nobody is filmed who has not signed one.

Sessions run inside our own space rather than in customer homes or businesses. That rules out bystanders who never agreed to anything, and keeps the premises under our control.

Where this goes

Contract capture now. A standing library after.

Today every hour is shot against a specific brief. That is the right way to start, because it forces the floor to be useful before it is large.

The end state is a library deep enough that most requests are answered from footage that already exists, and only the new task needs a shoot.

End of floor

Every task worth doing is worth recording once.

Build the floor, run the tasks, capture them, and turn the output into the training set physical AI has been missing.

This browser could not start WebGL, so the 3D floor is hidden. The page still reads top to bottom.