AI learned from the internet. The physical world was never on it.
A model can read every sentence ever published and watch every video ever
uploaded, and still not know how much force a drawer needs. Physical intelligence is not
written down anywhere; it is performed, by people, and then it is gone.
Language models
Billions of words
Already collected, already public. The corpus existed before the
models did.
Vision models
Billions of frames
Images and video scraped from an internet that was already there.
Again, found rather than made.
Physical AI
No corpus
Nothing equivalent exists. Egocentric capture of ordinary human work
has to be recorded on purpose, by someone, with consent.
The first two corpora were found. The third has to be made, and that asymmetry
is the point.
Check it yourself
This is the whole mechanism. There is nothing else behind it.
Every capture session writes one row into an append-only ledger. Eleven fields
join with a pipe, that string is hashed, and the hash of the row before it is one
of the eleven; so changing any earlier row breaks every hash after it. The record
below is an example, because a real one names a real person. The digest is not an
example: it is produced by the same function that runs in production.
The eleven fields, in the order they are joined
loading the published example…
The pre-image: those eleven, joined by |
loading…
SHA-256 of that string
loading…
That is the entire claim, and it does not require us to be honest,
only to be consistent, which is checkable. So check it: the button below
hashes the pre-image above in your browser, with the Web Crypto API your
browser ships, and compares the result to the digest we published.
Sampled every tenth frame by the pipeline, so the path is
drawn point-to-point between real measurements and nothing between them is
invented. Positions and velocities are derived; they are the
tracking model's own output, not an instrumented ground truth.
01 · Capture
One person, one camera, sixty-one seconds.
measured
A Capture Partner records the task on a head-mounted camera against a brief that fixed the task, the environment and the rig before anything was recorded. Drag the handle to move between the raw frame and the annotated one.
About this clip, before you go further
This clip predates our deployed consent flow, so it carries no consent receipt. That is the one thing wrong with the artifact this whole page is built on, and it is the next thing we fix.
Sixteen fixed windows, at their true frame positions.
derived
The recording is divided into sixteen chunks and each is given a verb-noun code. On this clip that division is a fixed window of 120 frames, not a boundary found in the footage, and the codes come from a fixed pool rather than a model. Content-aware segmentation is specified and not built. The bars below sit at their true frame positions.
Two hands, measured through sixty-one seconds of real work.
measured
Hand landmarks are computed every tenth frame — a three-hertz track, not thirty. Drag the timeline to move between sampled frames. The scrubber never interpolates: every position it shows is a row of the artifact, and there is nothing between two measurements to show.
The file, and the hash that proves it is the file.
measured
The finished record is content-addressed: rehash the annotation file you download and you get the digest below. The trajectory below is drawn complete rather than animated, because a path easing into place invents positions between the ones we measured.
Specified per brief before capture, across
6 capture domains.
Every brief states its own rig, its annotation depth and where it stands.
Hours of egocentric capture, specified
1,505
Target
Volume is agreed in a brief and recorded against it. Simulated
scenes are catalogued separately and are not counted here.
Delivery containers, each independently verified
5
Each has a writer and its own verifier, and every export is read
back off disk before it ships.
Frames annotated end to end and published in full
1,843
One real recording at 30 fps, with its annotation and its
checksum. Every engine figure on this page comes from that file.
Mean per-hand confidence
0.87
Arithmetic mean over the kinematic track. The worst value is
published beside it rather than the mean alone.
Routes in the engine behind it
173
Across
24 durable stores — the
consent ledger, the holdings register, the money book and the rest.
Catalogue
One annotation standard, across every domain we scope.
Each domain below is a capture specification you can commission, and each ships
against FH-Ego v1, so a dataloader you write once works across every
delivery.
Scene, task label, frame count, frame rate, usability verdict,
and the provenance block: clip id, content hash, cohort, QA verdict, and which
fields were derived rather than measured.
Tier 2 · Action
What happens, and when
Frame bounds per action segment, hand laterality (left / right /
both), scene semantics, camera motion, a verb-noun action code and a dense
natural-language description.
Tier 3 · Kinematics
Where the hands are
21 3D joint landmarks per hand, camera-relative wrist position,
frame-to-frame velocity, and a per-hand confidence;
so you can filter on quality rather than trust it.
A head-mounted camera sees the task from where the hands are.
Teleoperation gives you a robot's own proprioception, and it costs one person
driving one rig for every hour of demonstration it produces. Egocentric capture
gives you human hands solving the task in the environment the task actually
happens in, and it scales with people rather than with rigs. There is a real
case against it (no joint
torques, no action labels for free, a domain gap you have to close), and the
published work arguing both sides is a category argument, not a Fourthuman
result. We set it out on the method page.
We publish what reaches the person holding the camera.
The rate that reaches the person holding the camera is
published on the page addressed to them — so it is
the one cost inside a Fourthuman quote you can check before you have spoken to us.
What moves the number · 1
The environment
An Indian household kitchen and an active distribution centre are
not the same shoot. Access, scheduling and how many people have to consent all
change with the room.
What moves the number · 2
The session
How long a session runs, how many of them, and how many Capture
Partners the brief needs to reach the volume you are asking for.
What moves the number · 3
The annotation tier
Tier 1 is session metadata. Tier 3 is 21 3D joints per hand with
per-hand confidence, and it costs more to produce and more to verify. You pick
the tier in the brief.
A brief is scoped first and priced against that scope, so the number follows the
specification rather than a list. The structural reason a bespoke brief is
affordable here at all is the India–US wage differential.
Every clip passes automated face redaction server-side at ingestion, before any human review or delivery. YuNet runs where the host provides it — a DNN we chose because egocentric faces are off-axis, motion-blurred and partial, the angles a classical cascade misses. Where it does not, an older OpenCV cascade runs instead. The stage fails closed: if the detector cannot load, the clip fails the stage rather than passing through, so there is no unredacted path to delivery. The pass covers faces. What it covers, and how far, is set out in full on the redaction page. Every capture session produces a digital consent receipt persisted with a SHA-256 audit hash, revocable by the Capture Partner under the data-deletion rights in India's DPDP Act 2023. We hold no third-party attestation — no SOC 2 report, no ISO 27001 certificate — because that requires an external audit we have not commissioned. We will link the report here when we do.
We use the Fourthuman Ego Schema v1, a 3-tier annotation schema: Tier 1 (Session): task_id, scene, task_label, usability flag. Tier 2 (Action): frame bounds, hand laterality (left/right/both), scene_semantic, camera_motion, verb-noun action code, dense NL description. Tier 3 (Kinematics): 21 3D joint landmarks per hand, camera-relative Cartesian position vector P=[X,Y,Z], and frame-to-frame velocity, each hand carrying a per-hand confidence for the frame. All exported as structured JSON alongside video streams.
What data formats do you deliver datasets in?
We deliver LeRobot v2.1 (pi0/openpi and GR00T pipelines; our output loads with the real lerobot package), HDF5 (EgoDex-style paired .hdf5+.mp4, ALOHA / RoboMimic-lineage labs), and the open FH-Ego v1 JSON schema. Zarr v2 (UMI / diffusion-policy layout) and TFRecord (RLDS-style step schema; not tfds.load-compatible) are produced on request. Every export is read back by an independent verifier before delivery. We do not offer MCAP or WebDataset: MCAP is a robotics logging container no training dataloader consumes, and WebDataset is an internal sharding optimization, not a distribution format.
What makes you different from a general annotation vendor?
We work only on Physical AI, not general-purpose annotation. We combine a capture network in India — paid at a rate we publish rather than quote privately — with automated 3D annotation pipelines. The India–US wage differential is the structural cost advantage; we publish the floor a Capture Partner is paid from rather than a comparison against a US teleoperation figure we have not independently verified. Unlike general-purpose BPO annotation vendors, we own the full stack: recruitment, capture hardware, QA and privacy pipeline, and AI annotation. Delivery is LeRobot v2.1, HDF5 (EgoDex-style) and the open FH-Ego v1 JSON schema, with Zarr v2 and TFRecord (RLDS-style) on request. We don't sell raw footage; we sell annotated, privacy-cleared, format-ready training data.
How does your pricing compare to US-based teleoperation data collection?
We publish the floor a Capture Partner is paid from rather than quoting a US teleoperation benchmark we have not independently verified — the range and each brief's own rate sit in the Capture Partner app, because they depend on the brief. The India–US wage differential is the structural cost advantage and it is real; what we won't do is put a number on someone else's cost base and present it as measured. Ask for a quote against your capture spec and compare it with what you pay today. We capture value by delivering annotated, format-ready data, not raw footage.
Get started
Tell us what you’re training. We’ll scope the dataset.
Describe the policy and the embodiment you are training it for.
We scope the capture specification, price it against your volume and format,
and come back with what we can commit to, including the parts we cannot.
One of the two founders reads every one of
these.