No machine should drive, dig, or harvest on a world model that can’t tell 8 meters from 12 meters.
World models are trained on billions of video frames, and almost none of them carry true metric depth. Today's monocular depth models are very good at shape, but they still guess at scale.
There's an irony here. The largest stereo imager ever built is already in everyone's pockets. It's the mobile phone.
Hammerhead Mobile harnesses those phone cameras to capture metric depth wherever a phone can go.
Hammerhead Mobile In Action
Not a rendering. This is Hammerhead Mobile running today on a standard cell phone, turning its main and ultrawide cameras into a synced stereo pair and producing dense metric depth across the frame, outdoors, in ordinary light.


200 MP + 50 MP
Main and ultrawide, synced
14 fps
Real-time stereo
Dense metric depth
Across every frame
[ ground truth ]
ImageNet labels came from people clicking.
Metric depth labels can come from geometry.
The point isn't to replace a depth sensor. The point is ground truth, captured for free.
Pair metric stereo depth with the phone's IMU and GPS and you get scaled 4D reconstructions (3D over time) from ordinary video. Use those to supervise monocular depth and world models, and the model learns metric scale it can then apply to all video, including video shot with a single camera.
[ the sensor ]
Four cameras, six possible stereo pairs
On the Samsung Galaxy S25 Ultra, the rear camera island holds four cameras, each with its own lens, resolution, and field of view.

The baseline is short, but the pixel density is extreme.
18.5 mm
Between adjacent lenses, 37 mm end to end
200 MP
Wide camera resolution
~10,900 px
Wide-camera focal length, about 5x a typical machine-vision stereo camera
How accurate is phone-based stereo?
About 0.5% error at 5 m and 2% at 20 m for the best pair.
Mid-range depth, and that's where most of the world that matters for training sits.
Stereo depth error grows with distance squared and shrinks with baseline times focal length in pixels. Assuming 0.2 px matching precision, and taking the coarser camera of each pair as the limit:
Pair
Baseline
5 m
10 m
20 m
30 m
Wide + ultrawide (full FOV)
9 cm
37 cm
1.5 m
3.3 m
Wide + 5x (center FOV)
2.5 cm
10 cm
40 cm
89 cm
Ultrawide + 5x (center FOV)
5 cm
18 cm
73 cm
1.6 m
Doubling the baseline doesn't help if one camera is coarse. The ultrawide's lower angular resolution cancels the extra length, which is why the wide + 5x pair wins.
[ why nobody does this ]
The math has been possible for years.
The problem is calibration.
Factory calibration is stale before the shutter closes. Locking OIS and AF fixes the geometry but ruins the camera.
The Problem
OIS moves the lens every frame.
Autofocus changes the effective focal length.
Mismatched cameras differ in FOV, resolution, distortion, and color response.
Rolling shutter and loose sync between modules.
Heat and flex in a 7 mm chassis.
The fix
Calibrating a phone is the same challenge as calibrating wide-baseline stereo on a flexing truck or tractor, scaled down 100x. So NODAR solves it the same way:
01
Solve the geometry from the scene itself in every frame.
02
Rectify mismatched cameras to a common virtual camera.
03
Use the extra cameras as a consistency check.
[ prototype ]
Where it started
Our first prototype ran at 3.6 fps on an inexpensive POCO phone, using its main and ultrawide cameras at 1280x720. About half the pixels returned depth; the misses were mostly bare walls and ceiling, where there's little texture to match. The hardest part was sync: the phone's own timestamps disagree slightly, by up to about half a millisecond depending on the manufacturer.
3.6 fps
Early prototype
14 fps
Today

Left: main camera. Center: depth (0.35 m to 28 m). Right: ultrawide, cropped to match.
[ how it works ]
Put the phone in the cab
Phone video has a bias. It comes from where people walk and drive: cities, homes, highways. The land that feeds the world, and the mines, forests, and job sites that build it, is nearly invisible to world models.
The fix costs about as much as a phone mount. Clip a phone inside the windshield of a tractor or construction truck, facing forward, and it becomes a metric 3D recorder for the ground the models have never seen: the same field, pass after pass, season after season.

Concept rendering: Hammerhead Mobile for Android in an excavator cab.
Any machine
Any brand, any age. No integration with the machine’s controls.
Pose
Position and orientation from the phone’s GPS and IMU.
Action labels
Every frame is time-synced to the machine’s own logs: speed, steering, implement position.
Metric 3D plus action plus outcome is about as good as world-model training data gets.
[ the system ]
Feeding the model takes two inexpensive devices
A phone on the windshield and a listen-only CAN logger on the diagnostic port, recording on the same clock.

Conceptual rendering: a horizontally mounted smartphone serves as the central data acquisition hub, connected to the excavator's vehicle network via a USB-C to CAN bus adapter.
The State
FROM THE PHONE
Metric depth, images, GPS, and IMU: what the machine saw.
The Action
FROM THE CAN LOGGER
Joystick, throttle, and travel commands: what the operator did.
At the Edge
ON THE PHONE
Heightmaps instead of photos: about 100x less data leaves the cab.
Without CAN, the model sees what happened but never learns what the operator asked for. That gap between command and response is where soil, load, and skill show up.
[ what the data buys you ]
