Phone mounted inside an excavator windshield showing a live depth map, cabled to a CAN logger

[ mobile ]

Hammerhead Mobile

The stereo camera already in your pocket.

An app that turns your phone's cameras into a calibrated stereo pair. Mount it in any cab and it records metric 3D, pose, and operator commands on one clock: ground truth for world models, captured on land they have never seen.

Phone mounted inside an excavator windshield showing a live depth map, cabled to a CAN logger

[ mobile ]

Hammerhead Mobile

The stereo camera already in your pocket.

An app that turns your phone's cameras into a calibrated stereo pair. Mount it in any cab and it records metric 3D, pose, and operator commands on one clock: ground truth for world models, captured on land they have never seen.

No machine should drive, dig, or harvest on a world model that can’t tell 8 meters from 12 meters.

World models are trained on billions of video frames, and almost none of them carry true metric depth. Today's monocular depth models are very good at shape, but they still guess at scale.

There's an irony here. The largest stereo imager ever built is already in everyone's pockets. It's the mobile phone.

Hammerhead Mobile harnesses those phone cameras to capture metric depth wherever a phone can go.

Hammerhead Mobile In Action

Not a rendering. This is Hammerhead Mobile running today on a standard cell phone, turning its main and ultrawide cameras into a synced stereo pair and producing dense metric depth across the frame, outdoors, in ordinary light.

Phone on a tripod running Hammerhead Mobile, showing a color depth map of the curved driveway and stone wall in front of it
Phone on a tripod running Hammerhead Mobile, showing a color depth map of a backyard with a playhouse, slide, and trees

200 MP + 50 MP

Main and ultrawide, synced

14 fps

Real-time stereo

Dense metric depth

Across every frame

[ ground truth ]

ImageNet labels came from people clicking.
Metric depth labels can come from geometry.

The point isn't to replace a depth sensor. The point is ground truth, captured for free.

Pair metric stereo depth with the phone's IMU and GPS and you get scaled 4D reconstructions (3D over time) from ordinary video. Use those to supervise monocular depth and world models, and the model learns metric scale it can then apply to all video, including video shot with a single camera.


[ the sensor ]

Four cameras, six possible stereo pairs

On the Samsung Galaxy S25 Ultra, the rear camera island holds four cameras, each with its own lens, resolution, and field of view.

Rear of a Samsung Galaxy S25 Ultra with each of its four cameras and the laser AF labeled


The baseline is short, but the pixel density is extreme.

18.5 mm

Between adjacent lenses, 37 mm end to end

200 MP

Wide camera resolution

~10,900 px

Wide-camera focal length, about 5x a typical machine-vision stereo camera

How accurate is phone-based stereo?

About 0.5% error at 5 m and 2% at 20 m for the best pair.
Mid-range depth, and that's where most of the world that matters for training sits.

Camera

Resolution

Pixel

Equiv. focal length

Wide

200 MP, 1/1.3"

0.6 µm

24 mm, OIS

Ultrawide

50 MP

0.7 µm

~13 mm (120°)

3x Tele

10 MP

1.12 µm

67 mm, OIS

5x Periscope

50 MP

0.7 µm

111 mm, OIS

Stereo depth error grows with distance squared and shrinks with baseline times focal length in pixels. Assuming 0.2 px matching precision, and taking the coarser camera of each pair as the limit:

Pair

Baseline

5 m

10 m

20 m

30 m

Wide + ultrawide (full FOV)

18.5 mm

−25°C to +55°C (compute unit);
−40°C to +105°C (cameras)

9 cm

37 cm

1.5 m

3.3 m

Wide + 5x (center FOV)

18.5 mm

−25°C to +55°C (compute unit);
−40°C to +105°C (cameras)

2.5 cm

10 cm

40 cm

89 cm

Ultrawide + 5x (center FOV)

37 mm

−25°C to +55°C (compute unit);
−40°C to +105°C (cameras)

5 cm

18 cm

73 cm

1.6 m

Doubling the baseline doesn't help if one camera is coarse. The ultrawide's lower angular resolution cancels the extra length, which is why the wide + 5x pair wins.

[ why nobody does this ]

The math has been possible for years.
The problem is calibration.

Factory calibration is stale before the shutter closes. Locking OIS and AF fixes the geometry but ruins the camera.

The Problem

OIS moves the lens every frame.

Autofocus changes the effective focal length.

Mismatched cameras differ in FOV, resolution, distortion, and color response.

Rolling shutter and loose sync between modules.

Heat and flex in a 7 mm chassis.

The fix

Calibrating a phone is the same challenge as calibrating wide-baseline stereo on a flexing truck or tractor, scaled down 100x. So NODAR solves it the same way:

01

Solve the geometry from the scene itself in every frame.

02

Rectify mismatched cameras to a common virtual camera.

03

Use the extra cameras as a consistency check.

[ prototype ]

Where it started

Our first prototype ran at 3.6 fps on an inexpensive POCO phone, using its main and ultrawide cameras at 1280x720. About half the pixels returned depth; the misses were mostly bare walls and ceiling, where there's little texture to match. The hardest part was sync: the phone's own timestamps disagree slightly, by up to about half a millisecond depending on the manufacturer.

3.6 fps

Early prototype

14 fps

Today

Side-by-side main camera, depth map, and ultrawide views from the early Hammerhead Mobile prototype

Left: main camera. Center: depth (0.35 m to 28 m). Right: ultrawide, cropped to match.

[ how it works ]

Put the phone in the cab

Phone video has a bias. It comes from where people walk and drive: cities, homes, highways. The land that feeds the world, and the mines, forests, and job sites that build it, is nearly invisible to world models.

The fix costs about as much as a phone mount. Clip a phone inside the windshield of a tractor or construction truck, facing forward, and it becomes a metric 3D recorder for the ground the models have never seen: the same field, pass after pass, season after season.

Concept rendering of a phone mounted inside an excavator windshield, cabled to the cab’s data port

Concept rendering: Hammerhead Mobile for Android in an excavator cab.

Any machine

Any brand, any age. No integration with the machine’s controls.

Pose

Position and orientation from the phone’s GPS and IMU.

Action labels

Every frame is time-synced to the machine’s own logs: speed, steering, implement position.

Metric 3D plus action plus outcome is about as good as world-model training data gets.

[ the system ]

Feeding the model takes two inexpensive devices

A phone on the windshield and a listen-only CAN logger on the diagnostic port, recording on the same clock.

Labeled diagram of a phone mounted in an excavator cab and cabled to the CAN port, with an inset of the app’s depth, IMU, GPS, and CAN data screen

Conceptual rendering: a horizontally mounted smartphone serves as the central data acquisition hub, connected to the excavator's vehicle network via a USB-C to CAN bus adapter.

The State

FROM THE PHONE

Metric depth, images, GPS, and IMU: what the machine saw.

The Action

FROM THE CAN LOGGER

Joystick, throttle, and travel commands: what the operator did.

At the Edge

ON THE PHONE

Heightmaps instead of photos: about 100x less data leaves the cab.

Without CAN, the model sees what happened but never learns what the operator asked for. That gap between command and response is where soil, load, and skill show up.

[ what the data buys you ]

A world model is a learned simulator.

Show it the scene and an action, and it predicts what happens next. For an excavator loading a truck, the scene is the terrain around the machine, and the action is how the operator moves the joysticks.

Consequence prediction

Predicts the outcome before the move: "Dig here and the bucket fills about 90%." "This pass will stall the engine."

Ranking options

Simulates several candidate motions and suggests the best one, with a reason the operator can understand.

Coaching

Learns from your best operators and coaches everyone else in real time, while the human stays in control.

"If the bucket enters here at this angle, how full will it be, and will the engine bog down?"

That is worth money long before full autonomy.

Operator skill varies enormously: one Volvo CE study measured up to 3x the productivity between occasional and professional operators.

[ the fleet math ]

How much data does it take?

Less than you might think, and a modest fleet gets you there in one season.

~1,500 hours

Per machine, per year

✕

10 machines

A phone in each cab

=

15,000+ hours

Labeled data, one season

goal

DATA NEEDED

Proof of concept: one machine type, one task

50 to 200 hours

−25°C to +55°C (compute unit);
−40°C to +105°C (cameras)

Useful operator assist across sites, soils, and operators

1,000 to 5,000 hours

−25°C to +55°C (compute unit);
−40°C to +105°C (cameras)

General world model across machine types and tasks

10,000 to 100,000+ hours

−25°C to +55°C (compute unit);
−40°C to +105°C (cameras)

Rough planning numbers, not established requirements.

Volume isn't the bottleneck. Diversity is: different operators, soils, weather, and job types. Fifty machines on twenty sites beats five machines on one.

If you run a fleet working land the world models have never seen, a phone mount in every cab may be the cheapest data program you’ll ever start.

More questions? Check out the FAQ.