|

Ground Truth

|

3D Vision Brief: Metric Ground Truth at Internet Scale

Leaf Jiang

CEO

The stereo camera is already in your pocket.

World models are trained on billions of video frames, and almost none of them carry true metric depth. Today's monocular depth models are very good at shape, but they still guess at scale. A model that can't tell 8 m from 12 m isn't ready to drive, dig, or harvest.

The data has a second gap. It comes from where people walk and drive: cities, homes, highways. The land that feeds the world, and the mines, forests, and job sites that build it, are nearly invisible to world models.

There's an irony here. The largest stereo imager ever built is already deployed in the billions. It's the phone.

Put the phone in the cab

Clip a phone inside the windshield of a tractor or excavator, facing forward, and its rear cameras become a metric 3D recorder for ground truth the models have never seen: the same field, pass after pass, season after season. Phone stereo vision works from about 1 to 30 m, the range that matters directly ahead of a working machine.

Add a listen-only CAN logger on the diagnostic port, and every frame also gets an action label: the joystick, throttle, steering, and travel commands the operator gave, on the same clock as the images. The phone's GPS and IMU supply pose. The logger records raw CAN on almost any machine with a diagnostic port, with no integration. Decoding it is easiest for the OEM, which already has the message definitions.

Below is a concept rendering of NODAR Hammerhead Mobile, an app we're developing to turn any fleet into a source of metric ground truth. It runs our stereo pipeline on the phone, reads the CAN logger, and syncs the two on a single clock.


Setup: Hammerhead Mobile on Android in an excavator cab. The rear cameras capture terrain in 3D, RGB video, IMU, and GPS, while a USB-to-CAN adapter records machine state and joystick commands on the same clock.

From data to a better operator

A world model is a learned simulator: show it the scene and an action, and it predicts what happens next. For an excavator loading a truck, the scene is the terrain around the machine, and the action is how the operator moves the joysticks. A good model can answer "if the bucket enters here at this angle, how full will it be, and will the engine bog down?" On a tractor, the same kind of model learns how commanded speed and implement depth play out as wheel slip, ruts, and draft load in a given soil.

That is worth money long before full autonomy. Operator skill varies enormously: one Volvo CE study measured up to 3x the productivity between occasional and professional operators. A model that has learned from the best operators can coach everyone else in real time, with suggested dig points, fill estimates, and passes left to fill the truck, while the human stays in control.

Assist also paves the road to autonomy, with each step funding the next:

  1. Shadow mode. The model runs silently and compares its suggestion with what the expert did, so evaluation carries zero risk.

  2. Advice. On-screen or audio suggestions; the phone can be the display.

  3. Single-function automation, such as auto return-to-dig or bucket leveling.

  4. Supervised autonomy.

Today's grade control gets the geometry right. Learned prediction is the next layer: it knows what happens when the bucket meets the dirt. That's why the CAN data matters. Without it, the model sees what happened but never learns what the operator asked for, and the gap between command and response is where soil, load, and skill show up.

For a world model, ground truth is depth, pose, action, and outcome, all on one clock.

From cell phone to metric 3D sensor

Take the Samsung Galaxy S25 Ultra. Its rear camera island holds four cameras. Adjacent lenses sit 18.5 mm apart, 37 mm end to end. That's a short baseline, but the pixel density is extreme. At full resolution, the main camera's focal length is about 10,900 pixels, roughly 5x what a typical machine-vision stereo camera offers.

Article content

Modern phones carry several stereo pairs on the back. Our first prototype uses the main and ultrawide cameras of an inexpensive POCO phone. Samsung results will follow.

Stereo depth error grows with distance squared and shrinks with baseline times focal length in pixels. Assuming 0.2 px matching precision, and taking the coarser camera of each pair as the limit:

Stereo depth error vs. range for different camera pairs

Doubling the baseline doesn't help if one camera is coarse. The ultrawide's lower angular resolution cancels the extra length, which is why the main + 5x pair wins: about 0.5% error at 5 m and 2% at 20 m.

This is mid-range depth, not long-range perception. It doesn't need to be: most of the world that matters for training sits inside that envelope.

Is it good enough for real machines?

Here's how that stacks up against what popular machines need. The requirements below are my estimates from each machine's working geometry and speed, not industry standards. The phone curves show the main + 5x and main + ultrawide pairs at full resolution.

Article content

The pattern is clear. The phone covers the slow, close-in work: digging, loading, grading, and tillage. That's exactly where world-model data is missing. Keep in mind that the main + 5x pair only sees the center of the scene (about 18 degrees wide), and these are full-resolution numbers. 4K video is about 4x worse, and our 720p prototype is about 12x worse, so treat these as best case. Closing that gap is the full-resolution work described below. Fast machines that need to see 50 to 100 m ahead are a different job. That calls for a wide baseline, which is what Hammerhead does on vehicles.

Why hasn’t this been done before? Calibration!

Phones have used their cameras as binocular pairs for years: dedicated 3D phones in 2011, dual-camera portrait mode since 2016, and spatial video today. All of it works at arm's length and only needs depth good enough to blur a background. Metric depth at 20 m needs something harder: calibration that holds in every frame.

  • OIS moves the lens every frame.

  • Autofocus changes the effective focal length.

  • Mismatched cameras: different FOV, resolution, distortion, and color response.

  • Rolling shutter and loose sync between modules.

  • Heat and flex in a 7 mm chassis.

Factory calibration is stale before the shutter closes. Locking OIS and AF fixes the geometry but ruins the camera.

This is the same problem as wide-baseline stereo on a flexing truck or tractor, scaled down 100x. The fix is also the same: solve the geometry from the scene itself in every frame, rectify mismatched cameras to a common virtual camera, and use the extra cameras as a consistency check. That's how we approach it at NODAR, and it carries over to a phone.


Early prototype, 3.6 FPS. Depth from the main and ultrawide cameras of an inexpensive POCO phone, running NODAR's stereo pipeline at 1280x720. Left: main camera. Center: depth, 0.35 to 28 m. Right: ultrawide, cropped to match.

How fast is fast enough? For metric depth labels, about one stereo frame every second or two is plenty: at tractor speed, consecutive frames still overlap. Teaching dynamics needs more, roughly 2 to 10 frames per second, while the CAN logger records actions at 10 to 20 Hz. At 3.6 FPS, the prototype already meets the first bar and sits inside the second.

Article content

Hammerhead Mobile runs on Android first, where GPU compute support is strongest. iOS will follow.

How it scales

How much data? For scale, Wayve trained its GAIA-1 driving world model on 4,700 hours of driving data. These are rough planning numbers for off-road machines:

One machine works about 1,000 to 1,500 hours a year, so ten instrumented machines produce enough for an assist model in a few weeks. Volume isn't the bottleneck. Diversity is: operators, soils, moisture, weather, and job types.

So what does it cost? The usual way to collect off-road data is a dedicated data machine: a sensor kit bolted on, an engineer to integrate it, and an operator to run data passes. Here's a rough comparison, using my estimates:

That's 30 to 70x cheaper on hardware, and several hundred times cheaper per hour of data once you count machine and operator time, because the phone rides along on work that's already paid for. The trade is range and precision: a purpose-built data-acquisition kit sees farther and more accurately. For the slow, close-in work where the phone meets the requirements, it's good enough, and cheap enough to put in every cab.

The phone also does the heavy lifting at the edge. It turns stereo pairs into a compact depth map or terrain heightmap before anything leaves the cab, so it stores and transmits roughly 100 to 200x less data. It processes the 200 MP frames at full resolution on the phone and uploads a ~1 MP foveated depth map or heightmap, and no cloud GPU is spent on stereo matching.

Sim helps with rare and dangerous cases, but it is weakest exactly where earthmoving is hardest: bad weather, soil, hydraulics, and human behavior. The practical loop is real first. Phone depth reconstructs the real site, CAN load and measured terrain change calibrate the simulator, and the calibrated sim generates the variations. Metric depth from a cheap phone becomes an engine for calibrating the simulator, not only a dataset.

What we’re still solving

  • Camera sync. The two cameras in a phone don't share a hardware trigger, and timing error varies by phone maker, up to about half a millisecond. On a vibrating cab, that's enough to blur depth. Next: estimate each model's timing offset from motion and correct for it in software.

  • Full-resolution streams. The best numbers in the table need both cameras at full resolution at the same time, and phone operating systems often cap that. Next: test on the S25 Ultra, with full-resolution stereo stills between video frames as the fallback.

  • Hard scenes. About half the pixels returned depth indoors, mostly because walls and ceilings have no texture. Dirt, gravel, and crops are richly textured, which is good news. Sky, standing water, dust, and night are not. Next: field data from real job sites, with every pixel tagged by confidence.

  • Life in a cab. Windshield glass, glare, vibration, and summer cab heat throttling the phone. Next: a season of in-cab testing.

  • Is it really ground truth? Labels are only as good as their accuracy. Next: check phone depth against surveyed points and a wide-baseline Hammerhead on the same machine.

  • Who owns the data. Operators, owners, dealers, and OEMs all have a stake, and operators are on camera. Consent and data rights need to be designed-in from day one.

Start with one cab

This blog asks how machines see in 3D and what it costs. Here the hardware for metric ground truth is a phone and a CAN logger per machine. What's missing is software that makes the cameras agree on geometry.

For OEMs, the same software runs on cameras already on the machine, and the CAN data is already logged. The phone is the fastest way to start; the fleet is the way to scale.

If you run a fleet working land the world models have never seen, put a phone in a cab and help us solve the problems above. It may be the cheapest data program you'll ever start. We're looking for two or three fleets to pilot this season. Reach out to discuss.