3D Vision Brief: Stereo depth got cheaper, lighter, and 4x more precise. What changed?

We recently ran a side-by-side test for a humanoid robotics application. We compared a popular off-the-shelf stereo camera (1 MP, 95 mm baseline, active IR) with an inexpensive 5 MP binocular camera running our Hammerhead software. I want to share the results, mostly because the gap is simple physics, and it applies to anyone choosing a depth sensor for a robot.
Here's the first thing we noticed. In the same scene at the same moment, the floor and the lid of a Pelican case came out flat in the 5 MP depth map, while the 1 MP camera showed ripples across both. The 5 MP point cloud also had enough detail to read the word "nodar" knitted into a beanie and a sock.

Left: reference stereo camera. Right: 5MP binocular camera with NODAR Hammerhead software
The two setups:

How we measured range precision
We placed a ChArUco board, a known flat target, at 5 m and measured how far each depth map deviated from a plane. The reference camera surface deviation came in at 33.1 mm RMS, and the 5 MP setup at 7.1 mm.

Charuco deviation maps
Next, we moved a speckle target from 1 to 6 m, captured 20 frames at each distance, and computed the RMSE about the fitted plane. The 5 MP setup's error was about 4x lower across the whole range: 8.3 mm vs. 38 mm at 6 m. It was also steadier from frame to frame.

Depth precision vs. range
Why the gap is about 4x
Stereo depth error scales roughly as Z² / (f × B): distance squared, divided by focal length times baseline. The 5 MP setup has a 3.2x longer focal length in pixels and a 1.26x longer baseline, and together those give 4.0x. The measurements landed about where the math predicted. There's no magic in that part.
The hard part is keeping a narrow-angle, high-resolution stereo pair calibrated. The finer your angular resolution, the smaller the misalignment that corrupts depth. A sub-millimeter shift in the mount is enough, and the error grows with the square of distance. Robots vibrate, get bumped, and warm up. What you calibrated on the bench slowly stops being true. Calibration gets even harder as robots move toward ultra-lightweight, somewhat flexible materials. Furthermore, robots fall, a lot, and that wreaks havoc on stereo camera calibration.
Hammerhead recalibrates continuously from the live images. In one sequence, the cameras started out misaligned, and the depth map of the Pelican case was unusable. By the second frame, the alignment corrected, and the case's shape returned.

Left: Frame 1. Right: Frame 3
Without that correction, the errors show up downstream as SLAM drift, phantom obstacles, and missed grasps.
If you've worked with stereo before, you might assume 5 MP depth needs a large GPU and a matching power budget. That used to be true. Dense matching at this resolution meant heavy, slow algorithms that could eat most of an embedded board.
That has changed. Everything in this test ran on an NVIDIA Orin Nano, the smallest module in the Orin family. We've also recently ported Hammerhead to TI's TDA4 and NXP's Arm-based processors. So the software fits on the kind of embedded silicon that's already inside many robots and vehicles, instead of forcing a compute upgrade just to get depth.
This test was for a humanoid, where the cameras sit on the head. On an inverted pendulum, mass at the top costs more than its weight: our rough estimate is that each gram on the head acts like about 4 g of dynamic load on the joints below. Going from 116 g to 25 g isn't the headline, but the team cared about it.
In fairness, the 1 MP reference stereo camera does some things better:
It ran at 30 fps in this test; the 5 MP binocular camera firmware was limited to 10 fps.
Its IR projector lets it work in the dark and on blank, textureless surfaces. Passive stereo needs some light and some texture.
It computes depth on board, so you don't need a host processor. The 5 MP setup needs one, even if it's small. Our $138 figure is camera hardware only, and Hammerhead is licensed separately.
For close-range work in mixed lighting, it remains a solid default for large objects that don’t need 5MP resolution.
Takeaway
If you need depth beyond a couple of meters, or fine surface detail up close, focal length and baseline matter more than almost anything else, as long as you can keep the pair calibrated. And the processing no longer requires a big GPU. If you're working through that trade-off on your own robot, I'm happy to compare notes.
