Solving Depth Perception for Robotic Bin Picking

Bin picking depends on depth data that holds up pick after pick, tote after tote. A few millimeters of error and a gripper misses its mark or an object gets flagged as unreachable when it isn’t. Solving for accuracy at close range, across a range of bin sizes and object types, is a baseline and optics problem before it’s a compute problem.
To verify this, we built a test system with a 75mm baseline, paired with Sony IMX490 sensors (5.4MP) and 6mm lenses for roughly 70 degrees of field of view, running on Hammerhead, our patented stereo vision platform. That combination covered a volume of interest at 1.7m, sized for standard tote picking.


The tradeoff with a short baseline is usually accuracy. Sensor resolution is what recovers it. We experimented with a tripod at 1.9m and measured around 5mm precision at a 1.7m scene height. Reconstruction held up cleanly across the objects we tested.
Occlusion is the other piece of the baseline tradeoff, separate from accuracy. Wider separation between the two cameras means more parallax, and past a point that parallax means one camera sees a surface the other doesn’t, leaving a gap in the depth map right where a pick point is needed. A short baseline keeps both cameras converged on nearly the same view of the working volume, so that gap doesn’t open up even in tight or subdivided spaces.
That’s not an argument that short is always right. Open a bin up with no subdivision and no occlusion risk, and a wider baseline uses the same sensors to buy back accuracy, since depth error scales with baseline as well as resolution. The point of the HDK Nano is that baseline doesn’t have to be a fixed assumption made at the sensor’s design stage. It’s a parameter we can tune to the application during setup.
Same sensors, same platform. Baseline gets optimized for the job once, not designed around a one-size-fits-all default.
