4 ms·
Stereo vision gives you two slightly different perspectives on a scene. A single point in the 3D scene will be seen in a slightly different position in each 2D
by learn_more 9y ago
Stereo vision gives you two slightly different perspectives on a scene. A single point in the 3D scene will be seen in a slightly different position in each 2D image. The offset between the two positions give you the depth information. The more offset, the closer the point is to the observer.
The difficult part however is pairing the points between the images. You and I do that easily, because when we see, for example, the top left corner of our refrigerator, we can easily identify that same point in both images, probably because we recognize the objects we are looking at. A computer has more trouble pairing the points.
Lidar is a way a cheating, by temporarily shining a dot on the locations so that the points can be found and coordinated easily from both images.
- mturmon 9y agoJust to take this a step farther: the distance-measurement error characteristics of lidar and stereo are very different. To get accurate range through stereo, you need two calibrated cameras, rigidly mounted (with respect to each other), and with a high resolution (pixels x pixels in each image of the stereo pair). All this implies very expensive sensors, large camera rigs, and lots of image processing. To get acceptable stopping distances for a vehicle going 60mph using stereo only to detect obstacles requires things like 1.5-meter camera bars, 2048 pixels per image in the stereo pair, and onboard computing to compute stereo correspondence at 30 or 60 fps. It's hardware-intensive! TBH, I forget how the stereo range error scales with range - I think it is linear with range, but may be super-linear. This can be a problem for mapping. I think lidar is superior in this regard, in other words, its error scaling is sub-linear with range. If you're used to using only stereo, the concept of having a lidar for that point-range measurement looks pretty magical. Of course, stereo vision offers some advantages relative to lidar - it's a passive measurement, for example.
- joshvm 9y agoStereo error is quadratic with range. Bad times. You don't need high resolution, just a wide baseline for accuracy at distance. Resolution will only get you so far because you can't get (and you wouldn't want) sensors with pixels below a micron in size. You usually can't change the focal length much because you have constraints on field of view, therefore the only viable option is more baseline. e_z = e_d * Z^2/(b*f) e_d is the disparity error (i.e. matching error) in metric units; i.e. a few microns, Z is distance, b is baseline, f is focal length It's relatively easy to estimate the accuracy you can achieve on a car because the maximum baseline is fixed to < 2m typically. You can plug in reasonable fields of view, sensors etc. Typically you can assume 0.25 px matching accuracy, assuming the algorithm does sub-pixel interpolation. Example: e_d = 2.2µm, Z=30m, b=1.5m, f=3mm (for a 2000px wide sensor that's 70 degrees FOV): e_z = ~40 cm. That's not too bad - enough to identify something big. At 100 m you'd have a disparity of 20 pixels. If you had a search radius of 256 px you'd have a close-range of 8m. Nowadays the compute problem isn't too bad - you just throw a GPU at it (or an FPGA, but you have memory limitations there). LIDAR has a more or less constant error with range, provided you've got a high enough signal to noise. However this does mean that at short distances, LIDAR is quite poor relatively.
- mturmon 9y agoThanks for this. I wasn't sure that it was as bad as quadratic, but there it is. Your calculations show why lidar is such a persuasive alternative.