4 ms·
> For example, one of the distance estimation algorithms used in the Cornell paper, developed by two researchers at Taiwan's National Chiao Tung University, rel
by ThatGeoGuy 7y ago
> For example, one of the distance estimation algorithms used in the Cornell paper, developed by two researchers at Taiwan's National Chiao Tung University, relied on a pair of cameras and the parallax effect. It compared two images taken from different angles and observed how objects' positions differ between the image—the larger the shift, the closer an object is.
The shift or disparity between sensors doesn't really matter. We've known that wider convergence angles begets better object point estimation since the 70s. Yet, even the KITTI dataset doesn't attempt to take advantage of this, and uses two rather average cameras with a (relatively) short baseline of 0.06m (see: http://www.cvlibs.net/datasets/kitti/setup.php http://www.cvlibs.net/datasets/kitti/setup.php). That's 6cm!!! You have the entire width of the car to separate these cameras by.
> This technique only works if the software correctly matches a pixel in one image with the corresponding pixel in the other image. If the software gets this wrong, then distance estimates can be wildly off.
Again, yeah. But the problem is twofold: you need to detect / match similar points between two images, but the fundamental setup of your system can limit your precision and accuracy. Use a wider-angle lens with better convergent geometry. Every publication based on the KITTI dataset doesn't even address some of the most basic criticisms from photogrammetry.
Which leads to probably why LiDAR gives such a distinct advantage in most of these data sets. You solve two problems:
1) You solve the correspondence problem trivially because LiDAR doesn't need to match points between cameras, and there's no baseline / convergence criteria that the final point precision depends on.
2) Robust geometric data is well-modelled, well understood, and provides an easier criteria for machine learning systems (particularly ones running over KITTI, as in the article) to converge on than just using stereo-imagery with a baseline of 6cm. You get the scale of the system for free and your calibration troubles are whisked away as LiDAR systems tend to be better-calibrated and more stable than most lens systems or configurations you'll find in the cheap off-the-shelf cameras that many autonomous driving startups are using.
I guess I come off a little negative by looking at this, but my first reaction to Musk saying that nobody should or will want to ever use LiDAR for this is that he doesn't know a damn thing about what he's talking about.
- tlb 7y agoA 6 cm baseline is enough for humans to make adequate distance estimates. Besides the correspondence problem, a longer baseline makes it hard to keep the cameras aligned as the vehicle bounces and flexes. You can't mount them separately to the car -- a chassis can easily twist by a degree or two. So you need a stiff mounting bar between them, which you can either put outside the car like a roof mount (ugly, and it gets buffeted by wind) or inside (also ugly).
- flor1s 7y agoWhy even limit yourself to two cameras? If I recall correctly multi-view geometry benefits from having as many cameras as possible. In the future we will all have walls covered with a checkerboard pattern in our garage to calibrate the cameras on our self driving cars. :)
- dreamcompiler 7y agoGreat points. It would make perfect sense to have two baselines: One of a few cm for nearby objects and one car-width for good depth resolution of distant objects (which humans can't do, but humans have much better world models than computers, so better depth perception on the part of computers might close that gap a bit.) I also think lidar or radar will always be necessary. The Tesla fatality last week happened because a big white truck pulled out in front of the car. With a big blank surface, stereo pixel correlation is impossible, but it's trivial for lidar or radar to read such surfaces.