8 ms·
As someone who works in the industry (disclaimer these are my own views and don't reflect those of my employer), something about the framing of this article rub
by ikhatri 3y ago
As someone who works in the industry (disclaimer these are my own views and don't reflect those of my employer), something about the framing of this article rubs me the wrong way despite the fact that it's mostly on point. Yes it is true that different companies are choosing different sensing solutions based on cost and the ODD in which they must operate. But I think this last sentence just left a sour taste in my mouth "But the verdict is still out as to which is safer".
It is not an open question and I hate it when writers frame it this way. Camera only (specifically, monocular camera only) systems literally cannot be safer than ones with sensor fusion right now. This may change in the future at some point, but it's not a question right now it is a fact.
Setting aside comparisons to humans for a second (will get back to this), monocular cameras can only provide relative depth. You can guess the absolute depth with your neural net but the estimates are pretty garbage. Unfortunately, robots can't work/plan with this input. The way any typical robotics stack works is that it relies on an absolute/measured understanding of the world in order to make its plans.
That isn't to say that one day with sufficiently powerful ML and better representations we would be totally unable to use mono (relative) depth. People argue that humans don't really use our stereoscopic depth past ~10m or so and that's a fair point. But we also don't plan the way robots do. We don't require accurate measurements of distance and size. When you're squeezing your car into a parking spot you don't measure your car and then measure the spot to know if it'll fit. You just know. You just do it. And it's a guesstimate (so sometimes humans make mistakes and we hit stuff). Robots don't work this way (for now), so their sensors cannot work this way either (for now).
- shadowgovt 3y ago"But humans can do it with one eye closed and" ... and I want to grab the guy who says that by the collar and scream in their face "The whole point is to build something that can do better than a human."
- ActorNightly 3y agoSelf driving isn't a sensor problem, its a software problem. From how humans drive, its pretty clear that there exists some latent space representation of immediate surroundings inside our brains that doesn't require a lot of data. If you had a driving sim wheel and 4 monitors for each direction + 3 smaller ones for rear view mirror, connected to a real world car with sufficiently high definition cameras, you could probably drive the car remotely as well as you could in real life, all because the images would map to the same latent space. But the advantage that humans have is that we have an innate understanding of basic physics from experience in interacting with the world, which we can deduce from something simple as a 2d representation, and that is very much a big part of that latent space. You wouldn't be able to drive a car if you didn't have some "understanding" of things like velocity, acceleration, object collision, e.t.c So my bet is that just like with LLMs, there will be research published at some point that given certain frames in a video, it will be able to extrapolate the physical interactions that will occur, including things like collision, relative distances, and so on. Once that is in place, self driving systems will get MASSIVELY better.
- cj 3y ago> Self driving isn't a sensor problem, its a software problem. Taking things to the extreme, perhaps it’s actually a networking problem. Cars should have the ability to send signals and hear signals from the cars around it. Imagine if Car A could improve its own understanding of the environment using inputs/sensor data from nearby Car B.
- hedora 3y agoFor this to work, either (1) the network has to be reliable, and all cars have to be trustworthy (both from a security and fault tolerance perspective), or (2) the cars have to be safe even when disconnected from the network, such as during an evacuation. We already know for sure that we can’t solve (1), which means we have to solve (2). Therefore, car-to-car communication is, at best, a value add, not the enabling technology.
- iknowstuff 3y agoNot much would change. The idiotic idea of removing traffic lights in favor of self driving cars zipping past each other forgets about those pesky pedestrians we should be designing cities for.
- cratermoon 3y agopedestrians, cyclists, skateboarders, and all the other road users that the US car-centric society has determined are "hazards" to driving.
- cj 3y agoWhen I wrote the comment, I was envisioning the current world, but with some bluetooth type protocol that cars could use to send beacons to help other cars near it. The most basic example of how this could be helpful is if the car ahead of you turns a sharp corner and crashes into a truck stopped in the road. Without car-to-car networking, you won't brake until the crash is in your line of sight. Have you ever seen those youtube videos of massive car pile ups on highways caused by a crash, and then a cascade of additional crashes afterwards? E.g. icy conditions or dense fog. What if the original crash could communicate to cars behind it, wouldn't that be helpful if the crash isn't yet in the driver's (or car's) line of sight? I agree "not much would change" overnight. It's just another input for the car's software to have at its disposal. With the current hardware on the roads, I don't think it's technically possible for autos to achieve legitimate self-driving (if that's even the goal anymore?) - there are way too many edge cases that are way too difficult to solve for with just software.
- rbanffy 3y agoMinor nitpick > monocular cameras can only provide relative depth While the environment awareness is nowhere near as good as two or more cameras would be, if you consider the output over time, you get valuable information about the change rate of the environment, i.e. how fast that big thing is getting bigger, which may indicate one should actuate the brakes. Of course, I'm with the crowd that answers the question with a "how many can we have?" question. The more, the merrier. And the more types, the better - give me polarized light and lidar, sonar, radar, thermal, and whatever else that can be plugged in the car's brain to make it better aware of what happens (and correctly guess what's going to happen) outside it.
- hedora 3y agoMonocular cameras are a strange strawman. Is anyone seriously considering them? Binocular cameras provide absolute depth information, and are an order of magnitude cheaper sensors than the other options. Since this technology is clearly computationally limited, you should subtract the budget for the sensors from the budget for the computation. According to the article, the non-camera sensors are in the $1000’s per car range, so the question becomes whether a camera system with an extra $2000 of custom asic / gpu / tpu compute is safer than a computationally-lighter system with a higher bandwidth sensor feed. I’m guessing camera systems will be the safest economically-viable option, at least until the compute price drops to under a few hundred dollars. So, assuming multi-camera setups really are first to market, the question then is whether the exotic sensors will ever be able to justify their cost (vs the safety win from adding more cameras and making the computer smarter).
- trihexaflexagon 3y agoAt the risk of stating the obvious, stereovision in practice has a few interesting challenges. Yes, the main formula is deceptively simple: d = b*f / D (d - depth, D - disparity, b - baseline, f - focal length), but in practice, all 3 terms on the right require some thinking. The most difficult is D - disparity, it usually comes from some sort of feature matching algorithm, whether traditional or ML-based. Such algorithms usually require some texture surfaces to work properly, so if the surface does not have "enough" texture (example would be a gray truck in front of the cameras), then the feature matching will work poorly. In CV research there are other simplifying assumptions being made so that epipolar constraints make the task simpler. Examples of these assumptions are coplanar image planes, epipolar lines being parallel to a line connecting focal points and so on. In practice, these assumptions are usually wrong, so you need, for example, to rectify the images which is an interesting task by itself. Additionally, baseline b can drift due to changes in temperature and mechanical vibrations. So is the focal length f, so automatic camera calibration is required (not trivial). Don't forget some interesting scenarios like dust particles or mud on one of the cameras (or windshield if cameras are located behind the windshield) or rain beading and distorting the image thus breaking the feature matcher and resulting disparity estimates. Next, to "see" further, a stereo rig needs to have a decent baseline. For example, in a classic KITTI dataset, the baseline is approximately 0.54m which is much larger than, for example, human eyes (0.065m). Such baseline, 54cm, together with focal length, which, if I remember correctly, is about 720px in case of KITTI vehicle cameras, would give about 388m in the ideal case of being able to detect 1 pixel disparity. But detecting 1px of D is very difficult in practice - don't forget you will be running your algo on a car with limited compute resources. Say, you can have around 5px of D, that means max depth of around 77m - comparable to older Velodyne LiDARs. Some of the issues I mentioned are not specific to stereovision (e.g. you need to calibrate monocular cameras as well and so on), just wanted to point out that stereovision does not magically enable depth perception. The solution would likely be a combination of monocular and stereo cameras, combined with SfM (Structure from Motion) and depth-from-stereo algorithms.
- wilg 3y agoCan you elaborate on your reasoning? I’m shaky on some of logic here. “Monocular” cameras > no “absolute” depth > less safe The last leap is not well justified. Also, cars with vision based driving have multiple cameras. Whats the difference between a “binocular” camera and two “monocular” cameras? How does a “binocular” camera get better depth information? Is using multiple cameras to drive sensor fusion? Why is absolute depth a strict safety win? How do you know how the sensor details translate to the final safety of the full system? If this is just a handwavey upper bound on safety, how do you know that such a system can’t be safe enough for its design goals? If humans with only one eye are able to drive, why wouldn’t mono surround vision be at least as good as that?
- green_man_lives 3y agoI am just a hobbyist but I can answer some of these. > Whats the difference between a “binocular” camera and two “monocular” cameras? For the camera itself, nothing. They are probably referring to the implementation. You can have two cameras side by side but unless you are using homography to estimate depth from the two images, then your setup is monocular. > How does a “binocular” camera get better depth information? A pixel in two images (with known separation) will have a geometric relationship that can be used to extract depth information. This is a lot faster than alternative methods with a single camera and multiple images. > Is using multiple cameras to drive sensor fusion? This is really just a question of semantics. > Why is absolute depth a strict safety win? Why is it better to have two eyes than one? You can be more certain about what you are seeing. > If this is just a handwavey upper bound on safety, how do you know that such a system can’t be safe enough for its design goals? If you had a system with infinite compute you could probably do enough math to calculate absolute depth with 100% certainty. I believe you can already extract absolute depth with something called bundle adjustment-- but it requires multiple images since you are relying on parallax effects. It is also computationally expensive. > If humans with only one eye are able to drive, why wouldn’t mono surround vision be at least as good as that? Computers are not humans.
- michaelt 3y ago> Why is absolute depth a strict safety win? How do you know how the sensor details translate to the final safety of the full system? If you can get reliable depth information, the algorithm needed to avoid hitting stationary and slow-moving objects is extremely simple. Is the stationary object in our path, of nontrivial size, and about to enter our minimum stopping distance? If yes, do we have a swerve planned that will let us safely avoid it? If no, emergency stop. Because this logic is simple and well defined you can audit the implementation to the high standards applied to things like aircraft autopilot systems. And it'll work even if the stationary object is something that didn't appear in your training data - you know the algorithm will work the same even if that concrete barrier is painted with some cheery flowers, or if that fire truck is airport yellow instead of the normal red. Of course, this relies on the assumption you can get reliable depth information. If your depth sensor gets confused by a cloud of dust while driving in the desert, or gets blinded by the light of the setting sun, or is unable to detect a barbed wire fence, things are no longer quite so simple.... > If this is just a handwavey upper bound on safety, how do you know that such a system can’t be safe enough for its design goals? Personally I would say that in freeway driving, a self-driving car should be able to avoid 100% of collisions with clearly visible stationary objects in dry, well lit conditions when all system components are in normal working condition.
- andix 3y agoAs far as I know humans can also safely drive only with one eye. It’s perfectly legal in most countries. But I agree that current software (Tesla?) is not able to do that in the same way. So it may need more sensors until the software gets better. In theory cameras should also be able to see more than humans. They can have a wider angle, higher contrast, higher resolution and better low-light vision than the human eye.
- cratermoon 3y agoA human with one eye can use slight head movements and eye to gain a sense of depth. Perhaps the mono cameras need some kind of mount that allows them to not only look around but also move in 3 dimensions. That seems more complex than just having binocular cameras, though.
- ikhatri 3y agoYup! This kind of reconstruction is known as multi-view reconstruction. Though the cameras don't need to have a movable mount, they're already on a car which moves! The car moves and gives them a new "perspective" at every frame. That's how some monocular systems already work. Here's an example of one such system: https://github.com/nianticlabs/manydepth https://github.com/nianticlabs/manydepth That said, I think what you're referring to is more extreme perspectives that shift in ways the car cannot drive and you are correct that this would aid in reconstruction. This is how NERF models do their 3D reconstruction (https://nerfies.github.io/ https://nerfies.github.io/).
- m463 3y agolike pigeons. Can't cameras do this by just comparing frame 1 to frame 2?
- cratermoon 3y agoCompare how?
- BulgarianIdiot 3y ago> Setting aside comparisons to humans for a second (will get back to this), monocular cameras can only provide relative depth. You can guess the absolute depth with your neural net but the estimates are pretty garbage. Stereoscopic vision in humans only works for nearby objects. The divergence for far away objects is not sufficient for this. You may think you can tell something is 50 or 55 meters away through stereoscopic vision, but you can't. That's your brain estimating based on effectively a single image. That said reality is not a single image, it's a moving image, a video. Monocular video can still be used to estimate object distance in motion. Eventually AI will be good enough to work better than humans with just a camera. The problem is we're not there yet, and what Tesla is doing is irresponsible. They should've added LIDAR and used that to train their camera-only models, until they're ready to take over.
- cypress66 3y ago> You can guess the absolute depth with your neural net but the estimates are pretty garbage. I'm not sure what kind of systems you're referring to with "monocular cameras", but if you look at the visualization in a Tesla with FSD Beta, it's actually really good at detecting the position of everything. And that's with pretty bad cameras and not a lot of compute. Only rarely you'll see Tesla's FSD mess up because of perception, the vast majority of times they mess up is just the software being dumb with planning.
- green_man_lives 3y agoI am implementing Monocular vSLAM as a side project right now. I am working with some optimization libraries like GTSAM but having some issues. Do you know any good resources for troubleshooting this kind of stuff? It's pretty easy to see, even as someone with very little experience, the benefits of stereo vision over monocular. In addition to the depth stuff it's a lot easier/faster to create your point clouds from disparity maps.
- pja 3y ago> People argue that humans don't really use our stereoscopic depth past ~10m or so and that's a fair point. The second paper I reference in this comment https://news.ycombinator.com/edit?id=36232198 https://news.ycombinator.com/edit?id=36232198 claims that humans can maintain stereopsis out to 250m. That’s a huge difference from 10m & if true suggests that human drivers might well use 3D vision when driving.
- ikhatri 3y agoI think I said this already in one of my comments but I'm not a neuroscientist and I don't claim to be. That's why I think it's kind of pointless and silly for me (or any other engineer) to sit here and make arguments about what humans do and don't do in their brains. IMO it's better for us to focus on what the robots can and cannot do right now, and focus on solving those problems :) Thanks for the sources though, those papers are definitely neat and I'll be taking a look when I get a chance.
- GloriousKoji 3y agoAdding in the car speed and direction information to the monocular camera images gets you an absolute/measured understanding of the world.
- williamcotton 3y agoLet’s say you are driving down the street in a suburban neighborhood. You see a kid throw a ball into the street. You see from how his body moved that it is a lightweight ball and that it doesn’t require drastic (or any) measures to avoid. Or you see that it is a very heavy object and requires evasive maneuvers. How exactly does a certain type of sensor help with this? Isn’t the problem entirely based on a software model of the world?