8 ms·
A big part of the presentation so far was an engineer describing the huge difficulty of stitching together the multiple cameras into one vector space that can b
by 2bitencryption 5y ago
A big part of the presentation so far was an engineer describing the huge difficulty of stitching together the multiple cameras into one vector space that can be the input to the network, instead of treating each camera individually.
Seems like the biggest problem was each pixel from a camera does not tell you how far away it is, so even if you know the camera is X feet off the grounding pointing at Y degrees, you don't know if there's a wall in front of the camera or not. So you can easily reconstruct a 2D space from all the cameras just by knowing their positions, but you can't simply translate that back to 3D.
I could barely follow the solution to this problem, but it seems to approximate the distance somewhat okay... but it required yet another universe of features in the neural net pipeline just to achieve it.
The funniest part is... isn't this exactly what LIDAR is amazing at?? Or what about having each "camera" actually be two cameras that can achieve depth via parallax?
- ModernMech 5y agoYes. This is exactly what LIDAR is for. This is the reason every single team in the 2007 DARPA Urban challenge that finished the race was equipped with a Velodyne laser. We knew it was a critical enabling technology 14 years ago, and it baffles me Tesla is eschewing it. Tesla's lack of this technology is the main reason I feel they will never achieve what they claim to without a huge breakthrough in AI.
- tru3_power 5y agoWhat was their reasoning for not using Lidar?
- Siecje 5y agoCost and reliability. But Elon says it is because they don't need it.
- rnaud 5y agoToo expensive + Your eyes already do the distance estimation and they are basically cameras so we can do this just as well or better.
- mdorazio 5y ago> Your eyes already do the distance estimation and they are basically cameras so we can do this just as well or better. This is not really accurate. Our brains do the distance estimation, and they use all kinds of tricks and contextual clues to do it, not just parallax (~2.5 inch parallax for objects more than ~100ft away isn't all that helpful). And this is the whole problem with vision-only autonomous driving - ML capabilities are nowhere near the capabilities of a human brain.
- awal2 5y agoMonocular depth estimation has gotten really good recently though[0]. Not saying this one paper/method is 100% sufficient, but we're closing the gap in this one capability (depth estimation from pure vision) quite rapidly. [0] https://roxanneluo.github.io/Consistent-Video-Depth-Estimation/ https://roxanneluo.github.io/Consistent-Video-Depth-Estimati...
- ajross 5y ago> Our brains do the distance estimation, and they use all kinds of tricks and contextual clues to do it So, first, this presentation we're talking about is literally about problems like that. Second: your brain isn't nearly as good as you think it is, it's just constructing a coherent story to fool you into thinking it is. Try this on the highway sometime as a passenger: close your eyes and recite the distances to the vehicle in front of you and the one to either side. I bet you anything a Tesla is going to do that better. Third: they pretty much cracked this already. They stopped shipping radar on US Model 3's and Y's in the spring, have shipped hundreds of thousands of them now, and there's not a hint of signal that something is off with distance measurements with the cars. My car doesn't get this perfectly (you can actually watch the animations on screen bounce around a bit as the estimates change) but I think it objectively does better than I do. Distance/Lidar framing is old news, basically. Vision works fine. The worst bugs remaining with all the FSD Beta footage on Youtube are almost entirely pathing and planning issues. The car sees its environment just fine.
- ctdonath 5y agoThat’s exactly what Tesla described tonight. 8 video cameras on a moving platform provide enough eyes, parallax, context, etc to build accurate 4D (!) models in real time.
- doctoboggan 5y agoI think originally it was because of the cost of the equipment. Elon and team saw how expensive lidar was and thought they needed to be able to solve the problem without it. However lidar has dropped drastically in price, so at this point I think they aren't using it just for ego reasons (i.e., we claimed it was possible to get self driving without lidar in the past, so we have to continue down that path)
- dharmab 5y agoAlso they marketed autonomous driving features as future software updates (including a discount/price hike for early/later adopters). Early adopters would be unhappy if future features required additional hardware.
- mdoms 5y agoNot just marketed, but sold, and for a very pretty penny.
- Robotbeat 5y agoThey upgraded hardware in the past to support FSD customers, so I’m not sure they’re completely against it. If it was, say, just a $500 decision to make it completely viable, I don’t think Tesla would hesitate. But the $500 lidars can’t do what needs to be done.
- kitsunesoba 5y agoWasn't another factor the physical size of the lidar units? Part of Tesla's schtick is making normal or even attractive looking cars (as opposed to the "alien bug" aesthetic EVs were synonymous with at the time) and that's a lot harder to do with a big lidar unit on top.
- mandeepj 5y ago> so at this point I think they aren't using it just for ego reasons If their engineering is driven by ego (most of us think - it is) then - are they really engineers? We all know it’s coming from Top - Elon, in this case.
- Aliabid94 5y agoToo expensive. They also removed radar because of the global chip shortage so they're all in on vision.
- tablespoon 5y ago> Too expensive. They also removed radar because of the global chip shortage so they're all in on vision. That doesn't make much sense to me. Tesla also needs chips for computer vision.
- tim333 5y agoYeah, from their talk it seemed more that it was difficult to integrate the information from the radar and visual sensors.
- pcbro141 5y agoMaybe the government's Autopilot probe will force them to change their mind on Lidar.
- twiceaday 5y agoI think Elon basically says that long term they will only need vision so they will spend all their time focusing on vision from the start. Maybe lidar will succeed first, but it will be a worse success than vision: redundant / expensive / intrusive.
- babelfish 5y agoElon’s ego
- deleted 5y ago[deleted]
- tw04 5y agoI for the life of me cannot find the quote, but at some point I believe he said that if humans can judge distance/drive with nothing but vision, cars/computers will be able to as well. It was around the time he made the comment that anyone using lidar is doomed, which is much easier to dig up as it was a headline everywhere: https://arstechnica.com/cars/2019/08/elon-musk-says-driverless-cars-dont-need-lidar-experts-arent-so-sure/ https://arstechnica.com/cars/2019/08/elon-musk-says-driverle... Hoping someone else has the link to the discussion of humans using nothing but vision, my google-fu is lacking this evening.
- deleted 5y ago[deleted]
- reubenswartz 5y agoIt doesn't work well in rain, snow, fog, dust, etc. It's great in good conditions, but you either have a car that can drive only in good conditions, or you need a car that can drive safely without LIDAR. (Or someone needs to invent a better LIDAR.)
- sgustard 5y agoSee, for example, Andrej's talk starting at the 6-minute mark. Lidar requires a pre-rendered detailed map of the lanes, traffic lights and obstacles. Vision can operate in any novel environment the car is not preprogrammed for and is thus more scalable. https://www.youtube.com/watch?v=a510m7s_SVI https://www.youtube.com/watch?v=a510m7s_SVI
- abc_lisper 5y agoIdk.. animals seem to do fine without a lidar
- lttlrck 5y agoIf Tesla switches to stereoscopic vision with fully articulated cameras that can move independently of the vehicle (not fully independent, but limited 3 degrees of motion), and then they manage to integrate the output into something meaningful then maybe it would be comparable.
- cycrutchfield 5y agoThey already have multiple cameras; it's not clear why articulation of cameras would be required as they do not have foveas that require such articulation; they already do integrate the output of each camera into a view of the vehicle's surroundings.
- shkkmo 5y agoAnimals tend to use a combination of four tools to detect distance. Stereo vision, parallax, focus detection and scene context. Independent movement is not required for any of these (though it can help generate parallax.) Humans (and other animals) can learn to do pretty well with one eye and no head movement given the appropriate training regime. It makes the software problem harder but it certainly is possible.
- pyinstallwoes 5y agoMaybe it's more like spider eyes that see a blended vectorspace of reality with less movement?
- ModernMech 5y agoAnimals don't drive cars... edit: Since this got such a negative reaction I'll amend it with more of an argument. If an animal can do X with eyes, this in no way implies we can do X with binocular cameras, as it completely discounts 1) our eyes are preprocessors for our brain in a way that cameras are not. That both sensors capture light doesn't mean they are equivalent. 2) robot brains are not there, in any way, shape, or form.
- ec109685 5y agoHe talks about the approach here. They actually used radar and LIDAR to train the depth sensing neural network: https://t.co/osmEEgkgtL?amp=1 https://t.co/osmEEgkgtL?amp=1
- goldbattle 5y agoWhile taking two sequential images from the same camera in can provide this depth information also. But simple use of stereo cameras can solve a magnitude of problems (standing still and low parallax motions). Traditional stereo and even machine-learning based methods have have great success and accuracy for many years and could easily be an alternative to LIDAR also. I really don't know why this isn't leveraged more (maybe it is and I am unaware?).
- FredFS456 5y agoTesla has multiple cameras looking out the front, with some offset. I believe that in past talks they did talk about using that parallax to achieve better depth.
- jsight 5y agoThat was basically what I was thinking as well. There's a reason that Mobileye's camera only solutions tend to have a lot more cameras than Tesla does. Stereoscopic vision would be a big help, and tesla only has it in some directions and with relatively low resolution.