11 ms·
One question I've always had about Tesla's sensor approach: why not use binocular forward facing vision? Seems like it would be a simple and cheap way to get re
by eightails 5y ago
One question I've always had about Tesla's sensor approach: why not use binocular forward facing vision? Seems like it would be a simple and cheap way to get reliable depth maps, which might help performance in the situations which currently challenge the ML. Detecting whether a stationary object (emergency vehicle or child or whatever) is part of the background would be a lot easier with an accurate depth map, or so it seems to me.
Plus using the same cameras would help prevent the issues with sensor fusion of the radar described by Tesla due to the low resolution of the radar.
I know the b-pillar cameras exist, but I don't think their FOV covers the entire forward view, and I don't think they have the same resolution as the main forward cameras (partly due to wide FOV).
I'd love to hear why I'm wrong though.
- simondotau 5y ago> why not use binocular forward facing vision? Because Tesla have demonstrated that it's unnecessary. The depth information they are getting from the forward-facing camera is exceptional. Their vision stack now produces depth information that is dramatically superior to that from a forward-facing radar. https://www.youtube.com/watch?v=g6bOwQdCJrc&t=556s https://www.youtube.com/watch?v=g6bOwQdCJrc&t=556s (It's also worth noting that depth information can be validated when the vehicle is in motion, because a camera in motion has the ability to see the scene from multiple angles, just like a binocular configuration. This is how Tesla trains the neural networks to determine depth from the camera data.)
- jtc331 5y agoWhich raises the question of why it was so easy to demonstrate it failing at CES.
- simondotau 5y agoBecause the software running in release mode is a much, much older legacy stack. (Do we know if the vehicle being tested was equipped with radar or vision only?)
- 9935c101ab17a66 5y agoHow can it be unnecessary if they are having all these issues? The phantom brake events are no joke.
- simondotau 5y agoWhat I was talking about was largely doesn’t apply to the Autopilot legacy stack currently deployed to most Tesla cars. Personally I wish Tesla would spend a couple of months cleaning up their current beta stack and deploying it specifically for AEB. But I don’t know if that’s even feasible without affecting the legacy stack.
- 2muchcoffeeman 5y agoIt makes intuitive sense since you can say, play video games with one eye closed. Yes you lose field of view. Yes you lose some depth perception. But you don’t need to touch your finger tips and all your ability to make predictive choices and scan for things in your one-eyed field of view remains intact. In fact, we already have things with remote human pilots. So increasing the field of view with a single camera should intuitively work as long as the brains of the operation was up to the task.
- simondotau 5y agoAlso there are plenty of humans who are blind in one eye and they can still drive a car without difficulty.
- clouddrover 5y ago> Because Tesla have demonstrated that it's unnecessary. The depth information they are getting from the forward-facing camera is exceptional. Sure! Here's a Tesla using its exceptional cameras to decide to drive into a couple of trucks. For some strange reason the wretched human at the wheel disagreed with the faultless Tesla: https://twitter.com/TaylorOgan/status/1488555256162172928 https://twitter.com/TaylorOgan/status/1488555256162172928
- simondotau 5y agoThat was an issue with the path planner, not depth perception, as demonstrated by the visualisation on screen. The challenge of path planning is underrated, and it's not a challenge that gets materially easier with the addition of LIDAR or HD maps. At best it allows you to replace one set of boneheaded errors with another set of boneheaded errors.
- clouddrover 5y agoNo! It was an issue with the trucks! They shouldn't have been in the way in the first place! Don't they know a Tesla is driving through? They mustn't have been able to see it since they lack exceptional cameras.
- simondotau 5y agoApologies, I thought you were being serious.
- clouddrover 5y agoThat's okay. I didn't think you were being serious so that makes us even.
- drawkbox 5y ago> Their vision stack now produces depth information that is dramatically superior to that from a forward-facing radar. RADAR is more low fidelity though, blocky, slow and doesn't do changes in direction or dimension very well. RADAR isn't as good as humans at depth. Only benefit of RADAR is it works well in weather/night and near range as it is slower to bounce back than lasers. I assume the manholes and bridges that confuse RADAR are due to the low fidelty / blocky feedback. LiDAR is very high fidelity and probably more precise than the pixels. LiDAR is better than humans at depth and at distance. LiDAR isn't as good at weather, neither is computer vision. Great for 30m-200m. Precise depth, dimension, direction and size of object in motion or stationary. See the image at the top of this page and overview on it. [1] > High-end LiDAR sensors can identify the details of a few centimeters at more than 100 meters. For example, Waymo's LiDAR system not only detects pedestrians but it can also tell which direction they’re facing. Thus, the autonomous vehicle can accurately predict where the pedestrian will walk. The high-level of accuracy also allows it to see details such as a cyclist waving to let you pass, two football fields away while driving at full speed with incredible accuracy. [1] https://qtxasset.com/cdn-cgi/image/w=850,h=478,f=auto,fit=crop,g=0.5x0.5/https://qtxasset.com/quartz/qcloud4/media/image/sensorsmag/1524348581/Sensors_Insights_1_2018-04-24_Hero.png/Sensors_Insights_1_2018-04-24_Hero.png?VersionId=AxmNH4lrZxFUSGV7hc_N3QLsU.hQ9Rn5 https://qtxasset.com/cdn-cgi/image/w=850,h=478,f=auto,fit=cr... [2] https://www.fierceelectronics.com/components/lidar-vs-radar https://www.fierceelectronics.com/components/lidar-vs-radar
- deleted 5y ago[deleted]
- leobg 5y agoThey use three forward facing cameras, actually. And they do get a 3D representation. https://mobile.twitter.com/sendmcjak/status/1412607475879137280?s=69420 https://mobile.twitter.com/sendmcjak/status/1412607475879137... https://youtu.be/j0z4FweCy4M?t=3780 https://youtu.be/j0z4FweCy4M?t=3780
- eightails 5y agoSure, but they're not getting that 3d map from binocular vision. The forward camera sensors are within a few mm of each other and different focal lengths. And the tweet thread you linked confirms it's a ML depth map: > Well, the cars actually have a depth perceiving net inside indeed. My speculation was that a binocular system might be less prone to error than the current net.
- leobg 5y agoSure. You're suggesting that Tesla could get depth perception by placing two identical cameras several inches apart from each other, with an overlapping field of view. I'm just wondering if using cameras that are close to each other, but use different focal lengths, doesn't give the same results. It seems to me that this is how modern phones are doing background removal: The lenses are very close to each other, very unlike the human eye. But they have different focal lengths, so depth can be estimated based on the diff between the images caused by the different focal lengths. Also, wouldn't turning a multitude of views into a 3D map require a neural net anyway? Whether the images differ because of different focal lengths or because of different positions seems to be essentially the same training task. In both cases, the model needs to learn "This difference in those two images means this depth". I think with the human eye, we do the same thing. That's why some optical illusions work that confuse your perception of which objects are in front and which are in the back. And those illusions work even though humans actually have an advantage over cheap fixed-focus cameras, in that focusing the lens on the object itself gives an indication of the object's distance. Much like you could use a DSL as a measuring device by focusing on the object and then checking the distance markers on the lens' focus ring. Tesla doesn't have that advantage. They have to compare two "flat" images.