3 ms·
Yes, the vision only approach isn't something that you do in robotics. There is usually a hierarchy of sensors, mainly for redundancy. Example: Bumper sensory
by osuairt 4y ago
Yes, the vision only approach isn't something that you do in robotics.
There is usually a hierarchy of sensors, mainly for redundancy.
Example: Bumper sensory at the wheel base, sonar / Lidar at the mid, and a camera at the top for advanced sensing.
For the sake of cost cutting Tesla has done away with their radar sensors at the front of the vehicle. It would be a substantial cost overhead, but have very real repercussions when it comes to safety, while also providing a "ground truth" to what at least the front facing cameras are seeing.
I don't think Lidar is a practical sensor for them to adopt, because it is quite bulky and has limited viewing angles, but I would expect them to have adopted some novel, lower cost radar solution.
Apart from the lower cost of the camera, I think Elon's rationale for having a camera only FSD is not valid, has made the problem needlessly complex and unsafe. He believes since we have eyes, and we can drive a car, then it should be sufficient to drive the car, but we only use eyes because these are the sensors we were born with, it is the best we have. In my mind, Elon's approach is like looking at a horse, and saying to yourself, that you want to build a car based on a horse, where instead of wheels, you have four mechanical legs, and those mechanical legs are limited is so many ways, but they should still at least "work", but there is no reason to limit locomotion in that way. The same with the vision system on a FSD, the whole spectrum of light is available, with any number of configurations, providing data at rates and with precision far beyond what a camera system can do.
- agoose77 4y agoIIRC there was a presentation from Karpathy talking about the challenges with sensor fusion, particular in resolving divergence between e.g. the vision and the radar stack: https://www.youtube.com/watch?v=NSDTZQdo6H8&t=1949s https://www.youtube.com/watch?v=NSDTZQdo6H8&t=1949s My background is in physics, but I find myself having a growing appreciate for the vision-only stack. It's really challenging building a formal understanding of the world that is robust to outliers that are so numerous as navigating in an urban environment. With vision, you have multiple kinds of information that are highly correlated (colour, spatial distribution, depth, etc) that are self-consistent. Whereas, fusing radar with vision, where object responses to radar are highly geometry & material dependent, is a much harder task. I'm really not an expert, so this reads more as an opinion than an experienced view, but I can see the merits in doubling down on vision.