4 ms·
> I just think that if the software is unable to make decisions based on visual data alone without up to date high resolution maps it'll never achieve true FSD
by jowday 5y ago
> I just think that if the software is unable to make decisions based on visual data alone without up to date high resolution maps it'll never achieve true FSD in the general case (not geo locked). You'll end up trapped in a local max otherwise because there are just too many conditions in the real world that vary.
My contention is that there’s no way to actually solve for the general case with currently existing technology. The amount of novelty in the real world is too great for any system to account for it without disambiguating via HD maps or remote support.
>You have to solve the vision problem.
This isn’t a vision problem specifically - even if you had LIDAR and high resolution imaging radar and 8 A100s on every Tesla, “true generalized self driving” wouldn’t be achievable without HD maps with our current understanding of Machine Learning.
>My understanding was that Tesla did not rely on the same stuff that Waymo and Cruise require.
Tesla maps individual traffic light elements, stop signs, and lane markings, but will attempt to drive even if the area isn’t mapped.
Disparities in FSD performance in different areas is largely attributable to some areas being better mapped than others - the mapping data has a huge effect on its performance. There are key elements of the driving task (including recognizing and reacting to every single type of sign other than a stop sign) that FSD can’t do and relies entirely on maps for.
- Retric 5y agoNovelty isn’t nearly as big of a problem as you might think. One of Wamo’s famous videos was someone on an electric scooter chasing a duck in the middle of the street. That’s very odd behavior, but the car followed the rather simple option of just not hitting them and going forward when possible. Cars really don’t need to identify what something is just it’s location and movement which is a vastly easier problem. A trash can rolling down the street can be treated just like an oil drum doing the same thing etc.
- jowday 5y ago> Cars really don’t need to identify what something is just it’s location and movement which is a vastly easier problem. A trash can rolling down the street can be treated just like an oil drum doing the same thing etc. You’d think that, until you encounter something like a turn restriction sign with a bizarre conditional restriction that it’s never seen before. At which point the car needs to OCR the text, parse the semantic meaning, and apply to the scene.
- Retric 5y agoOr treat that turn restriction as applying 100% of the time.
- jowday 5y agoAnd now we’re already making concessions about the car’s abilities. There are 10 MPH speed limit signs on Market Street in SF that specify in incredibly small text “when behind trolleys”. Assuming we take your approach, the car will just always go down market at 10 MPH. Imagine if it’s a negative turn restriction - IE, it’s permitting turns except for during certain hours and conditions. Now the car is treating it as always permitted and turning into traffic. An edge case, but something it’s going to encounter in the real world.
- Retric 5y agoAnd now your moving the goalposts. We are talking extreme edge cases in some random small town not common signs in a major city. They can always get updates on what some random sign in some random location means as long as their safe and don’t block traffic that’s all that’s needed. Also, negative restrictions can again default to full restrictions. Permitting a car to say park in a snow lane doesn’t require a car to park in the snow lane.
- jowday 5y agoI don’t think I’m moving the goalposts - we were discussing whether autonomous driving (which I take to mean L4-L5 driving without the need for a human in the loop) is possible without geofences or HD maps. “Edge cases in some random small town” are exactly the sort of thing you need to worry about without a geofence. Not to mention these sorts of edge cases are way more common in large cities than small towns - one of the examples I gave was down a central avenue in San Francisco. >They can always get updates on what some random sign in some random location means as long as their safe and don’t block traffic that’s all that’s needed. What if it truly fails to parse the sign accurately and does something illegal or dangerous? What does sending an update out look like? Does a human take a look at a crop of the sign and review it? Why not just map it in that case?