7 ms·
The amount of jitter in the estimates makes me nervous, especially when the model thinks something is present in one frame and not there in the next.
by ipunchghosts 7y ago
The amount of jitter in the estimates makes me nervous, especially when the model thinks something is present in one frame and not there in the next.
- bepvte 7y agoThis is a problem even in production autopilot in Teslas. When stopped at a stop light, you can see cars "dancing" and rotating randomly in a jittery fashion. Today during an auto lane change, the system blinked a truck from two lanes across in and out of my target lane, causing the car to cancel going into an empty lane after quarter ways entering it, twice.
- ChrisClark 7y agoThe dancing cars were fixed several months ago in an update. The computer used to just recognize cars, and then align them according to the lanes it sees. At a stoplight it when it had trouble seeing the lanes clearly the cars would rapidly change orientation. Since they updated the neural net to also recognize the vehicle orientations the dancing has stopped. I have seen a lane change cancel recently, a couple weeks ago though. You can tell the car is a 'nervous' driver. It plays it way too safe, but I guess that's a good thing at this point.
- modeless 7y agoI'm pretty sure that they "fixed" the dancing cars problem by applying a low pass filter to the data before sending it to the visualization, just so people would stop complaining about it. I think there's still a lot of jitter in the underlying data.
- gowld 7y agoThat's a good fix, but they should apply the filter to the data used in the driving logic also.
- Rebelgecko 7y agoIf the filter adds significant latency that could go poorly
- solinent 7y agoTo elaborate, the filter also may not improve the accuracy, just the perceived accuracy. To be correct but one second late is to be completely inaccurate. The system is trying to estimate the current position of the car, but also predict future positions. So a little bit of imprecision is fine since it improves accuracy related to predicting the future positions of the cars. A slight move in one direction may indicate a lane change, so it is always useful to be aware of that so as not to accelerate past a car whose measurement appears to be more inaccurate, since they actually might be moving. If you did the same thing with a human's "sixth sense" perception of the positions of the cars, you'd definitely find that they move a lot compared to their actual positions when the head is turned since our ability to merge our vision and our inertial sense is not very good for the most part. The same issues arises with AR/VR, it's useless to know a more accurate position of the user if it's not the present position, because then that will definitely lead to motion sickness.
- microcolonel 7y agoSeems like it would make more sense to model the inertia. Cars don't randomly accelerate at 100,000m/s/s in some direction they aren't pointed. Though they should have a model for detecting obstacles in the view regardless of inertia, because sometimes something really does appear in front of you in a thirteenth of a second. You could probably model inertia with n prior frames of probability fields.
- eru 7y agoModelling inertia seems like a special case of a low pass filter? A very useful and physically plausible special case, of course.
- modeless 7y ago> Cars don't randomly accelerate at 100,000m/s/s in some direction they aren't pointed What if they are hit by a truck? Maybe not 100,000 m/s^2 but if you assume that cars can't accelerate in directions they aren't pointed, you will be wrong at the worst possible time.
- microcolonel 7y agoThat's why I elaborate, and why I chose that number. The only way something actually accelerates like that is an error, or it's an error.
- modeless 7y agoA threshold that high will be useless as it will miss most errors. A threshold low enough to catch most errors will reject some valid data. A naive approach like that will not work. A better approach would be to include temporal data in the inputs to the neural net so it can learn how to do the prediction and filtering itself using all the context available in the input imagery, instead of processing each frame completely independently and feeding low-dimensional symbolic results into some other system. But you'd need a very large dataset and a very large neural net.
- mannykannot 7y ago
- chillingeffect 7y agoDo you know if, when a vehicle disappears, the system assumes the vehicle continues moving as it was when last spotted?
- bdamm 7y agoI'm pretty sure that the visualization is only showing highly confident classifications (not sure about the SUV/Pickup thing). Under the hood the algorithm is locating all kinds of objects that could be but are not displayed on the screen as some kind of unknown box. Probably the reason Tesla isn't showing this is because the location and size of objects are uncertain and people would freak out if they saw all that traffic (some of it quite close) jumping around.
- taneq 7y agoIf it’s a debug visualisation, why would it not be displaying everything? Of course, it’s in a public release so it’s probably, er, ‘tidied up’ a bit.
- orasis 7y agoDancing isn’t completely gone - especially with large trucks.
- xeromal 7y agoI'm pretty sure the truck thing is because images aren't stitched yet. That's coming in an update.
- plexicle 7y ago"The dancing cars were fixed several months ago in an update." On the latest software and with HW 2.5, this is not true. It's still very much there.
- rootusrootus 7y agoAgreed, with HW3 it's the same, there's still lots of jitter and dancing cars. As of today, with software updated about three days ago.
- YZF 7y agoModel 3 owner here: the dance where cars spun around and landed on top of you is gone but detected cars are still quite jittery. I notice that when I'm stopped and also when I'm driving. This is quite noticeable in the transition between different regions in the car (presumably when the vehicle is handed over between different cameras or sensors). Even something relatively simple as the traffic aware cruise control will sometimes slow down for no apparently reason or simply turn itself off in the rain. Given the combination of the visualizations and the performance of cruise control and autopilot I think Tesla is very far away from fully autonomous driving under all conditions. But they'll probably keep getting better at the semi-autonomous/good conditions/freeway "self-driving"/"augmented driving"...
- jfim 7y agoThat's not unusual for computer vision (or any kind of sensor really), that kind of data is normally filtered and smoothed after that, and merged with previous frames or other sensors. What would be worrying is the model misclassifying an object, not detecting it at all, or having the bounding box consistently off.
- eru 7y ago> That's not unusual for computer vision (or any kind of sensor really), [...] Including human vision. The raw sensory data is pretty messy, and with some ingenious experiments some researchers can get a glimpse of exactly how messy.
- jfim 7y agoOh definitely. The selective attention test [0] really shows how humans can perceive certain things in a way that's different from how computers perceive them. The research on GANs also shows how computers can be fooled by things that wouldn't confuse humans. [0] https://www.youtube.com/watch?v=vJG698U2Mvo https://www.youtube.com/watch?v=vJG698U2Mvo
- TaylorAlexander 7y agoAgreed. It also said it was running at 13fps. Not stoked about a vehicle going 70mph updating at 13fps.
- LeoPanthera 7y agoTesla says their "Hardware 3", which is what you get if you buy it now, can process all cameras at 60fps.
- TaylorAlexander 7y agoI find it odd that they would publish a video showing performance numbers from out of date hardware. I mean I believe you - I watched the presentation in their custom processor and it’s quite impressive. Just weird that they’re showing old performance numbers. Perhaps this video is old.
- grecy 7y agoThere is speculation going around the "big rewrite" Elon mentioned last week is actually porting the code to run natively on the new hardware. Speculation says it's just been running in an emulation layer, but now they're about to unleash the full potential of the hardware. If true, it makes sense the video would also have been captured using this emulation layer, explaining why it's not latest-and-greatest-fast.
- TaylorAlexander 7y agoAh that would make some sense. Certainly I’m expecting extremely good performance from the new computer once everything is running natively.
- mike_d 7y agoIf that is true you should call in to question the integrity of a company that would run life-critical software on a non-RTOS.
- platz 7y agoWhy can't they just whack it with some kind of Bayesian latent space model. A big jump should have to require more evidence than the history of the previous posterior
- iamaelephant 7y ago"Just". Pro tip, these guys work on these problems all day every day. If you think you've solved one of their major problems after 18 seconds of consideration then you're probably missing a large amount of context.
- platz 7y agoI was inviting you to tell me why
- gzer0 7y agoSome care is needed when choosing priors in a hierarchical model [such as Bayesian], particularly on scale variables at higher levels of the hierarchy. The usual priors such as the Jeffreys prior [1] often do not work, because the posterior distribution will not be normalizable and estimates made by minimizing the expected loss will be inadmissible. [1] In Bayesian probability, the Jeffreys prior is a non-informative (objective) prior distribution for a parameter space; it is proportional to the square root of the determinant of the Fisher information matrix. Why is this of relevance? It has the key feature that it is invariant under a change of coordinates for the parameter vector. That is, the relative probability assigned to a volume of a probability space using a Jeffreys prior will be the same regardless of the parameterization used to define the Jeffreys prior. This makes it of special interest for use with scale parameters. Why is this an issue? Accordingly, the Jeffreys prior, and hence the inferences made using it, may be different for two experiments involving the same theta parameter even when the likelihood functions for the two experiments are the same—a violation of the strong likelihood principle.
- Gravityloss 7y ago
- tantalor 7y agoYour human eyes+brain do the same thing.
- Gene_Parmesan 7y agoOur brains are purpose-built to really see only the very core center of our vision; the brain then creates an approximate model of the surrounding space, but a lot of what is in that model is influenced by what the brain "expects" to see. So I think yes, we can also suffer from some similar inaccuracies when the objects in question are in our periphery. However the main difference is that, when we are consciously looking directly at something, we can almost always tell with 100% certainty what we're looking at, up to a considerable distance. I can see a car pulled over to the side of the highway a solid half mile ahead sometimes, and have plenty of time to respond. Computer vision doesn't have this additional strength. As always though, the strength that computer vision has over us is it never gets tired or distracted, and it never operates in "default mode" where sensory inputs don't get full (or even much at all) conscious attention.
- erikpukinskis 7y agoNope. You can “see” all kinds of things that aren’t there, because your brain has yet to notice anything forcefully telling you otherwise. This happens all the time and you’d have no way to notice it. Even when some new information forcefully comes into play, your brain is often able to adjust your memory so you believe you knew it along, so long as the initial percept is fresh enough and had enough uncertainty. All of this feels to you like a perfect unbroken stream of direct seeing but it is an illusion. You don’t see anything directly, you get fuzzy spurts of probability and turn it into your world in your mind. A world that’s likely to be unrecognizable to the next person.
- tantalor 7y ago> we can almost always tell with 100% certainty That's just your brain again. You might mistake a bike for a lamp post, and switch between beliefs several times, before you figure it out, then convince yourself you knew it the whole time.
- 7y ago
- pbreit 7y agoKind of like a human?
- Shivetya 7y ago(TM3 owner) While in motion what is presented to the driver there is little to no jitter. Where you get it mostly is when stopped and the car seems to adjusting between cameras to determine where an object adjacent truly is. sometimes there is no jitter and other times its a bit odd. in motion the car drives just fine with the caveat they have not enabled signal recognition. I use TACC and at times full AP on my daily commute which includes road speeds from 35 to 55. I particularly like it on rainy days. I treat it like having a high school kid being chauffeur... I am a back seat driver who just happens to be in the driver's seat. as for visual representation like in the video or waymo's demo videos, like many other things in life when you see how the sausage is made it is a wonder how we all survive it. The key difference between Tesla and Waymo is Tesla is not geo fenced, same with Cadillac's supercruise which is not available except on interstate. who has the best solution, I am not willing to place a bet on that yet