3 ms·
The idea of sensor fusion is to have some kind of error model for all sensors (and frequently for navigation also a dynamic model of the vehicle) so that you ca
by bafe 3y ago
The idea of sensor fusion is to have some kind of error model for all sensors (and frequently for navigation also a dynamic model of the vehicle) so that you can weigh the various sensor differently based on the uncertainty. The most primitive incarnation of this would be a simple linear Kalman Filter, but you can use similar concepts with more complex non-linear observation and dynamic models
- nihzm 3y agoExpanding a bit for those who don't know how a Kalman filters (or any Bayesian recursive estimator for that matter) work here's the essential idea. Each sensor is given an uncertainty model, that is for example a stochastic model by adding e.g. Gaussian noise to the "true" value it is measuring. Further, you have another model that describes the dynamics, e.g. equations from physics can tell you where the car will go if you know the current speed, position, etc.. 1. The kalman filter computes a probabilistic prediction of what it thinks is happening by using the dynamics model. That is, based on what it knows so far, where will the car (probably) be when the next measurement comes up? 2. When measurements from various sensors come in, the Kalman filter uses Bayes's theorem to compute a mean (posterior), in which each measurement is weighted by the probability that the measured value is correct (using the uncertainty models; "correct" here means "in agreement with the prediction"). In other words, sensor that are inaccurate (large variance) will be considered less in the computation of the mean, while more accurate sensors are given more importance. Once the mean of the measured quantities are computed they are used again in step 1 and the whole thing is repeated. As you can see, that disagreements are justified by the inaccuracies of the sensors, and the process of performing a probabilistic weighted average solves the problem. For the Kalman filter in particular, it can be be show (mathematically proved) that this process minimizes the variance (uncertainty) of the measured quantities (which btw. is an amazing result if you think about it).
- bafe 3y agoThanks for the excellent summary on the Kalman filter! I admit I was too lazy to write any details.
- sebastos 3y agoThat is the general idea, but now you've just descended to the next level of the iceberg. While it's true that sensor fusion is a Good Idea in lots of contexts, the whole paradigm rests on a fundamental assumption that you can model the problem domain. Perception pushes up against the boundaries of what Bayesian filtering can handle because even the most complex error models are hopelessly simplistic compared to the system generating your measurements. Your model would have to capture the bounds of uncertainty for what your stereo pair might say about the depth of any point in any scene it will ever see. A simple example of why this gets complicated: if I have a point in one camera and a point in another camera and I know they correspond to the same real-world spatial point, I can calculate some distance, and the statistics of that calculation can be captured by a halfway-reasonable error model. But how did I know they corresponded to the same point in the first place? Well, because they look the same according to some image feature... or because some deep neural network told me so.. etc. There just aren't very good ways to model just how haywire ^that^ process can go. So at the end of the day, once you let this evil into your perception system, using statistics to blend your sensors together is undermined, and all of your precious covariances just turn into tuning knobs you can twiddle. The dirty secret is that almost all robotic perception systems are hiding unprincipled, un-modelled heuristics in the data association process. This is kicked under the rug because it doesn't really fit into traditional estimation theoretic frameworks. In a lot of papers you'll see academics push it aside by just calling that the "front-end", which they brush aside as a little widget you put on the front. If you're lucky they'll do ablations across a couple different options. Of course, this is just one level deeper down the iceberg. It goes far deeper. Even if you could model the statistics of a depth camera well, the statistics of "what are all the objects in your scene about to do" is another couple of orders of magnitude more un-modellable. Often engineers will do something like attach a "constant-velocity" model to the agents in a scene. Imagine trying to bin all of the reasons you might stop walking in a straight line into a bubble that describes how "noisy" that picture of the world is! Now you can begin to appreciate just how hopeless it is to explicitly model uncertainty in the world around us.
- joshuamorton 3y agoYou're basically describing the no-free lunch theorem, but in practice we can define fairly robust generic models for certain things, like "physics". That's why optical mice work pretty well. You can start to add unjustified assumptions and that'll make the world model weaker, yes. But starting with pretty basic assumptions like "you can segment an object from an series of images because each solid object will move on its own trajectory", or even more basic like "objects have edges" and then a few dozen samples per second, and suddenly you have a fairly robust way to detect things. Same for predicting where something will go. If you can estimate an objects current velocity, acceleration, and jerk with reasonable precision, you don't really need a highly predictive heuristic for the world model. For decision making you need more robust heuristics, like "the car to my left has right of way at the stop sign", but you don't need that level of heuristic to identify that there is a car and that it isn't part of the pavement and that it is currently sitting still.