4 ms·
I appreciate the kind reply, but I don't think you understood the gist of my complaint. I have a degree (just undergrad) in math, and I've implemented Kalman f
by civility 9y ago
I appreciate the kind reply, but I don't think you understood the gist of my complaint. I have a degree (just undergrad) in math, and I've implemented Kalman filters, Kalman smoothers, information filters, particle filters and so on at least a dozen times. I know what operations to perform, and I even have an intuition about why they work.
When I complain about the Bayes theorem way of describing or thinking about Kalman filters, I really mean that I don't understand what almost anything in that expression means. The notation is intuitive if read as English, but opaque to what operations are being performed on what mathematical objects. Capital P means probability of an event or predicate, and A is some abstraction of the thing I'm trying to estimate, and B is the abstraction of a new observation or measurement.
Think of it this way: I have some knowledge of my prior estimate, and I'm able to characterize it as a Normally distributed random variable. It has a mean and a covariance in 3D space (let's call those x and P). There is a multivariate function for that pdf (lowercase p); it's an exponential of the negative square of the Mahalanobis distance normalized to have a unit volume. The same thing applies for my measurement, except it's only a 2D mean and covariance (let's call those z and R). I already know that I can use my H matrix, some matrix inverses and multiplications to combine this information. It's really just a case of multiplying the pdf for the state estimate and the measurement and re-normalizing to a unit volume. In other words, I believe both things are true, and they are independent, so I can multiply their pdfs and re-normalize to get a combined result. The rest is just linear algebra.
Now let's get to Bayes. P(A) is supposed to be the probability of an "event" or predicate. I can squint sideways and translate that as integrate my pdf for the prior estimate (Normal with mean x and covariance P) over a some unspecified bounds and convert that to a probability of my estimate being in those bounds, but I'm not sure that's what is intended. Again a similar thing applies for B (Normal with mean z and covariance R). It's frustrating that the bounds are never stated, because they could be radically different spaces in the numerator and denominator. I guess I'll give that a pass because maybe the notation would be too cumbersome if it was included, but this simplification seems to never be stated in any of the books or papers I've read on it.
Next, P(B|A) has to be probability of another "event", and if I squint again, the best I'm able to come up with is that it means a new Normal pdf with mean = H'z and covariance = H'RH. However, I haven't seen that spelled out any where, and so really that's just assuming the conclusion I want, which is super questionable. It also doesn't help me understand the left side of the Bayes equation - that vertical bar there seems to mean something different. When I look to the definition of conditional probability, I don't see anything about the vertical bar applied to multivariate Normal distributions.
If you translate it all to English, it reads as a coherent sentence, and that's fine. This is the "prior state estimate", that's the "observation", and this other thing is that's the "a posteriori" etc... However, if Bayes really helps with the understanding the math, the vertical bar has to mean something specific as an operator, and one would hope it meant the same thing on the left and right sides of the equation.
Similarly, in order for all of those P( ... ) to become scalars so multiplication and division are well defined, you need some integration bounds to turn the event into a probability. But at this point, I'm not even sure if that's what is intended by the notation - maybe they aren't scalars, and multiplication and division mean something radically different here. I honestly don't know.
- abstrakraft 9y agoProbability notation generally works best for the people who already understand the concept in question. Let me take a crack at your question. The equation in question is P(A|B) = P(B|A) P(A) / P(B). In modern Kalman filter literature, this would be stated as something like: P(x_k | z_k) = P(z_k | x_k) P(x_k) / P(z_k). It is generally left as implicit in these sorts of equations that everything is also conditioned on the sequence z_1 to z_{k-1}. In this equation, x_k is a free variable in the state space (possibly multi-dimensional, so a vector), while z_k is the measurement, which is a realization of the random variable distributed as N(H(x_k), R_k). The result is a PDF over the free variable x_k. So let's tackle the terms one by one: 1. P(z_k | x_k) - this is the probability that we measured z_k, given that the true object is at x_k. This is the aforementioned normal distribution N(H(x_k), R_k). 2. P(x_k) - this is the prior probability of the state estimate, generally after propagation through the motion model from P(x_{k-1} | x_{k-1}). In the Kalman filter, this is also Gaussian, and conveniently has the same mean as the term above (see note below). 3. P(z_k) - this is the denominator that someone else mentioned earlier can be effectively ignored, which is right - you only need it to normalize the numerator. If you must compute it, it can be factored as the integral over the entire state space of P(z_k|x_k)*p(x_k). Given z_k, this is a number, not a function. 4. P(x_k | z_k) - The result, which is a Gaussian PDF. You can arrive at it numerically by plugging in specific values for x_k, in which case (1) and (2) are numbers. Or symbolically, in which case (1) and (2) are functions, and you'll end up with the form of a Gaussian PDF. Note: The original article quotes a distribution for the product of two Gaussians with arbitrary means. It does not state that this is an approximation, which is exact only in the case of equal means. This is why unbiased measurements are one of the Kalman assumptions.
- gugagore 9y agoI think the crux of the complaint of is the imprecision in saying e.g. "P(z_k | x_k) - this is the probability that we measured z_k, given that the true object is at x_k" Technically, you can only give a useful answer about the probability density of the random variable Z_k at the value z_k, conditioned on X_k = x_k. In the Bayesian interpretation of the Kalman filter, you never have an event "I measured z_k" (that event has probability 0, of course). I agree that the probability notation is the issue here. Look at how wikipedia shows Bayes' rule for continuous random variables on both sides of the |. : https://en.wikipedia.org/wiki/Bayes%27_theorem#Random_variables https://en.wikipedia.org/wiki/Bayes%27_theorem#Random_variab... That's the kind of explicit and precise notation I would use to help someone understand the Kalman filter from a Bayesian perspective. Once you use that definition of Bayes' rule, then you can substitute the definitions of the multivariate normal pdf, Do The Math, and derive the Kalman filter recursive updates.