4 ms·
Improbable Inspiration: Bayesian Networks (1996)
- new23d 6y ago> Then, in the late 1980s--spurred by the early work of Judea Pearl, a professor of computer science at UCLA, and breakthrough mathematical equations by Danish researchers--AI researchers discovered that Bayesian networks offered an efficient way to deal with the lack or ambiguity of information that has hampered previous systems. The "mathematical equations by Danish researchers", for those interested, are most likely this paper: Lauritzen, S.L. and Spiegelhalter, D.J. (1988), Local Computations with Probabilities on Graphical Structures and Their Application to Expert Systems. Journal of the Royal Statistical Society: Series B (Methodological), 50: 157-194. https://doi.org/10.1111/j.2517-6161.1988.tb01721.x https://doi.org/10.1111/j.2517-6161.1988.tb01721.x Direct PDF Link: https://www.eecis.udel.edu/~shatkay/Course/papers/Lauritzen1988.pdf https://www.eecis.udel.edu/~shatkay/Course/papers/Lauritzen1...
- ta988 6y agoI highly recommend Judea Pearl's book: https://en.m.wikipedia.org/wiki/Causality_(book) https://en.m.wikipedia.org/wiki/Causality_(book) it really is a fascinating read.
- CapriciousCptl 6y agoInteresting to see this underpinned the Office Help System one way or another— probably being the infamous paperclip.
- mensetmanusman 6y agoWay ahead of their time. They just needed 25 more years of Moore’s law...
- cultus 6y agoThere's been enough progress in approximate Bayesian methods that many things can be done thousands of times faster than back then, as well. The reputation of Bayesian methods as being slow is undeserved nowadays.
- p1esk 6y agoCan someone point me to any examples where Bayesian neural networks are successfully used for any practical applications? Like where they are better than regular non-Bayesian NNs? By better I mean better accuracy.
- RavlaAlvar 6y agoBayesian network is not Bayesian neural network
- minkowski 6y agoNot answering your question, but just to point out to readers that this article is about graphical models, not Bayesian neural networks.
- tachyonbeam 6y agoIt's kind of unfortunate that ML has become completely synonymous with neural networks in many people's mind.
- nextos 6y agoVery small datasets and/or where a good uncertainty estimate of predictions is really important.
- tirthapatel 6y agoAFAIK, Bayesian networks are extensively used in biological sciences and economics. Not sure if this will be useful, but I found a survey that discusses these applications: https://www.frontiersin.org/articles/10.3389/fncom.2014.00131/full https://www.frontiersin.org/articles/10.3389/fncom.2014.0013...
- JHonaker 6y agoBayesian network is a synonym for directed graphical model. Any time you see graphical models, they’re usually BNs. Undirected graphical models are very closely related too (all directed models can be represented as undirected models, but not all undirected models can represent directed models), but they’re usually not referred to as BNs. They’re used all over the place. One school of causal inference is heavily steeped in BNs/DAGs. This shouldn’t be surprising because the creator of BNs, Judea Pearl, is heavily involved in causal inference now.
- dmarchand90 6y agoI like how in the mid 90s neural networks were almost a write-off. "But the neural nets won't help predict the unforeseen. You can't train a neural net to identify an incoming missile or plane because you could never get sufficient data to train the system."
- nextos 6y agoThey are almost orthogonal concepts in some regards. Bayesian models (and in particular Bayesian networks or graphical models) and neural networks are about different things. The former try hard to capture uncertainty and causality. The later are all about non-linearity. For example, Pyro implements tons of facilities to have Bayesian models augmented with neural networks. It makes a lot of sense from a modeling perspective to model the big picture using a Bayesian model (generally a graphical model) and then use neural networks for some components. You capture the overall causal structure, but you are also outputting really precise predictions. For example, a deep markov model. There are tons of unexplored ideas combining both, and in general I think this is the future of deep learning and one component towards AGI.
- radomir_cernoch 6y agoIndeed. Mathematically speaking, a graphical model merely formalizes conditional independence. Their advantage is a statistical interpretation, which is also a factor that makes algorithms (like belief propagation) harder to parallelize on GPUs.
- JHonaker 6y agoBelief propagation is hard to parallelize because, in its basic form, it’s a sequence of serial computations. If the graph is a tree, though, then it’s actually pretty easy to parallelize. The only issue is you have to wait for all the incoming messages from your children to arrive before you push your message upward. Once you reach the root, you can fully parallelize the downward pass (in a tree). The other main issue is that in graphs with multiple paths to a single node, you can’t quite do this forward backward pass and get the _exact_ answer. You can approximate it by just passing messages in these cycles that arise, but you’re not actually guaranteed to converge to the right marginal probability anymore. This is called Loopy Belief Propagation. Additionally, there’s no real true ordering of nodes for a passing order. There’s not a natural sequence like there is from leaves to root and back when you can go around and around in circles. It surprisingly works reasonably well in a lot of cases anyway though. BP and the various approximate versions are super interesting. The original algorithms really only work on low dimensional discrete spaces or things with analytic solutions to integrating from conditional/joint to marginal distributions. However, there’s been some really cool stuff coming out in the last 5ish years about using particle based approximations to work on more complicated/continuous spaces.