11 ms·
Probabilistic programming does in 50 lines of code what used to take thousands
- newpattern 11y agoAnybody here on HN have experience with probabilistic-programming ? This looks quite disruptive if it works.
- eximius 11y agoReally? "Disruptive"?
- vezzy-fnord 11y agoConstraint solvers and similar tools can look pretty magical to people who use them for the first time, so I don't blame the OP. What's more interesting is that we're seeing heightened interest in these techniques again after they were ostensibly sidetracked in favor of statistical methods.
- sidrajaram 11y agoI'm no expert, but afaik probabilistic programming isn't a new method or technique. It is just wrappers around existing statistical techniques, as an attempt to divorce the details of inference algorithms with model specifications. I'm not buying in just yet, because although it's nice to talk about model specification as completely independent processes, the availability of fast inference algorithms sometimes dictates what models you should choose. Sometimes less exact models with a larger parameter space that allows you to crunch orders of magnitude larger datasets (with approximate inference algorithms) yield more useful results than better specified models...and sometimes not. The thing if one still needs to know the whens and whys of picking certain models over others, and can't just gloss over the inference details.
- sgt101 11y agoThe questions are "when does it fall off a cliff" and "does it work for more than one thing". Normally the answers are "as soon as you stop doing toy problems" or "no". But, this (and Church) look very interesting.
- troels 11y agoMaybe I just don't get the terminology right, but isn't this exactly a statistical method?
- deleted 11y ago[deleted]
- oh_sigh 11y agoSure. If someone here has nothing more than high school math under their belt, they may not understand that this isn't a huge deal. For example, I refer you to the paper wherein some biologists "discovered" integrals in 1993. http://care.diabetesjournals.org/content/17/2/152.abstract http://care.diabetesjournals.org/content/17/2/152.abstract
- nanofortnight 11y agoIf you want to know about the history of similar language paradigms, look at logic programming. Similar to probabilistic programming, it is declarative, utilises an engine at run time, the user builds a model of the word then queries it, and it uses fewer lines of code to express for problems within a limited domain. Hardly "distruptive". You should really think of them as DSLs that people actually use. If you're intrested in such things, perhaps you should take a look at PRISM, which embeds a probabilistic framework within B-Prolog. http://rjida.meijo-u.ac.jp/prism/ http://rjida.meijo-u.ac.jp/prism/
- nabla9 11y agoProbabilistic programming is way to increase flexibility of statisticall modelling. You only write the code that generates data and give statistical model and possibly some parameters to fix approximations used in estimation and get out model from that. For example, Stan-language supports MCMC modelling using Hamiltonian dynamics. http://mc-stan.org/ http://mc-stan.org/
- king_magic 11y agoJust like UML and Prolog.
- BenoitEssiambre 11y agoI've played with it a bit and, in my opinion, the principles behind it at least, the streamlined and optimized simulation of bayesian generative models is the best chance we have to solve artificial general intelligence. Reading probmods.org and dippl.org made me go from being very pessimistic I would see it in my lifetime to a solid maybe.
- murbard2 11y agoIt's a very powerful statistical technique, yes, but I doubt it will be enough for AI. The problem is that you need to sample efficiently from the posterior distribution, and for anything AI related, MCMC isn't going to cut it. Let's take a step back and consider a logical problem. Can you put N socks in N-1 boxes such that no box contains more than one socks? Obviously not, it's the pigeonhole principle. Convert that question into a Boolean circuit, and throw a SAT solver at it. It will die. In fact, using only first order logic, a proof of the pigeonhole principle requires an exponential number of terms. Looking at logical propositions alone is too myopic to solve the problem, you have to formulate higher level theories about the structure of the problem to solve it (in this case, natural numbers). The same goes for probabilistic programming. As long as the paradigm is to treat the problem as a black box energy function to sample from, it is doomed to be inefficient. Try writing a simple HMM and the system will choke, even though there are efficient algorithms to sample from such a model. If you look at deep learning techniques, they take an interesting approach which is to learn to approximate the generative distribution and the inference distribution at the same time. This is the basis of the work around autoencoders, deep belief networks, and it guarantees that you can tractably sample from your latent representation.
- BenoitEssiambre 11y agoI have been out of the machine learning field for years and I haven't look into deep learning methods even though they are intriguing (just not sufficiently bayesian to pique my curiosity enough to spend my limited free time on it). I do believe that automatically building a hierarchical structure (as I assume happens in deep learning) is the way to abstract away complexity but I think this is achievable within a generative bayesian monte carlo approach. For example, I have been toying with using MCMH similarly to how it is used in dippl.org to write a kind of probabilistic program that generates other programs by sampling from a probabilistic programming language grammar. A bit like with the arithmetic generator here: https://probmods.org/learning-as-conditional-inference.html#example-inferring-an-arithmetic-expression https://probmods.org/learning-as-conditional-inference.html#... but for whole programs. After the MCMH has converged to a program that generates approximately correct output, you can tune grammar hyperparameters on the higher likelihood output program samples so that next time it will converge faster. I don't know if this counts as approximating "the generative distribution and the inference distribution at the same time" under your definition but my hope is that the learned grammar rules are good abstractions for hierarchical generators of learned concepts. Of course the worry is that my approach will not converge very fast but there are reasons to think that having a suitably arranged hyperparametrized probabilistic grammar might use the Occams' razor inherent in bayesian methods to produce exponentially fewer, simpler grammar rules that generate approximations when it doesn't have enough data to converge to more complex and precise programs and that these simple rules which rely on fewer parameters might provide the necessary intermediate steps to then pivot to more complex rules. These smaller steps help MCMH to find a convergence path. Not sure how well it will work for complex problems however. My idea is still half baked. I have ton's of loose ends and details I have not figured out, some of which I might not even have a solution, as well as little time to work on this (this is just a hobby for me). Anyways, all that to say that probabilistic programming can go beyond just hardcoding a generative model and running it.
- ub 11y agoNo experience but looks like they are organizing a summer school on probabilistic programming languages. http://ppaml.galois.com/wiki/wiki/SummerSchools/2015/Announcement http://ppaml.galois.com/wiki/wiki/SummerSchools/2015/Announc...
- pliny 11y agorelevant paper: http://mrkulk.github.io/www_cvpr15/1999.pdf http://mrkulk.github.io/www_cvpr15/1999.pdf
- e12e 11y agoThanks for that. As far as I can tell the compiler for the Picture language hasn't been published?
- stevenspasbo 11y agoIs that 50 lines of code, or 50 lines of using a library that's thousands of lines of code?
- aet 11y agoI agree -- I think lines of code is the wrong measure. A better measure would be how large is the set of problems that occur in practice that normally would take thousands of lines which can now solve be solved with 50 lines of code using a probabilistic programming language.
- nanofortnight 11y agoGenerally it's anything that you can express using a probabilistic graphical model. It's mainly useful in engineering (real engineering), control systems, robotics, inference, machine learning and artifical intelligence circles.
- baldfat 11y agoWhat do we say about languages built on C? Is it 100 lines of code but there are hundreds of thousands of lines of code for that higher level language you just coded? I don't think libraries count in terms of code. We all use code to program. Standing on the shoulder that preceded us. Using a library and a function should just count for the most part.
- jarrettc 11y agoTrue, but the parent commenter is getting at something important. The article suggests that researchers have found a new, much more concise way to express the solutions to difficult problems. That's different from a library, which merely packages pre-built solutions to a finite set of problems. It's like the difference between a complete kitchen that fits in your pocket and an iPhone app that lets you order a burrito. The article suggests something like the former. A library which encapsulates 1000 lines of code into a single function call is like the latter.
- jb55 11y agoIf anyone wants to jump into this, Josh Tenenbaum and Noah Goodman put together this amazing interactive book for learning probabilistic programming with Church: https://probmods.org/ https://probmods.org/
- rrtwo 11y agoGreat project but very weird choice of language. How probable is it (pun intended) that the person coming to learn about probabilistic programming would already know functional programming?
- Retra 11y agoIsn't functional programming a standard part of any computer science curriculum? Why would you expect programmers not to know it?
- g0wda 11y agoNo it's not, in most of the world.
- Widdershin 11y agoNot everyone does a computer science degree.
- saryant 11y agoI was very surprised to learn from two Stanford CS alums that functional programming was not a requirement in that program.
- vidarh 11y agoEven the places where it is a part of the curriculum, it is often such a small part that unless people specifically take courses related to functional programming you can't expect them to be able to actually use it. Or remember much of it for that matter. Heck, I spent months on a binge reading functional programming research papers, and it still doesn't mean I know any functional languages other than very superficially.
- cheatsheet 11y ago> “It goes beyond image classification — the most popular task in computer vision — and tries to answer one of the most fundamental questions in computer vision: What is the right representation of visual scenes? Can someone knowledgeable in graphics research explain the context that this question comes from? If I am reading the question correctly, I infer that the question suggests that there exists a right way to reproduce the visual experience of reality. To me, this sounds like a question that is equally valid to have no answer (or many answers) in aesthetics, art, and philosophy, etc.
- darkmighty 11y agoThe question is fundamental to all kinds of recognition: recognizing the invariants of the scene, the data that distinguishes it from other scenes, which is very close to the definition of Shannon information. For example, if you can extract a 'Mesh' from a 2D picture, you can generate many other view points, and that mesh can be considered a good representation. If you are more sophisticated however (and perhaps have a larger "dictionary"), you can instead extract 'There are two wooden chairs 1m from each other, ...'. That's the sense in which the representation is fundamental to computer vision -- it distills what the system knows (or what it wants to know) about scenes. The more concise the representation without loss of information the smarter your system is (and past a point becomes a general AI problem).
- nabla9 11y agoYou can figure out statistical structure of natural images (or just faces) and derive efficient representations with similar properties as to those observed in the visual system of the brain. See for example: Natural Image Statistics — A probabilistic approach to early computational vision https://www.cs.helsinki.fi/u/ahyvarin/natimgsx/ https://www.cs.helsinki.fi/u/ahyvarin/natimgsx/
- rasz_pl 11y agoThink about Dreaming. "seeing" during a dream state works by experiencing pure data representation of the real world. People fluent in lucid dreaming can tell you something funny happens when you try to thorough examine objects while sleeping. Constructed worlds tend to be skin deep, and fall apart when poked. Everything is build with ideas drawn from your experience. Its Plato's Allegory of the Cave all the way down. Imagine "watching" a movie compressed using your very own prior knowledge. Every scene could be described in couple of hundred lines of plaintext. Today we do this by reading a book :) What if we could build an algorithm able to render movies from books?
- dang 11y agoAlso https://news.ycombinator.com/item?id=9363496 https://news.ycombinator.com/item?id=9363496 from yesterday.
- contingencies 11y agoProbabilistic prediction: this is primarily going to be used for robots that monitor, kill or assist with killing people.
- Retra 11y agoAlternatively, one could monitor, save, or assist with saving people.
- murbard2 11y agoYes, you can specify extremely powerful statistical models in only a few lines of code using probabilistic programming. However, at this point, unless you design your program in a very specific way and use a lot of tricks, your sampler is very unlikely to converge, and you won't get any meaningful result without a gargantuan amount of computing power.
- eli_gottlieb 11y agoAh, so they wrote a DSL just for doing Bayesian inverse-vision models, and apparently they've now got it to accuracy rates competitive with most other major vision methods? Good job!