4 ms·
Tangentially, I have a concern with neural-symbolic hybrids, surely someone can address it. Do you think that symbols produced by a neural network (in a complex
by skdotdan 7y ago
Tangentially, I have a concern with neural-symbolic hybrids, surely someone can address it. Do you think that symbols produced by a neural network (in a complex enough task) will ever be comprehensible by a human? Because my intuition here is that the symbols will look super complex and even random, and actually they will be just a combination that "just works" but we won't know why, just as other deep learnng models. Instead of arbitrary floats (weights), we will have arbitrary chains of symbols.
- nicklovescode 7y agoYou might enjoy https://distill.pub/2018/building-blocks/ https://distill.pub/2018/building-blocks/. I’m on the ml interpretability team at OpenAI. Happy to answer any questions you have!
- skdotdan 7y agoIt looks super interesting, thanks! If I have any questions I'll let you know ;) (By the way, probably you have the coolest job on earth)
- taliesinb 7y agoIt depends how those symbols are encoded. Techniques like attention, and systems like transformers that are built on top of them, are often produce highly interpretable execution traces simply because their patterns of activity are very revealing of how they are going about solving the task. It's harder to interrogate their learned weights in any free-standing way, of course. But the neuro-symbolic concept learning paper I gave already shows the potential translucency of these kinds of hybrid systems: the linguistic interface (its a VQA task) allows one to simply look up the feature vectors and programs associated with particular phrases or English nouns. Similarly, the scene is parsed in an explicitly interpretable way, with bounding boxes for the various objects on which reasoning will commence. This 'bridge' theme between natural language and the underlying task space is really powerful, and it probably makes sense to figure out how we include them for systems that have nothing to do with natural language. https://arxiv.org/abs/1901.11390 https://arxiv.org/abs/1901.11390 contains another great example of how interpretable such models can be, especially if they are generative. Take a look at those segmentations! Lastly, https://arxiv.org/abs/1604.00289 https://arxiv.org/abs/1604.00289 lays out this vision in a lot more detail.
- skdotdan 7y agoMany thanks for your answer!
- hadsed 7y agoThink of it like discovering a new human language. We use logic along with knowledge about the world to triangulate words that deal with concepts. ML models work in the same way. Difference is they aren't human so it's harder to just use your empathy muscles (though even humans cut off from the rest of the world can come to pretty wild perspectives and ways of thinking). But they are logical, in some respect, and they are modeling our human perspective on some process in the world. But as a sibling comment posted in a link to the Distill paper, we need a lot of tools to make that process easier. For example, very often researchers will probe single neurons and find they do something we find conceptually understandable, like a neuron detecting when you're inside a parenthetical when generating text so that the parenthesis is eventually closed. I'd expect the "symbols" to be very similar, because after all the neurons are symbols too. Both require you to either relate it to a small concept or put some together to create bigger concepts (or both).
- skdotdan 7y agoThanks!
- YeGoblynQueenne 7y agoThat's a good point that is not often discussed. I've actually done some work into this kind of problem- how to explain automatically defined predicates in symbolic, logic-based machine learning: https://github.com/stassa/metasplain https://github.com/stassa/metasplain So that's "metasplain" a little program that explains "invented predicates", which is what I say above, predicates that are automatically constructed by a symbolic machine learning system in the process of learning. Think of them as invented features that are relevant to the learning task. These are given automatic names so they're difficult to read, especially if you have lots of them. Metasplain starts by automatically assigning meaningful names to invented predicates by combining the (not invented) symbols of their literals, then asks the user for improved names. It can go all the way automatically, without interaction, but the results are a bit meh. With a human in the loop you get the best of both worlds. And that's what I think is the best way to solve interpretability problems in machine learning: instead of automating them, which is like trying to create a chicken so you can get an egg so you can hatch a chicken, put the human back in the loop and make it easy for her to provide meaningful explanations, even if she's not an expert.
- skdotdan 7y agoThanks, this direction looks promising and makes a lot of sense.