8 ms·
Wittgenstein’s theories are the basis of all modern NLP
- mlucy 8y agoIt's really difficult to overstate how important embeddings are going to be for ML. Word embeddings have already transformed NLP. Most people I know, when they sit down to work on an NLP task, the first thing they do is use an off-the-shelf library to turn it into a sequence of embedded tokens. They don't even think about it; it's just the natural first step, because it makes everything so much easier. In the last couple years, embeddings for other data types (images, whole sentences, audio, etc.) have started to enter mainstream practice too. You can get near-state-of-the-art image classification with a pretrained image embedding, a few thousand examples, and a logistic regression trained on your laptop CPU. It's astonishing. (Note: I work on https://www.basilica.ai https://www.basilica.ai , an embeddings-as-a-service company, so I'm definitely a little bit biased.)
- jandrese 8y agoIt's an exciting time for sure. To the layman this feels like the first real progress we've had towards AI since the 70s. It seems like the field kind of wandered off into the realms of pure mathematics for a few decades with little tangible progress, but now we're getting stories every few weeks about how computers can recognize objects in pictures or compose new music or whatever.
- madhadron 8y agoWhat I find particularly neat are the non-Euclidean embeddings, such as hyperbolic spaces to generate hierarchies.
- minkzilla 8y agoDo you know of any good resources for learning about such things for someone with a cursory of Word2Vec ?
- Radim 8y ago"Implementing Poincaré Embeddings" (hyperbolic embeddings implemented in Gensim): https://rare-technologies.com/implementing-poincare-embeddings/ https://rare-technologies.com/implementing-poincare-embeddin...
- DoctorOetker 8y agojust like minkzilla I would love to read more about this
- akozak 8y agoFiguring out how to process context is important for NLP, no question. But I think this is probably wrong on Wittgenstein. I'm pretty sure his entire point in the Philosophical Investigations was that "meaning" is exactly NOT probabilities of symbol co-occurrence, or just names of objects in the world. Symbols acquire meanings from their use by humans. Accounting for context in NLP via probabilities of occurrence might be useful in better reproducing language, but we should be careful not to say that this is the essence of meaning and language.
- whatshisface 8y ago>Accounting for context in NLP via probabilities of occurrence might be useful in better reproducing language, but we should be careful not to say that this is the essence of meaning and language. Yes, and the article actually includes evidence in favor of this and against its own conclusion. It mentions that vector "cat" is closer to vector "dog" than vector "dog" is to vector "dogs," which makes sense if you interpret it as a measure of appearance in sentences but no sense at all if you force it into the mold of "the meaning of words."
- akozak 8y agoYes - but for me it was this paragraph: > And it’s now quite clear where the Wittgenstein’s theories jump in: context is crucial to learn the embeddings as it’s crucial in his theories to attach meaning. In the same way as two words have similar meanings they will have similar representations (small distance in the N-dimensional space) just because they often appear in similar contexts. So “cat” and “dog” will end up having close vectors because they often appear in the same contexts: it’s useful for the model to use for them similar embeddings because it’s the most convenient thing it can do to have better performances in predicting the two words given their contexts. I am actually fine to say that this approach is useful and convenient - and that we can fairly call measuring probabilities of co-occurrence measuring "context" in some sense. But "context" for Wittgenstein in his account of meaning was clearly not word or symbol occurrences. It was a much broader view of the way that language fits in with human intentions and behavior and the wide variety of uses for a word. I hate to quote Wikipedia, but from the PI article: "Wittgenstein argues that definitions emerge from what he termed "forms of life", roughly the culture and society in which they are used." https://en.wikipedia.org/wiki/Philosophical_Investigations#Meaning_and_definition https://en.wikipedia.org/wiki/Philosophical_Investigations#M...
- nostrademons 8y agoIt's interesting how different this is from 10 years ago, when Chomsky's theories were the basis of all modern NLP, or even 5 years ago, when most NLP used a hybrid of formal grammars + embeddings. I remember attending a tech-talk on part-of-speech tagging in 2011; the state-of-the-art then was a probabilistic shift-reduce parser where the decision to shift vs. reduce at each node was done by a machine-learned classifier.
- ppod 8y agoWittgenstein emphasized meaning as context and usage before Chomsky, but the actual method was first properly investigated by structural linguists such as JR Firth and Zelig Harris, who was Chomsky's supervisor. Good articles here: https://en.wikipedia.org/wiki/Distributional_semantics https://en.wikipedia.org/wiki/Distributional_semantics https://aurelieherbelot.net/research/distributional-semantics-intro/ https://aurelieherbelot.net/research/distributional-semantic...
- southerndrift 8y ago>As human beings speaking English it is quite trivial to understand that a “dog” is an “animal” and that is more similar to a “cat” than to a “dolphin” but this task is far from easy to be solved in a systematic way. Are they? A dog can be trained like a dolphin, unlike a cat. In the context of training, dogs are more similar to dolphins.
- philippoi 8y agohttps://drsophiayin.com/blog/entry/a_super-simple_method_for_training_cat_tricks/ https://drsophiayin.com/blog/entry/a_super-simple_method_for...
- AaronFriel 8y agoYes, I suspect you had to cherry-pick a dimension in which dog and dolphin are closer than dog and cat. That would defy conventional wisdom justifies that dogs are closer to cats than to dolphins, but that's also modeled by word vector embeddings. In the metric space, the distance between dog and cat might be lower than dog and dolphin in many dimensions, but higher in this specific one. A general distance function will have to take all of the dimensions into account, not just those cherry-picked. So the conventional wisdom _and_ your personal belief are both accounted for, and in the context of training "dog" and "dolphin" might be more similar. I still suspect that's not actually true, and I'd be really surprised if a survey of users found dog and dolphin to be closer than dog and cat in _any_ dimension.
- lettergram 8y agoFor those interested, I recently wrote a guide on using neural networks for NLP[1]. I wrote the guide with the explicate goal of trying to help the people understand NLP (sentence classification) without the need to understand the math. I cover word embeddings: https://austingwalters.com/word-embedding-and-data-splitting/ https://austingwalters.com/word-embedding-and-data-splitting... As well as FastText: https://austingwalters.com/fasttext-for-sentence-classification/ https://austingwalters.com/fasttext-for-sentence-classificat... Hope someone finds it useful. [1] https://github.com/lettergram/sentence-classification https://github.com/lettergram/sentence-classification
- kolbe 8y agoI am really struggling to find where Wittgenstein fits into any of this at all. >And it’s now quite clear where the Wittgenstein’s theories jump in: context is crucial to learn the embeddings as it’s crucial in his theories to attach meaning. That's not at all clear to me. The crucial part of W's tome is that two sentient beings are knowingly engaging in a game where they have 'agreed' on meanings. My guess from reading Philosophical Investigations is that W would only think NLP were possible in formal settings like law, where all players of the game know the rules quite well, and the program could be trained as if it were a player in that game.
- sp332 8y agoI think the point is that the only way to learn what a word means is to see how it is used. Trying to define a word from some kind of first principles, dictionary-style, is not going to be very effective. The best way for a computer to learn what words mean is to analyze a lot of real-world data. I would love for a computer to be able to ask questions, or at least surface marginal cases for more training, but that seems to be a very uncommon feature at least in these toy examples.
- akozak 8y agoThe issue isn't that this approach won't be useful in building systems we can interact with linguistically. The problem is in describing the system as having learned a meaning. It might seem pedantic or like something only philosophers of language would care about. But it gets to the core of how we should talk and think about the nature of AI as NLP gets more and more sophisticated.
- sp332 8y agoWell it may not be very satisfying, but Wittgenstein's point is that there isn't anything more to understand about the meaning of words than the ability to use words effectively. http://existentialcomics.com/comic/268 http://existentialcomics.com/comic/268
- atrudeau 8y agoThese older word embedding models (word2vec, GloVe, LexVec, fastText) are being superseded by contextual embeddings ( https://allennlp.org/elmo https://allennlp.org/elmo ) and fine-tuned language models ( https://ai.googleblog.com/2018/11/open-sourcing-bert-state-of-art-pre.html https://ai.googleblog.com/2018/11/open-sourcing-bert-state-o... ). These contextual models can infer that "bank" in "I spent two hours at the bank trying to get a loan" is very different from "The ocean bank is where most fish species proliferate."
- libertas 8y agoI would think that the tractatus would be more useful to an AI. But Witgebstein's remarkable ability to shift the paradign and over extend into a meta level of analysis seems similar to the way alpha mind and Leela play chess. The tools W uses to understand perception have a more probabilistic and irrational nature then the tools he uses in his previous work. As if he realized that human communication cannot be considered as a closed and finite system, hence I cannot see how his ideas are implemented in these applications, yet.
- sswaner 8y agoYes, and considering the Tractatus as a framework for a closed and finite set of linguistic rules, such as a domain specific language has great applicability. For example, I flipped open my copy (yes I keep a copy on my desk) and opened to 4.122: rules to indicate internal and external relations between objects. Almost reads like a system requirements document.
- perfmode 8y agoCan someone ELI5 the term "embedding"?
- mlucy 8y agoA word embedding transforms a word into a series of numbers, with the property that similar words (e.g. "dog" and "canine") produce similar numbers. You can have embeddings for other things, such as pictures, where you would want the property that e.g. two pictures of dogs produce more similar numbers than a picture of a dog and a picture of a cat.
- perfmode 8y agoAh. Sounds like a vector space. How does one select a basis?
- leereeves 8y agoIt is indeed a vector space. You don't really choose a basis, an ML tool like word2vec [1] does. And like most advanced applications of ML, exactly how it works is a mystery. 1: https://en.wikipedia.org/wiki/Word2vec https://en.wikipedia.org/wiki/Word2vec > The reasons for successful word embedding learning in the word2vec framework are poorly understood. Goldberg and Levy point out that the word2vec objective function causes words that occur in similar contexts to have similar embeddings (as measured by cosine similarity) and note that this is in line with J. R. Firth's distributional hypothesis. However, they note that this explanation is "very hand-wavy" and argue that a more formal explanation would be preferable.
- alextp 8y agoThe historic picture makes a little more sense (though this is not something a 5yo would understand). We call these things embeddings because you start with a very high dimensional space (image a space with one dimension per word type, where each word is a unit vector in the appropriate dimension) and then approximate distances between sentences / documents / n-grams in this space using a space with much smaller dimensionality. So we "embed" the high dimensional space in a manifold in the lower dimensional space. It turns out though that these low dimensional representations satisfy all sorts of properties that we like which is why embeddings are so popular.
- andybak 8y agoI really wish NLP didn't have two common meanings.
- jeromebaek 8y agoThe author has seriously misunderstood Wittgenstein's contributions to philosophy of language. >And it’s now quite clear where the Wittgenstein’s theories jump in: context is crucial to learn the embeddings as it’s crucial in his theories to attach meaning. Yes, Wittgenstein said context is important for meaning, but that is hardly his unique or even most important contribution to philosophy of language. Wittgenstein's real contribution is in showing that meaning cannot be pinned down like butterflies under glass -- that meaning spontaneously arises in each playthrough of a language-game, and that any effort to find a "canonical", "authoritative" definition is grasping at an illusion. But word embeddings try to do almost exactly what Wittgenstein says is an illusion -- trying to pin down a canonical n-dimensional vector for each word. To correspond with Wittgenstein's theory, there cannot exist any mapping from a word to a vector. Perhaps each vector can be dynamically changing in a by principle uncomputable way. But to get there we are going to need a lot more advances than the state of the art NLP.
- akozak 8y agoThat's a great way to put it! It doesn't mean the approach isn't useful for building systems that we can interact with linguistically, just that we shouldn't kid ourselves into thinking the model has captured meaning.
- visarga 8y ago> Perhaps each vector can be dynamically changing in a by principle uncomputable way. The BERT language model does dynamic (contextual) embeddings and is state of the art in NLP. https://towardsdatascience.com/bert-explained-state-of-the-art-language-model-for-nlp-f8b21a9b6270 https://towardsdatascience.com/bert-explained-state-of-the-a...
- jeromebaek 8y agoI don't think we are using the same definition of the word "dynamic" here.
- idoubtit 8y agoIn what sense are these theories a "basis" to NLP? Did they have any influence? Do they bring any practical contributions? I suspect a slight similarity between popular domains (Wittgenstein and NLP) was contrived into an article that seems very light on the W part. The "Wittgenstein’s theories" that appear here is just that "the meaning of a word is its use in the language". If such a plain concept was all of Wittgenstein’s theories, he would be long forgotten. For centuries, dictionaries have presented words through one or several explanations as well as quotes and examples. 150 years ago, Émile Littré wrote a wonderful French dictionary that contains 80,000 words and about 300,000 literary quotes. He knew no word has a simple and permanent meaning, and that one needs to know many real world contexts to get a fine view on a word.
- ttctciyf 8y agoA while ago, I commented here[1] to the effect that there's a good fit between the treatment of meaning in Wittgenstein's Philosophical Investigations (PI) and the neural net / connectionist approach, calling it "decentralised, statistical, emergent" in contrast to more cognitivist ideas. I was challenged to justify my buzzword-laden characterisation, and I reproduce my response here as it seems relevant to your question: > Right at the start of PI, the "Augustinian picture" of meaning is set out: "words in language name objects - sentences are combinations of such names" a picture where "every word has a meaning. This meaning is correlated with the word. It is the object for which the word stands." And so (later, #81) we "think that if anyone utters a sentence, and means or understands it, he is operating a calculus according to definite rules." This is the view of language which PI aims to - sorry, buzzword incoming - disrupt. > By contrast, PI puts the case that there is no central unifying model applicable to all instances of language use, rather: "I am saying that these phenomena have no one thing in common which makes us use the same word for all - but that they are related to one another in many different ways, and it is because of this relationship, or these relationships that we call them all language." (PI #65) > There follows a discussion of vagueness. In the Augustinian picture where the meaning of a word is an object, and a calculus of these objects is performed, it is difficult to avoid the consequence that meanings are exact. PI uses the example of defining the word "game": "How should we explain to someone what a game is? I imagine that we should describe games to him, and we might add: "This and similar things are called 'games'". And do we know any more about it ourselves? [...] But this is not ignorance. We do not know the boundaries because none have been drawn. [...] we can draw a boundary - for a special purpose. Does it take that to make the concept usable? Not at all! [...] One might say that the concept 'game' is a concept with blurred edges." > So much for "decentralised" and "statistical". As far as "emergent" goes, I think even with the large variance in readings of PI it's uncontroversial to say that it seeks to ground meaning and understanding in relation to "customs" or social practices rather than in some variety of metaphysical correspondence between language and reality required by different variations of the Augustinian picture. In this sense, meaning emerges from the use of words relative to these cultural forms. > Connectionism, as an investigative paradigm, (oops, buzzword!) is simpler (I believe) in that it doesn't require the identification of an actual realised model (or "mental mechanism") such as a neural encoding of a "language of thought" or cognitive frames, etc., in the brain - it "just" requires that a bunch of simple elements can result in complex rule-following behaviour without needing to explicitly encode the rules. Hopefully the quotes above will go some way to indicate how this programme is philosophically somewhat in tune with PI. > (Indeed the extensive sections on samples and teaching language games are eerily reminiscent of descriptions of training neural nets, now that I think of it... "How do I explain the meaning of 'regular', 'uniform', 'same' [...] if a person has not yet got the concepts? I shall teach him to use the words by means of examples and by practice - And when I do this I do not communicate less to him than I know myself." (PI #208)) 1: https://news.ycombinator.com/item?id=10158214 https://news.ycombinator.com/item?id=10158214
- mlthoughts2018 8y agoOne interesting concept I read in Wittgenstein was the idea of decomposing a word into its constituent parts. I’ll use the term broom for it because that was the classic example and also the motivation for David Foster Wallace’s novel “Broom of the System.” So you take “broom” and you could decompose it into “handle” and “bristles”. But then you could decompose it more, by recursively decomposing “handle” into “grains of wood” and “bristles” into “pieces of fiber” (or whatever). You keep doing this ad infinitum, I guess on down to the summation of a bunch of quarks or whatever. The question of interest to Wittgenstein was where does this process bottom out. What would it mean, either physically or semantically, to have a word identifying a concept that could not be broken down into further constituent parts. Wittgenstein was interested in this for the philosophy of language. But I got interested in it by thinking about the decomposition as a mathematical operator, D(“broom”) = {“handle”, “bristles”} and then asking what it could mean if this operator D had an “eigenvector” with an “eigenvalue” of 1, so that Dx = x for some non-decomposeable word x. In some ways, you can see how it could relate to things like word2vec and embedding representations if you could represent a decomposition operator, and define a hierarchical relationship of words as an ordering of how to more or less specifically decompose a word’s representation.
- NotAnEconomist 8y agoI've always wondered if 'exists' is something like that -- and hence why it can't be a property. (Well, if you believe that Kant guy.) You can sort of think of all objects -- broom, bristles, quarks, etc -- as being codata that decomposes to some version of "existence existing", an interference pattern of some fundamental object self-interacting.
- deleted 8y ago[deleted]
- KasianFranks 8y agoInaccurate. This is absurd. Epigraphy is the basis of all modern NLP/NLU. Add computational epigraphy, neuroscience, linguistics and cognition. Ref: Word2Vec is based on an approach from Lawrence Berkeley National Lab ""Google silently did something revolutionary on Thursday. It open sourced a tool called word2vec, prepackaged deep-learning software designed to understand the relationships between words with no human guidance. Just input a textual data set and let underlying predictive models get to work learning." “This is a really, really, really big deal,” said Jeremy Howard, president and chief scientist of data-science competition platform Kaggle. “… It’s going to enable whole new classes of products that have never existed before.” https://gigaom.com/2013/08/16/were-on-the-cusp-of-deep-learning-for-the-masses-you-can-thank-google-later/ https://gigaom.com/2013/08/16/were-on-the-cusp-of-deep-learn... Spotify seems to be using it now: http://www.slideshare.net/AndySloane/machine-learning-spotify-madison-big-data-meetup http://www.slideshare.net/AndySloane/machine-learning-spotif... pg 34 But here's the interesting part: Lawrence Berkeley National Lab was working on an approach more detailed than word2vec (in terms of how the vectors are structured) since 2005 after reading the bottom of their patent: http://www.google.com/patents/US7987191 http://www.google.com/patents/US7987191 The Berkeley Lab method also seems much more exhaustive by using a fibonacci based distance decay for proximity between words such that vectors contain up to thousands of scored and ranked feature attributes beyond the bag-of-words approach. They also use filters to control context of the output. It was also made part of search/knowledge discovery tech that won the 2008 R&D100 award http://newscenter.lbl.gov/news-releases/2008/07/09/berkeley-lab-wins-four-2008-rd-100-awards/ http://newscenter.lbl.gov/news-releases/2008/07/09/berkeley-... & http://www2.lbl.gov/Science-Articles/Archive/sabl/2005/March/06-genopharm.html http://www2.lbl.gov/Science-Articles/Archive/sabl/2005/March... A search company that competed with Google called "seeqpod" was spun out of Berkeley Lab using the tech but was then sued for billions by Steve Jobs https://medium.com/startup-study-group/steve-jobs-made-warner-music-sue-my-startup-9a81c5a21d68#.jw76fu1vo https://medium.com/startup-study-group/steve-jobs-made-warne... and a few media companies http://goo.gl/dzwpFq http://goo.gl/dzwpFq We might combine these approaches as there seems to be something fairly important happening here in this area. Recommendations and sentiment analysis seem to be driving the bottom lines of companies today including Amazon, Google, Nefflix, Apple et al."
- 8y ago