7 ms·
Abstractions my Deep Learning word2vec model made
- platz 11y agoreminds me of how Chinese words are made up of individual characters that have semantic meaning themselves.
- sho_hn 11y agoThe Korean Hangeul alphabet is an interesting compromise. It's an alphabet, but multiple letters are grouped together into syllabic characters when written. Those syllables in many cases map back to Chinese Han characters by way of sound values (a lot of the Korean vocab is Chinese in origin, even though the language have very distinct grammar), which means the boundaries between morphemes often match the character boundaries. This is reflected in the orthography, where in case of multiple options for how to distribute letters over characters, the option that keeps the same morpheme spelled consistently through use in different words is preferred. So while you can write phonetically as in Latin, the written language retains a high level of morphological information and things feel very Lego-like.
- robinsloan 11y ago"Lego-like" == what a great description.
- aswanson 11y agoThere seems to be a pattern in Asian languages that try to equate symbols that represent things to the physical respresentation. I recall a friend stating the Korean word/character for balance looked like a person holding jugs of water on each shoulder.
- sho_hn 11y agoSome Han characters are ideographic in nature, but not all of them. Korean used to be written with Han characters (using sets of very complicated rules for how to apply them to the language) prior to the invention of Hangeul, but other than a handful of them they aren't in widespread use any more outside specialized or educational contexts. However, some of the Hangeul letters are featural in design, e.g. the velar consonant ㄱ (g/k) is meant to be a side view of the tongue when producing its sound.
- aswanson 11y agoVery interesting. Can you refer me to some tutorials or texts on these features of those languages?
- lqdc13 11y agoI thought Word 2 Vec isn't "Deep Learning" as both CBOW and skip-gram are "shallow" neural models.
- dave_sullivan 11y agoHow about a single layer neural net trained with dropout? Not deep learning because there's only 1 layer, but the technique is fairly new and used in deep learning, popularized by some of the guys that popularized deep learning, usually mentioned in conversations about deep learning. It's really a shallow neural model though. Word2vec is similarly related, but you're right in that it is not a perceptron-based neural network with multiple layers (ie "deep neural network") Still, interesting blog post, worth reading and googling for more information. FWIW I just use the deep learning definition of "neural net or representation learning research since 2006" and find it fits better.
- deleted 11y ago[deleted]
- fchollet 11y ago> the technique is fairly new and used in deep learning, popularized by some of the guys that popularized deep learning, usually mentioned in conversations about deep learning Logistic regression with regularization is fairly new? 'Pioneered' by the same people as deep convolutional neural networks? Are you certain about this?
- agibsonccc 11y agoI think you need to define regularization. L1/L2? Yes those are old. Drop out itself IS fairly new[1]. As well as its newer cousin drop connect. I agree neural nets themselves are basically just a crazier parametric model. Many of the things we do to modify the gradient are applicable to logistic and other simpler regression techniques as well. [2]: https://cs.nyu.edu/~wanli/dropc/dropc.pdfhttp://www.cs.toronto.edu/~rsalakhu/papers/srivastava14a.pdf https://cs.nyu.edu/~wanli/dropc/dropc.pdfhttp://www.cs.toron...
- SilasX 11y agoSo that's the result? That you can find sorta clever vector equations in the results like "Obama + Russia - USA = Putin"?
- fiatmoney 11y agoYes. There is some nifty work on trying to structure an actual grammar based on the geometry of the embedding space, but "look at this cool thing we found" is a totally worthwhile thing to publish. (Embeddings themselves are useful for a whole lot more than that, though.)
- SilasX 11y agoIf these are the most interesting results, then it's not. Most of the work seems to be in the (human provided) search for happy combinations like these. Are they typical? Is the whole space like this? Is there a consistent metric for deciding that these are in fact clever and it's not just the human doing the work in making it make sense in a funny way?
- sp332 11y agoAs pointed out elsewhere, the results are pretty accurate, not nearly perfect but enough to open up new applications. http://arxiv.org/pdf/1301.3781.pdf http://arxiv.org/pdf/1301.3781.pdf Word2vec is exciting because it has a useable amount of accuracy for a relatively small amount of cleverness on the part of the human. You can just throw a huge corpus at it and it performs reasonably well.
- SilasX 11y agoWhat does "accurate" mean in this context? If there's a standard meaning of it, I'd be glad to judge it by that. But the metric the author offered, as their defining results, was some cute things with Putin and Obama. So yeah, if the results are accurate relative to some standard metric, I wasn't disputing that. I was questioning the reliability of the results the author presented, about cute Obama/Putin connections.
- bdamos 11y agoHow did you select words to compare? Did you have to try many poor combinations before selecting a "good" set?
- MrLeap 11y agoI've played with word2vec, and it took nothing to get really interesting combinations. The first thing I tried was computer : server :: phone : ? I didn't really have a great answer for that in my head before I ran it. Word2vec decided the closest match was "voicemail". It breaks down when you feed it total nonsense, but what would you expect it to do? :P I'm constantly super impressed by properties of the vectors.
- danieldk 11y agoMikolov, et al. 2013 [1] do a proper evaluation of this. E.g. they found that the skip-ngram model has a 50.0% accuracy for semantic analogy queries and 55.9% accuracy for syntactic queries. word2vec comes with a data set that you can use to evaluate language models. [1] http://arxiv.org/pdf/1301.3781.pdf http://arxiv.org/pdf/1301.3781.pdf
- rspeer 11y agoI would insist on a better dataset before really calling these "semantic analogies" (and don't just take my word for it: Chris Manning complained about exactly this in his recent NAACL talk). The only semantics that it tests are "can you flip a gendered word to the other gender", which is so embedded in language that it's nearly syntax; and "can you remember factoids from Wikipedia infoboxes", a problem that you could solve exactly using DBPedia. Every single semantic analogy in the dataset is one of those two types. The syntactic analogies are quite solid, though.
- danieldk 11y agoand "can you remember factoids from Wikipedia infoboxes", That's a simplification. E.g. I have trained vectors on Wikipedia dumps without infoboxes, and I queries such as Berlin - Deutschland + Frankreich work fine. Of course, even the remainder of Wikipedia is nice text in that it will contain sentences such as 'Berlin is the capital of Germany'. So, indeed, it makes doing typical factoid analogies easier. That said -- I am more interested in the syntactic properties :).
- lelf 11y ago> word2vec is a Deep Learning technique first described by Tomas Mikolov only 2 years ago but due to its simplicity of algorithm and yet surprising robustness of the results, it has been widely implemented and adopted. … And patented http://www.freepatentsonline.com/9037464.html http://www.freepatentsonline.com/9037464.html
- andrewtbham 11y agoIf you use word2vec in an agent, like siri or watson, how would google know?
- msoad 11y agoThat's not a question you can ask in a meeting at Apple!
- agibsonccc 11y agoAs someone who has a card at the table with this one, we asked that question on the list fwiw[1]. I don't think there will be any problems here. 1: https://groups.google.com/forum/#!searchin/word2vec-toolkit/patent/word2vec-toolkit/1hID9F74_Ho/WFlQNrRfulgJ https://groups.google.com/forum/#!searchin/word2vec-toolkit/...
- fchollet 11y agoThe "deep" in deep learning refers to hierarchical layers of representations (to note: you can do "deep learning" without neural networks). Word embeddings using skipgram or CBOW are a shallow method (single-layer representation). Remarkably, in order to stay interpretable, word embeddings have to be shallow. If you distributed the predictive task (eg. skip-gram) over several layers, the resulting geometric spaces would be much less interpretable. So: this is not deep learning, and this not being deep learning is in fact the core feature.
- ma2rten 11y agoI think you could see it as a sort of recurrent neural network.
- fchollet 11y agoHow? There is no recurrent data flow in word2vec. Word2vec maps words and their context with 2 embeddings, a dot product and a sigmoid. That's it.
- rndn 11y agoCouldn't you treat the network as a matrix and perform addition and subtraction on it in the same manner?
- Smerity 11y agoI'm uncertain whether you mean the algorithm or the output. The question is interesting to both however. The most common method for producing word vectors, skip-grams with negative sampling, has been shown to be equivalent to implicitly factorizing a word-context matrix[1]. A related algorithm, GloVe, only uses a word-word co-occurrence matrix to achieve a similar result[2]. You can also view the output as an embedding in a high dimensional space (hence the name word vectors) but more surprisingly you can learn a linear mapping between vector spaces of two languages, which lends it immediately useful in translation. From [3]: "Despite its simplicity, our method is surprisingly effective: we can achieve almost 90% precision@5 for translation of words between English and Spanish". [1]: "Neural Word Embedding as Implicit Matrix Factorization" http://papers.nips.cc/paper/5477-neural-word-embedding-as-implicit-matrix-factorization.pdf http://papers.nips.cc/paper/5477-neural-word-embedding-as-im... [2]: http://nlp.stanford.edu/projects/glove/ http://nlp.stanford.edu/projects/glove/ [3]: "Exploiting Similarities among Languages for Machine Translation" - page 2 has an intuitive 2D graphical representation http://arxiv.org/pdf/1309.4168.pdf http://arxiv.org/pdf/1309.4168.pdf
- tshadwell 11y agoWhen I see things like this, it makes me wonder how much data forms each of these vectors; if a single article were to say things about Obama, or humans and animals, would it produce these results?
- thisjepisje 11y agoAnyone tried this with the corpus of HN commentary?
- fchollet 11y agoI have, actually. Here's the code for the experiment, with a link to download the data: https://github.com/fchollet/keras/blob/master/examples/skipgram_word_embeddings.py https://github.com/fchollet/keras/blob/master/examples/skipg... I also recommend using Gensim for word embeddings.
- datacog 11y agoOP: Do you have some more results to share coming from your model?
- fauigerzigerk 11y agoI wonder if Obama + 2017 == Obama - President
- eatonphil 11y agoI am not getting the "Obama + Russia - USA = Putin" piece nor the "King + Woman - Man" bit either. Nothing particularly meaningful came up on a search for the latter. Could someone explain?
- mwsherman 11y agoIf I understand your question, the idea is that one can do “arithmetic” on concepts. Essentially the first equation asks “Obama : USA :: ? : Russia”. Similarly, “King : Man :: ? : Woman”. The way the corpus “talks about” Obama in relation to the USA is similar to how the corpus talks about Putin in relation to Russia. That the system can reveal this is amazing to me.
- eatonphil 11y agoOh cool, I see! Thanks for the explanation.