4 ms·
Such a fluffy article. Here is what he's publishing: http://arxiv.org/find/cs/1/au:+Hinton_G/0/1/0/all/0/1 http://arxiv.org/find/cs/1/au:+Hinton_G/0/1/0/all/0/
by Tobu 11y ago
Such a fluffy article.
Here is what he's publishing: http://arxiv.org/find/cs/1/au:+Hinton_G/0/1/0/all/0/1 http://arxiv.org/find/cs/1/au:+Hinton_G/0/1/0/all/0/1
This seems to be a good introduction to the topic: http://arxiv.org/pdf/1310.4546v1.pdf http://arxiv.org/pdf/1310.4546v1.pdf
This is about paragraph vectors: http://arxiv.org/abs/1507.07998 http://arxiv.org/abs/1507.07998 http://arxiv.org/abs/1405.4053 http://arxiv.org/abs/1405.4053
- rspeer 11y agoI have doubts about these results if they depend on paragraph vectors. The first paragraph paper vector (Le and Mikolov 2014) was irreproducible [1]. Not even Mikolov could reproduce it. Paragraph vectors also fundamentally involve training on your test set: the paper stresses the importance of all the vectors being learned jointly, without addressing why this is a problem for evaluation. The developers of gensim have been making an effort to make a version of doc2vec that can be applied to documents it was not trained on (of course it doesn't perform as well). They seem content to clean up after Google's messy publications, but in a fair world, they would be the ones getting the citations if they succeed at this. [1] http://stats.stackexchange.com/questions/123562/has-the-reported-state-of-the-art-performance-of-using-paragraph-vectors-for-sen http://stats.stackexchange.com/questions/123562/has-the-repo...
- varelse 11y agoIs it possible that the reason Quoc Le never published the code is that it got tangled up in the word2vec patent? At GTC 2015, Ren Wu stated (paraphrased) that the reason DNNs are advancing so quickly is because all the practitioners are sharing code, data, and techniques too quickly for the lawyers to gum up the works. The lawyers have since stepped up their game IMO, doing their best to protect us from that imminent Robot Apocalypse Elon Musk keeps warning us is right around the corner.
- rspeer 11y agoThat would be a pretty plausible explanation of why Le has gone completely silent about that paper. I assumed the worst when I read your comment; I figured that neural net semantics was going to stall for 17 years the same way that basic morphology did when Xerox patented FSTs. But it looks like word2vec is Apache-licensed, meaning the patent can only be used defensively. Phew.