Y
HN Search
Hacker News Search
new
|
comments
|
top
|
jobs
juxtaposicion
searching PlanetScale…
1.
▲
2.
▲
3.
▲
4.
▲
5.
▲
6.
▲
9 ms
·
61.
▲
by
juxtaposicion
13y ago
Hey SandB0x, thanks for the advice! I actually started off by using numpy.dot, which is precisely what's needed. The problem is that I need it to go even faster (currently takes a few seconds) but this function is already heavily optim
62.
▲
by
juxtaposicion
13y ago
Hey doctoboggan, awesome questions! >> What is the dimensionality of each word vector and what does a words position in this space "mean"? What is this dimensionality determined by? Each dimension roughly is a new way that w
63.
▲
by
juxtaposicion
13y ago
Thanks! The number crunching is indeed very intense. The full body of vectors has a million rows and thousand columns, which fills up all of the 10GB of available memory. When you punch in a query, it adds or subtracts the requested vectors
64.
▲
by
juxtaposicion
13y ago
Some of this is that underlying model is insufficiently trained, but some of this is disambiguation. Disambiguation in text is a very, very, hard problem. So some of your examples, when clarified, are a bit clearer: Paul McCartney - Beatles
65.
▲
by
juxtaposicion
13y ago
It's not perfect; the real limiting factor is the volume of text. For the research paper behind word2vec, to get accurate associations for common words (King, man, woman, etc.) required a news sources training text of order a billion
66.
▲
by
juxtaposicion
13y ago
Hi, creator here (Chris Moody). Great question. The underlying algorithm, word2vec, ( https://code.google.com/p/word2vec/ ) isn't built for streaming data which means that at the moment it assumes a fixed numbe
67.
▲
by
juxtaposicion
13y ago
Harvard - Boston + Silicon http://www.thisplusthat.me/search/Harvard%20-%20Boston%20%2B...