29 ms·
It's incredibly useful for search, given the property that similar words are close in the vectorial space. And given it's purely numbers, it's really fast to co
by _pctq 9y ago
It's incredibly useful for search, given the property that similar words are close in the vectorial space. And given it's purely numbers, it's really fast to compute.
To see an example, type "fuel" in the search input on this page: https://openvoyce.com//products/quuu https://openvoyce.com//products/quuu
You'll see many relevant results, none of them using the word "fuel". This is done purely with postgres, computing a L2 distance sort - no elasticsearch.
- lacksconfidence 9y agoWould you be willing to go into a little more detail about what this is actually doing? What is the shape of the database? Do you normalize each document into a single vector which is compared, or are you keeping per-word vectors? I'm imagining you probably don't have a database with a row for every word, but maybe you do? How do you pre-filter the list of documents to compute L2 for? If no pre-filtering, can this approach scale into millions of documents?
- _pctq 9y agoThe main trick is to do an average of word embeddings in a given document, an idea I took from the youtube paper on recommendation engine [1]. I have a separated service that contains the word embeddings, generated with word2vec. The idea is to generate an embedding for the document by making an average of the embeddings of the words it contains, each having a coefficient based on the word's rarity (so, a rarer word has more weight than a stop word). When saving a document, OpenVoyce is contacting this API and asks to generate an embedding for the document, then it only saves that in its own database (as a "cube" vector of 200 dimensions). From there, searching for something new is just about asking for an embedding for the search terms and using `cube_distance()` [2] as sort function, it does not require pre-filtering since stop words are already weighted off (although, there is some filtering in the API as it ignores words it doesn't know). It would still help to be able to define user specific stop words, though. For example, on Quuu's OpenVoyce, most suggestions are about adding new categories, so "category" should be considered a stop word, that's something I plan to implement. I can't tell yet how it scales to million of records because we're very far from there for now (there are 4500+ suggestions and comments on OpenVoyce at present day). My bet is that if the amount of data becomes a problem, it may be fixed by reducing the number of dimensions of the vectors. Oh, there's also something to know: the cube extension for postgres doesn't allow for more than 100 dimensions. This is something configurable, but only by editing a header file from the extension (that's the author's recommended method). I've detailed the problem and solution on my pg350d repos [3] [1] https://static.googleusercontent.com/media/research.google.com/en//pubs/archive/45530.pdf https://static.googleusercontent.com/media/research.google.c... [2] https://www.postgresql.org/docs/current/static/cube.html https://www.postgresql.org/docs/current/static/cube.html [3] https://github.com/oelmekki/postgres-350d https://github.com/oelmekki/postgres-350d
- PaulHoule 9y agoI think the results for "fuel" are incredibly bad. Sure, "Diesel" shows up, but so does "Essential Oils". Have you done any evaluation to show that this is better than a more a conventional search engine? The Lesson of TREC is that 95% of the things that will "obviously" improve search results will not.
- _pctq 9y agoThe goal is not to not show a single bad result, it's to show the good results.