12 ms·
A simple search engine from scratch
- franczesko 1y agoOn the topic of search engines, I really liked classes by David Evans. The task was also building a simple search engine from scratch. It's really for beginners, as the emphasis is on coding in general, but I've found it to be very approachable. https://www.cs.virginia.edu/~evans/courses/ https://www.cs.virginia.edu/~evans/courses/
- marginalia_nu 1y agoThe SeIRP-book, free online as a PDF, is also a fantastic resource on traditional search engines and information retrieval in general. [1] https://ciir.cs.umass.edu/irbook/ https://ciir.cs.umass.edu/irbook/
- franczesko 1y agoDue to dead links, this is more appropriate url: https://www.cs.virginia.edu/~evans/courses/cs101/ https://www.cs.virginia.edu/~evans/courses/cs101/
- StefanBatory 1y agoServer not found. Did HN gave it hug of death?
- franczesko 1y agoPlease see links to videos and notes - they still work. Udacity must have removed the course
- fuzztester 1y agothe actual course link on udacity gives a 404.
- ktallett 1y agoI always wonder if the days of search engines for specific topics could return. With LLM's providing less than accurate results in some areas, and Google, bing, etc being taken over by adverts or well organised SEO, there feels like a place for accurate, specialised search.
- datadrivenangel 1y agoThe curation of an index of resources is what's needed for niche search
- dcist 1y agoWestLaw and Lexis Nexis provide this for legal search, but quite frankly, these services are subpar. It's amazing that these two companies rake in hundreds of millions but they are both slower than Google, Bing, Yandex, or any LLM service (ChatGPT, Claude, Gemini, etc.) while scouring a universe of text that is orders of magnitude smaller. The user experience is also terrible (you have to login and specify a client each and every time you attempt to use the service and both services log you out after a short -- in my opinion -- period of inactivity, creating friction and needless annoyance to the user). There's an opportunity there.
- ktallett 1y agoI haven't personally used the mentioned services as they aren't in my field, however what is the accuracy of their results? Are they double checked? I don't find LLMs particularly accurate in my field (that's being kind), if anything I find they make up sources that simply don't exist. I mean poor UX has no excuse but slow speed can be reasoned if it makes the quality of the service better.
- ordersofmag 1y agoHere’s a place to start if you want to go down the rabbit hole of how search at places like this is approached. https://haystackconf.com/us2022/talk-12/ https://haystackconf.com/us2022/talk-12/ https://www.youtube.com/watch?v=9vCMFIJRiKk https://www.youtube.com/watch?v=9vCMFIJRiKk
- potato-peeler 1y ago[dead]
- sp0rk 1y agoThe SVG equation is very difficult to read if you're using a dark OS theme because the blog uses the OS preference for dark/light theme (and doesn't seem to give an option to change it manually, either.)
- tekknolagi 1y agoFixed, I think? Let me know
- DylanSp 1y agoWorks now (I noticed the same issue).
- dheera 1y agoOn the side, not criticizing OP but I hate the word "cosine similarity" and I wish people would just call it a "normalized dot product" because anyone who took sophomore-level university calculus would get it, but instead we all invented another word
- cosmicgadget 1y agoThis was a really nice read. Now I have no excuse not to upgrade my blog search. I do feel that I'll have a ton of long tail words like 'prank'.
- snowstormsun 1y agoNice idea, but this approach does not handle out of vocabulary words well which is one major motivation for using a vector-based search. It might not perform significantly better compared to lexical matching like tf-idf or BM25, and being slower because of linear complexity. But cool regardless.
- netdevphoenix 1y agoIt is supposed to be a simple search engine. Keyword: simple. As long as it does what it is meant to, as a simple search engine, it seems fine
- snowstormsun 1y agoUsing tfidf or bm25 would actually be simpler than a vector search. I understand this is just for fun, just wanted to point that out.
- LunaSea 1y agoTF/IDF does not support out-of-vocabulary keywords as far as I know.
- haasisnoah 1y agoHow would you handle those in wordvec? And isn’t a big advantage that synonyms are handled correctly. This implementation still has that advantage.
- cosmicgadget 1y agoOr since OP has both the cosine similarity matching and naive matching, a heuristic combination of the two since they address each other's weaknesses.
- janalsncm 1y agoVector based approaches either don’t handle OOV terms at all or will perform poorly, depending on implementation. If you limit to alphanumeric trigrams for example you can technically cover all terms but badly depending on training data.
- swyx 1y agothis embeds words with word2vec, which is like 10 years old. at least use BERT or sentencetransformers :)
- gthompson512 1y agoI have been thinking a bit lately about how much sense that makes compared to just using word vectors, since traditional queries are super short and often keyword based(like searching for "ground beef" when wanting "ground beef recipes I can cook easily tonight") and so lack most of the context that BERT or similar gives you. I know there are methods like using seperate embeddings for queries and such, but maybe a basic word based search could be more useful, especially with something like fastText for out of vocabulary terms.
- curtisszmania 1y ago[dead]
- kaycebasques 1y ago> The idea behind the search engine is to embed each of my posts into this domain by adding up the embeddings for the words in the post. Ah, OK! I never really grokked how to use word-level embeddings. Makes more sense now.
- skarz 1y agoIs 'grokked' a common verb now? I had never even heard the word until Musk's AI.
- StefanBatory 1y agoIt was a word before, as far as I remember. Saw it a few times here.
- skarz 1y agoWhat does it even mean?
- russfink 1y agoTo understand and comprehend something in fullness. To reach the depths of the concept, idea, or entity so deep that you are practically one with it. (This is per my recollection of the Heinlein story, where grokking one in fullness was the highest form of respect.)
- kaycebasques 1y agoA common verb "now"?? > Grok (/ˈɡrɒk/) is a neologism coined by the American writer Robert A. Heinlein for his 1961 science fiction novel Stranger in a Strange Land. While the Oxford English Dictionary summarizes the meaning of grok as "to understand intuitively or by empathy, to establish rapport with" and "to empathize or communicate sympathetically (with); also, to experience enjoyment",[1] Heinlein's concept is far more nuanced, with critic Istvan Csicsery-Ronay Jr. observing that "the book's major theme can be seen as an extended definition of the term."[2] The concept of grok garnered significant critical scrutiny in the years after the book's initial publication. The term and aspects of the underlying concept have become part of communities such as computer science. https://en.wikipedia.org/wiki/Grok https://en.wikipedia.org/wiki/Grok
- vojtechrichter 1y agoI really like people playing around with technology many take for granted, without understanding its core, underlying princliples
- leumassuehtam 1y agoThe author has a nice series on compiling a Lisp [0], but unfortunately his search engine fails to find it by querying it with "lisp" or "Lisp". [0] https://bernsteinbear.com/blog/compiling-a-lisp-0/ https://bernsteinbear.com/blog/compiling-a-lisp-0/
- tekknolagi 1y agoI wonder if that's just not in the top 10k words :/
- freilanzer 1y agoAnd it's unfinished since 2020.