3 ms·
I’ve been working on a search engine for Gemini space [1] and have been reading a lot of early search engine papers including this one. This Google paper is pre
by billyhoffman 5y ago
I’ve been working on a search engine for Gemini space [1] and have been reading a lot of early search engine papers including this one. This Google paper is pretty good but I prefer the description of Mercator (which became Alta Vista’s crawler)
https://courses.cs.washington.edu/courses/cse454/09sp/papers/mercator.pdf https://courses.cs.washington.edu/courses/cse454/09sp/papers...
There is a updated paper describing Mercators design about a year later when it moved to scale on multiple machines and became a continuous crawler:
https://www.hpl.hp.com/techreports/Compaq-DEC/SRC-RR-173.pdf https://www.hpl.hp.com/techreports/Compaq-DEC/SRC-RR-173.pdf
Also helpful is the entire Modern Informational Retrieval textbook which is available online and dives into indexing, full text search, etc
https://nlp.stanford.edu/IR-book/information-retrieval-book.html https://nlp.stanford.edu/IR-book/information-retrieval-book....
It’s fun because many challenges DEC, Google, and others had in the late 90s can largely be solved by modern machines (disk seek times, having to do appends in bulk, storaging reverse indexes in RAM, etc)
[1] https://en.m.wikipedia.org/wiki/Gemini_(protocol) https://en.m.wikipedia.org/wiki/Gemini_(protocol)
- marginalia_nu 5y agoMy experience is that just winging it goes a surprisingly long way. search.marginalia.nu's broad architecture is extremely similar to that of Google (as described in the OP), and I've arrived at it just by building what made sense without any prior insight into the field. Starting with something dumb and replacing it with something smarter whenever the dumb thing doesn't work does seem to go a long way.