3 ms·
It kind of depends how in depth you want to go. I wrote a small but somewhat complete search engine some time ago. The steps are basically: Have a queue with
by LordHeini 4y ago
It kind of depends how in depth you want to go.
I wrote a small but somewhat complete search engine some time ago.
The steps are basically:
Have a queue with urls.
Download a page from the queue
Use a html parser to remove markup and get the links of the page.
Add the links to a queue without the already visited ones.
Use a stemmer to clean up the text (porter stemmer or whatever).
Calculate an inverted index: https://en.wikipedia.org/wiki/Inverted_index https://en.wikipedia.org/wiki/Inverted_index
Save the stuff in an appropriate data structure (Hash or Tree).
Write a query engine for AND, OR and (all the words).
Calculate a simple page rank by counting the links to a page.
For just learning purposes it is not that hard but if you want to get all the crazy corner cases of the "real" web you will go insane.
The easy alternative:
Use a lib to crawl a page.
Plonk all the documents text into postgres with full text search or elastic search.