4 ms·
From my own attempt at a small personal search engine of the web, which doesn't yet use vector embeddings, each page averages to 10kB of indexed storage. That
by SyneRyder 2mo ago
From my own attempt at a small personal search engine of the web, which doesn't yet use vector embeddings, each page averages to 10kB of indexed storage.
That would make the final index size of 4 Billion pages about 40 Terabytes. Those charts seem to suggest that's just the size of the Bing index though, and that Google is actually 10x larger at 40 Billion pages. So that would be 400 Terabytes.
My little engine doesn't index the full HTML. If I did, each page averages to 170KB in size, and your index storage just grew 17x.
On a tiny scale - single digit millions - you can get ridiculously far with just SQLite.
- vivzkestrel 2mo ago- thank you for sharing that - another stupid question: how do you about loading data from millions of pages simultaneously - here is my silly thought process for this: - get multiprocessing library in python - combine that with asyncio and aiohttp - send a whole bunch of requests and save raw html to file storage? - some big questions... - how often do you scan the same website - what headers do you need to add in order to make it not look like some bot or should you actually reveal that you are a search engine bot - do you need rotating proxies? is something like brightdata or residential proxies used or am I overthinking this? - I am thinking of taking a small subset of 400 billion pages (like maybe just every blog ever listed on HN) and vector embedding all the text - what do you think the cloud infra side specifically on AWS (since I am highly familiar with it) would look like?
- SyneRyder 2mo agoSo, I clearly need to write a blog post about this one day - there are so many answers and they won't all fit in an HN reply. So I'll give some higher level concepts that I found helpful. Not to dodge your questions, but it gives a way to think about things. * Do Things That Don't Scale. There's a Paul Graham essay on this, but the general idea I took is that a naive, brute force solution will get you quite a long way, and you'll learn lots from it. PHP & SQLite on shared web hosting should never work, but, it goes SO much further than you'd think. * Wait Until Something Is A Problem. You'll eventually feel the parts that need scaling. Worry about them then. For a long time I didn't even have an inverted index, didn't even have FTS enabled on SQLite. Things still worked. (But you really do want an inverted index & full text search.) * Do You Actually Need This? You probably don't need to index the entire internet. You only want the parts of the internet that are relevant to you. You can use your browser's history to get a list of the websites / pages that are actually important to you. It is not that many. This is roughly the idea behind asciimoo's Hister project, which makes a local index of the internet while you browse the web. On headers & proxies & javascript etc... so far, I don't bother. If a website doesn't want me indexing them, I respect that. There's still large parts of the web that do want to be indexed. Read their robots.txt files so you can learn what websites would like you to do, especially Crawl-Delay. I have never used AWS, so I can't help there. My stuff is all just Apache, PHP & SQLite, and a bunch of self-made Go command-line programs. For further research, read all of the Marginalia Search blog, you'll learn lots from that. Remember Marginalia is run by just one person. Also look at Seirdy's list of Search Engines With Their Own Indexes. https://www.marginalia.nu/log/ https://www.marginalia.nu/log/ https://seirdy.one/posts/2021/03/10/search-engines-with-own-indexes/ https://seirdy.one/posts/2021/03/10/search-engines-with-own-...
- vivzkestrel 2mo agoyou sir are a legend! thank you very much for sharing those
- TrueProxies 1mo ago[flagged]