3 ms·
I wish you luck and I hope you succeed, but you make it sound much much easier than what it would be. First of all, you're going to drown in hardware costs, if
by acatton 2y ago
I wish you luck and I hope you succeed, but you make it sound much much easier than what it would be.
First of all, you're going to drown in hardware costs, if you run your own hardware. If you run on AWS, you will be the largest AWS customer. 2 years ago, when Google was still displaying result counts, I got 1.3 billions results for "sushi"[1]. This means that if you use a reverse index to lookup your results, the "sushi" entry will be ~19GiB large, assuming you use UUIDs. If you think 90% of this is spam, and you only index the non spam (detecting spam/seo is far from trivial, but let's say you figure it out), you still need ~2GiB just for mapping "sushi". With 755,865 words in the English dictionary, according to wikipedia[2], you'll need ~1.5 PiB (yes, pebi/peta, 1,536 TiB) just to store relationships for English pages. This is assuming you don't support other languages, you discard 90% of pages, and you don't cache the content of pages for re-indexing.
In addition to this, you also need to store the meta-data for each pages (vote counts from your voting system, whether it's serving different content, etc...). The order of magnitude has to be in the O(100TiB) from my conservative gut feeling. (still assuming you discard 90% of the web, and I'll assume you aggregate the metadata on the domain, not on the individual pages)
The second challenge is your ranking. Now that you've become the dominant search engine with your awesome ranking system, you will become the main target for swaths of motivated click-farms which are exploiting workers from low income countries. They will be trying to register accounts, vote and game your ranking. You can most likely detect this behaviour, but their behaviour will be very similar to a significant portion of your real users. So you'll be fishing in a pond with a rocket launcher, and some of your legitimate users will be collateral victims. Otherwise, you'll spend most of your time playing a cat-and-mouse game with the SEO spammers instead of improving your search engine and fixing bugs.
I'm also falling in the trap "i could rewrite that in a weekend" sometimes, but for a search engine, I would love to see decent competition, but it's near impossible.
[1] https://news.ycombinator.com/item?id=30925402 https://news.ycombinator.com/item?id=30925402
[2] https://en.wikipedia.org/w/index.php?title=List_of_dictionaries_by_number_of_words&oldid=1223696888 https://en.wikipedia.org/w/index.php?title=List_of_dictionar...
- adileo 2y agoAbsolutely spot on. Additionally, it's worth mentioning that a lot of content is now locked behind a few major platforms (eg. Facebook, LinkedIn, Medium, YouTube, etc.) or CDNs like Cloudflare, which often block crawling from non-Google IPs or well-known search engines. While the other costs mentioned here can be optimized with current hardware prices and a good database, anti-crawling measures necessitate thousands of IPs/proxies, making the process even more challenging and costly.
- chongli 2y agoAdditionally, it's worth mentioning that a lot of content is now locked behind a few major platforms (eg. Facebook, LinkedIn, Medium, YouTube, etc.) or CDNs like Cloudflare, which often block crawling from non-Google IPs or well-known search engines. I think this is fine. If I want to find something on one of those big sites I just go there directly. However if I want to search the web for a site I’ve never been to before then I’m stuck with the bad results of the current search offerings. It’s quite depressing!
- lelanthran 2y agoI was just idly threatening to do something, not actually starting a venture to do it. But, in any case, lets go with your numbers for running costs: in 2024 money, what is your estimation of the running costs? I ask because there have been a number of new search engines pop up, and they have nowhere near the expenditure power of google, yet they still have devoted followings. > The second challenge is your ranking. Now that you've become the dominant search engine with your awesome ranking system [snipped problems that follow] TBH, that's the best kind of problems to have. Lets become dominant first before we say there's no point in becoming dominant.
- acatton 2y agoMost search engine popping up are just a rebranding of bing and yandex. There is a good article about it: https://seirdy.one/posts/2021/03/10/search-engines-with-own-indexes/ https://seirdy.one/posts/2021/03/10/search-engines-with-own-... Regarding your question of cost, if you buy the cheapest hardware possible, and manage it yourself, because you don't want to pay AWS' premium, for storage (again, assuming to discard 90% of the pages, and only do english) you'll be at least at €80k (12 × €3000 disk servers with 300 × €150 HDDs) of upfront cost, then, if you colocate at Hetzner, just for storage, you'll bet at €500/month + ~€3k of electricity that you need to pay yourself. My gut feeling is that this would cost (only 10% of the English-speaking web) ~€150k of upfront cost and €5k/month to run. Assuming you buy the cheapest everything, and do your own admin sys. And this is not forecasting growth, serving ads, etc...
- immibis 2y agoKagi obviously manages it. The Internet is big, but most of it is spam, which you can discard. You don't need relationships between all pages, just all websites. You should track the quality of the website, not each page. Google result counts are fake.
- acatton 2y agoI don't know how Kagi manages it. I suspect like duckduckgo, they index a little bit on their own, and use a GBY [1] as a back up. According to Seirdy[1], they use Brave in the background. Brave just burn their cryptomoney to build a search engine. Don't get me wrong, I like what their doing, but it was easy for them to start since they bootstraped from Cliqz' index[2] And for the second part, you do need to store the relationship between keywords and pages, that's what I was talking about. You cannot store a relationship between "types of water" and "reddit.com" you need to store it between "type of water" and "reddit.com/r/hydrohomies/..." [1] https://seirdy.one/posts/2021/03/10/search-engines-with-own-indexes/ https://seirdy.one/posts/2021/03/10/search-engines-with-own-... [2] https://brave.com/blog/brave-search/ https://brave.com/blog/brave-search/