4 ms·
I think you might be underestimating the size of the document corpus that you'd be running over.
by humbledrone 8y ago
I think you might be underestimating the size of the document corpus that you'd be running over.
- z3t4 8y agoLets say there are 5 billion url's with an average of 10 KiB of data (if you take out JS/CSS/images etc), and one server has 50 GB or ram you would need 1000 servers, which is very small considered Google probably have one million servers deployed. I just tried to text search your comment on Google and it found your post! So Google is already doing full text search, and does it in less then one second (0.69 to be precise). There are probably many reasons why they don't allow Regex, probably because it would be very easy to "scrape" resources such as e-mail addresses, credit card numbers, etc. It would however be cool if Google would allow you to search structured data, for example find 100 recipes that has eggs in it :P Silly example, but the possibilities are endless!
- mcbits 8y agoMaybe you should clarify what you meant by "naive full text search" because that usually means scanning all documents character-by-character to match the characters in the query, which is definitely not what Google does.
- humbledrone 8y agoBut that's exactly my point -- when you get to the stage where you have 1,000 servers with 50G of RAM each, you have gotten to the point where an optimization like an inverted index is completely sensible. The design you propose has to do a full regex scan over 50 TB of RAM for every. single. user. query. For Pete's sake! This is definitely the realm where the computational costs make it worthwhile to spend engineering resources to optimize, especially if you are going to serve lots of users.