3 ms·
Interview questions: How would a search engine distinguish between the two kinds of queries, tens of thousands of times a second? And how would one architect
by puzzle 8y ago
Interview questions:
How would a search engine distinguish between the two kinds of queries, tens of thousands of times a second?
And how would one architect such a two-tiered system, particularly with an eye toward cascading failures?
- bshipp 8y agoI don't work at Google so I'm probably way off base, but if I was designing it I wouldn't bother telling the difference between the two types of queries. I'd break up the indices into digestible chunks, perhaps chronologically by year/month crawled, and then run all queries simultaneously (in parallel) against all those index chunks and combine the results at the end. Infinitely scalable and can be tweaked to ensure specific response times. And there'd definitely be no need to set some arbitrary date cut-off; just add a few more virtual machines. I'd bet that's what Google was doing, and then scaled back those machines to save money and boost profits.
- puzzle 8y agoThat's kind of how Google works, with multiple index tiers. Look up patents by Anna Paterson to get a few clues, assuming your lawyers won't bark at you. Still, you can't keep partial results around forever, unless you want to make searches a lot more expensive, having to add a lot of capacity just to deal with the buffer bloat. Each query touches at least a thousand machines. Adding "a few more virtual machines" isn't going to cut it, especially if you have to handle tens of thousands of requests per second.
- typon 8y agoMake it opt-in. Instead of requiring every search to be finished in less than 0.5 seconds, allow users to tick a box that says "Take your time" and pull indices from slow storage in that case. If I know I want something niche, I am willing to wait the extra few seconds or even a minute.
- aeorgnoieang 8y agoHell, even waiting an entire day (e.g with results sent via email) might be reasonable for some searches.
- puzzle 8y agoBehind the scenes, a search on Google involves at least a thousand machines. How much extra state (internal connections, memory for partial results, etc.) would such a new search type create? How do you deal with a new kind of hot spots now? What if millions of people suddenly activate such an option? What if a botnet does it?
- typon 8y agoAll good points :)
- ummonk 8y agoProgressive loading is a thing.