4 ms·
I've worked to Inktomi/Yahoo search ~2000-~2015. Any query that is accurate across the whole index is going to be very very expensive. In order to return fast
by lakis 5y ago
I've worked to Inktomi/Yahoo search ~2000-~2015.
Any query that is accurate across the whole index is going to be very very expensive.
In order to return fast results, there are caches and precomputed results in almost every level of a web search query.
But an accurate count implies that you will bypass all the caches and count every single document for every single query term. That's very very expensive in both CPU and memory.
We had special internal keywords that disabled caches and produced accurate results. All of them came with special warnings that querying too fast with these special options could bring down the whole search engine. A constant worry was the new employee that knew very little, learnt these keywords and then ops were dealing with crashes across ten of thousands of machines due to out of memory.
As to the fabrication. The number are not random numbers. They are estimates based on what we think is the best approximation of the real numbers.
We had done internal research to compute the best approximation function.
Also, it's very difficult to figure out what is real and what is fake from the outside.
I was involved in
https://www.nytimes.com/2005/08/10/business/worldbusiness/yahoo-says-its-search-index-is-bigger-than-googles.html https://www.nytimes.com/2005/08/10/business/worldbusiness/ya...
Yahoo was 100% correct on their claims. I run the queries myself. But external researchers could not verify the claims because all external queries were hitting caches. So they were trying unique queries in order to estimate the index size which end up having a lot of problems.