7 ms·
The assertion that this is because "indexing the whole Web is crushingly expensive, and getting more so every day" is a bit flawed. Since old content is very un
by Udo 8y ago
The assertion that this is because "indexing the whole Web is crushingly expensive, and getting more so every day" is a bit flawed. Since old content is very unlikely to be updated, it doesn't have to be re-crawled a lot. I'm certain Google has a score that tells it how often the content of a given site is likely to change. This argument of expense becomes even less durable when you consider that DuckDuckGo, a company with an infinitesimal fraction of Google's resources, is perfectly able to keep that kind of content in its database.
I agree with the observation that this is about shifting everything to current data, because people overwhelmingly care about things that happened a few days ago. There used to be a long tail of users searching for old data and references, but I suspect they're fading away. Biasing the index towards recency also has legal advantages for Google, because delisting old content makes it less likely to receive takedown requests in connection with "right to be forgotten" legislation.
- ghaff 8y agoYeah, it seems a natural consequence of the combination of vast amounts of recent content with the fact that people mostly want recent content. To pick one trivial example from yesterday, if I'm looking for help with an interface issue with some current version of a program, forum posts from 10 years ago are probably not useful. Information that people regularly access for whatever reason will tend to remain relatively visible. But, yeah, relatively obscure older content is just going to get drowned out unless you know exactly where and how to look. One might argue with Google's criteria around relevance. However, that older information is going to get harder and harder to find just in the natural course of things.
- jsnell 8y agoCrawling isn't the real problem, nor is the bulk storage for the crawled pages. What do you do with these pages after you've crawled them? You need to build an index out of them, and serve that index out of some kind of low latency storage (DRAM, Flash). That makes increasing the index size very expensive. The index size has to be limited, and selecting the right pages to include in the index is thus a core quality feature for a search engine.
- Udo 8y agoI'm having trouble imagining that Google would be more limited by the ratio of hardware power vs data size today than it was in the early days. If keeping the whole index in DRAM is now a requirement, then yes, I'd expect a hugely reduced overall dataset - but wouldn't that affect way more sites/pages than the comparatively few dropped historical records? I still suspect that this whole thing is more about bias (and personalization, be it correct or incorrect) in the results.
- jsnell 8y agoIt was a limit back in the day as well. Remember Google's "supplemental results"[0]? This has always happened. The only thing that's different is that a blogger was personally insulted by his output not being fully indexed, and decided to pitch it as history being erased. [0] https://searchengineland.com/google-dumps-the-supplemental-results-label-11830 https://searchengineland.com/google-dumps-the-supplemental-r...
- puzzle 8y agoGoogle's index has been in memory for most of its life now: http://glinden.blogspot.com/2009/02/jeff-dean-keynote-at-wsdm-2009.html http://glinden.blogspot.com/2009/02/jeff-dean-keynote-at-wsd... It's actually more complicated than just a single static index, which is also why it's unrealistic to expect a search engine to be deterministic at scale.
- 0815test 8y agoThe index only has to be on "low latency storage" if low-latency results to any query are required. While that's definitely true of the modal "Google Search", most of these queries for "long-tail, old content" as discussed in the OP don't really need that sort of quick response.
- puzzle 8y agoInterview questions: How would a search engine distinguish between the two kinds of queries, tens of thousands of times a second? And how would one architect such a two-tiered system, particularly with an eye toward cascading failures?