5 ms·
This is a great intro to full-text search. These techniques were the basis of some of the early web search engines (pre Google). More of this stuff is explore
by andyjpb 6y ago
This is a great intro to full-text search.
These techniques were the basis of some of the early web search engines (pre Google).
More of this stuff is explored in Witten Moffat and Bell's book "Managing Gigabytes" which dates from the era of those early search engines and when "Gigabytes" was something unimaginably large that certainly wouldn't fit in RAM and perhaps not even on a single disk.
- ramraj07 6y agoWhatever terabytes or petabytes of data we have in indices today, is it all served from memory? I remember Google saying at some point they went fully in memory searches, don't know if they have kept at it or if it's the standard.
- MR4D 6y agoA petabyte of ram is less than $5MM - using just ram costs and consumer ram as well [0]. So given how much money they make and how many servers they have, I think it’s highly probable. The caching of images and web pages might be more difficult. Plus, Google could easily just cache the the results for the top million queries without much ram at all by comparison. [0] - https://www.newegg.com/team-32gb-288-pin-ddr4-sdram/p/N82E16820331396?cm_sp=homepage_ss-_-p4_20-331-396-_-08042020 https://www.newegg.com/team-32gb-288-pin-ddr4-sdram/p/N82E16...
- elorant 6y agoWhat kind of motherboard could take a PB of RAM?
- MR4D 6y ago8,000 of them networked in a data center. Back in 2011 it was thought they had a million servers, so 8,000 is easy for them. [0] - https://www.datacenterknowledge.com/archives/2011/08/01/report-google-uses-about-900000-servers https://www.datacenterknowledge.com/archives/2011/08/01/repo...
- mrkeen 6y agoI'd wager the amount of crawlable information on the web has grown much larger than a Google-like company like has grown its hardware.
- GordonS 6y agoMaybe even Google feels this way, since they seem to have deindexed so much old content?