4 ms·
see, that's the problem with engineers today. AltaVista ran on 3x300MHz 64bit processors with 512M of RAM back in 1995. Resources cost peanuts these days. It's
by arpa 5y ago
see, that's the problem with engineers today. AltaVista ran on 3x300MHz 64bit processors with 512M of RAM back in 1995. Resources cost peanuts these days. It's just that we're so used to bloat and digital inflation, we can't even start considering unbloated implementations as we perceive them as "not modern". Apparently we are also stuck in the centralized/ownership mentality. If you crowdsource search indexing and processing, it scales alongwith amount of users. Oh and also another bane of the internet is the need to monetize. FFS, if email would be invented today, we probably would have to buy NFT poststamps.
- adtac 5y agoI'm sorry what? Have you seen the rate at which data is being created today? I mean, if you want to index the size of the 1995 web with your raspberry pi, go ahead, but it costs insane amounts of money to index the December 2021 web and keep the index up-to-date. edit / full disclosure: I work at google but nothing related to search
- deleted 5y ago[deleted]
- Closi 5y agoI think the words "minimally viable" are being ignored here. I think OP's point is, assume you only have 10/100 terabytes of space and limited compute ability - how would you approach the problem? I assume 90% of google's searches probably come from less than 1% of their total index, not to mention that Google is also keeping full cached versions of the whole website including images.
- adtac 5y agoI just eyeballed my browser history from the last 2-3 days and I'd estimate 15% is current/latest news related, some 25% is programming related, 15% is e-commerce stuff, the rest random crap. I'd imagine 10-100 TB can easily serve all of _my_ search space (even the links I didn't look at on page 10) from the past few years, but that's the thing -- it's just my search space. How do you serve the rest of the world? I wish I knew the answer :)
- foxfluff 5y agoWell Google doesn't know the answer either. Results are complete trash when you try find something niche or in a local language. The challenge isn't to index the entire web, it's to index the useful parts of it, and I think an index covering most of the useful web can be seeded quite easily with some community effort.
- deleted 5y ago[deleted]
- noogle 5y agoIs it really necessary to index EVERYTHING? It's true that we have much more data today than 26 years ago, but not all of these websites qualify or provide value (duplicate results, promotional content, outdated content). The challenge then moves to the curation, but it's no longer infeasible.