4 ms·
This is a good point and I'm surprised it hasn't had more attention. We've seen history repeat itself enough times that it's inevitable. Probably the best you c
by version_five 3y ago
This is a good point and I'm surprised it hasn't had more attention. We've seen history repeat itself enough times that it's inevitable. Probably the best you can do is download a few of the full precision 70B or bigger models (LLaMA etc) to at least have some state of the art weights around you could use yourself if you wanted if/when openAI goes to shit.
The nice thing about the language models is as least they're relatively compact, compared to google search which you'd never be able to host locally.
Though thinking about it, in a way LLMs are local copies of the "good" internet. The big ones mostly seem to know what information there is to know and have none of the crap and link/content farms that is now most of the internet.
- MPSimmons 3y ago>... compared to google search which you'd never be able to host locally That's something I haven't thought about for a long time. How big is the Google index, would you say? Petabytes or exabytes, or am I vastly overestimating?
- raincole 3y ago> over 100,000,000 gigabytes Accodrding to Google: https://www.google.com/search/howsearchworks/how-search-works/organizing-information https://www.google.com/search/howsearchworks/how-search-work...
- hansvm 3y agoProbably petabytes for the text index and exabytes for the whole index (including images and video). The entire textual internet is hundreds of terabytes uncompressed, and indexes tend to balloon storage size to make serving tailored queries faster and cheaper. That idea shifts a bit with images/video where vector/ML techniques dominate; you get a reduction in quality, but quality is good enough (mostly, probably, we think), and you have within a small constant multiple one way or the other of the underlying dataset size for the index size (as opposed to text, where 100x-1000x one way or another isn't uncommon). If the index is less than petabytes then it's impossible to serve many long tail queries even just based on text content. IME those queries _have_ gotten progressively worse results the last few years, so maybe the techniques have changed, but when the money printer is "search" and a petabyte of data is a fraction of a full-time engineer in cost to store, I doubt they would have cut costs that aggressively.
- jdshaffer 3y agoThat reminds me of the time my friend downloaded the ENTIRE Yahoo! search database to a 1.44MB floppy. I think that was back in 1993. laugh