3 ms·
I don’t trust or expect AI companies to serve books at all. That’s not what they are scanning them for.
by dpark 2mo ago
I don’t trust or expect AI companies to serve books at all. That’s not what they are scanning them for.
- swed420 2mo agoNot in a traditional sense, but obviously on the surface, they're using the info to regurgitate in some fashion and serve back. The point is that even under the best intentions, hallucinations occur. Then there's the fact that most models have an ideological bias programmed into them. The only expectation I have is for companies or anybody else to not destroy rare books. Is that such a tall order?
- dpark 2mo ago> they're using the info to regurgitate in some fashion and serve back. Sure, in the same sense that they regurgitate any other text they consume. LLMs by definition do not have the full training dataset available, though. It’s far larger than the resulting model. So they can’t reliably reproduce full text without an external source (or if it’s in the training data repeatedly). ChatGPT actually refused to give me a bible quote the other day, presumably because I ran into some general “book regurgitation” safety net. > The only expectation I have is for companies or anybody else to not destroy rare books. Is that such a tall order? Honestly, yeah. The idea that people or corporations should hold onto books forever because of a cultural “ick” about throwing out books is a bit ridiculous. Most books end up in landfills. They aren’t feeding Da Vinci manuscripts into this pipeline. They are feeding still-in-copyright books.
- swed420 2mo ago> LLMs by definition do not have the full training dataset available, though. That makes it even worse, then. This proves the original point. > The idea that people or corporations should hold onto books forever because of a cultural “ick” about throwing out books is a bit ridiculous. If we're building black and white straw man arguments, then sure, let's not archive anything.
- dpark 2mo ago> That makes it even worse, then. This proves the original point. I don’t know what the “original point” is here, but these AI companies are not providing “book excerpt services” and do not claim to. ChatGPT at least will refuse to provide detailed book excerpts (I hit a week or two ago myself). > If we're building black and white straw man arguments, then sure, let's not archive anything. It seems like you are the one creating the straw man. Do you have evidence that these companies are shredding actually rare books? The only cited concrete examples (in this thread anyway) are all rather boring. I seriously doubt they are shredding 200 year old books because why would they?
- pfdietz 2mo ago> It’s far larger than the resulting model. Is it? How many different books are we talking about, and how much information is that, after conversion to text and lossless compression? Images, maybe, but text?
- dpark 2mo agoThese models are trained on way more than just books. GPT-3 was trained on about half a terabyte of filtered plaintext and the training corpuses have grown significantly by then by all accounts.
- pfdietz 2mo agoI imagine that compresses by ~90%, and current top commercial models have a couple of trillion parameters, don't they?
- dpark 2mo agoThey aren’t trained on compressed plaintext so I’m not sure of the relevance there. But regardless it’s my understanding that’s modern models are trained with orders of magnitude more storage than their parameters require. But it’s possible I’m incorrect. This is getting to the fringe of my knowledge of concrete LLM details.
- pfdietz 2mo agoThe relevance is because the LLMs are storing information, not the explicit text, so we want to know how much actual information they need to store (this being an information theoretic argument). The representation in the parameters doesn't necessarily need 1 parameter per character, if the text is highly redundant.