3 ms·
What do you think the market is for people who want to read books or articles condensed into lines of text suitable for model training? I think it's 0. Maybe
by harshreality 2y ago
What do you think the market is for people who want to read books or articles condensed into lines of text suitable for model training?
I think it's 0. Maybe a handful if pirate libraries of readable ebooks and papers didn't already exist, but even a handful of copyright violators wouldn't be a serious commercial threat to the publishers.
So far, governments have been unable to do anything about pirate libraries. Complaining about AI datasets, poorly formatted for human consumption, seems misplaced.
Criticizing this as "stealing" or "piracy" is a vote for a future where only big tech, or the major publishers themselves, can train models on good datasets. Nobody else will have the money or market power to license that much material.
- zaptheimpaler 2y agoThe market for the entire contents of say OpenAI's entire codebase mangled and corrupted at parts might be 0 as well, doesn't make it not stealing. When NVIDIA's or Samsung's IP was leaked including binary firmware that is hard to read and understood by very few, that doesn't magically mean stealing and releasing it became legal.
- harshreality 2y agoIt doesn't magically mean anything, but you're comparing a case where there's a plausible argument for fair use (no commercial impact from the training set sharing, not done for profit, transformative), against... what specific instance are you talking about with NV or Samsung? Theft of trade secrets? Samsung leaking their own IP to chatgpt, through stupidity, isn't illegal. This doesn't even affect the major AI companies that are doing training, does it? Don't you think Google, OpenAI, Anthropic all have their own datasets by now? If they wanted to use bulk content from pirate ebook and academic paper libraries, they would've already mirrored most of those sites a long time ago.