9 ms·
Isn't it basically not possible for the input data set list to be listed? It's an open secret all these labs are using immense amounts of copyrighted material.
by pradn 1y ago
Isn't it basically not possible for the input data set list to be listed? It's an open secret all these labs are using immense amounts of copyrighted material.
There's a few efforts at full open data / open weight / open code models, but none of them have gotten to leading-edge performance.
- bee_rider 1y ago“Not possible” = “a business-destroying level of honesty”?
- tokioyoyo 1y agoThere is a "keep doing what you're doing, as we would want one of our companies to be on top of the AI race" signal from the governments. It could've been stopped, maybe, 5 years ago. But now we're way past it, so nobody cares about these sort of arguments.
- rcxdude 1y agoEven if training on the copyrighted material is OK, just providing a data dump of it almost certainly is not.
- alpaca128 1y agoNo need for a data dump, just list all URLs or whatever else of their training data sources. Afaik that's how the LAION training dataset was published.
- anonymoushn 1y agoproviding a large list of bitrotted URLs and titles of books which the user should OCR themselves before attempting to reproduce the model doesn't seem very useful.
- echoangle 1y agoAren't the datasets mostly shared in torrents? They probably won't bitrot for some time.
- Wowfunhappy 1y ago...no? They also use web crawlers.
- deleted 1y ago[deleted]
- bee_rider 1y agoThe datasets are collected using web crawlers, but that doesn’t tell us anything about how they are stored and re-distributed, right?
- Wowfunhappy 1y agoWhy would you store the data after training?
- bee_rider 1y agoAre you saying that you know they don’t store the data after training? I’d just assume they did because—why scrape again if you want to train a new model? But if you know otherwise, I’m not tied to this idea.
- Wowfunhappy 1y agoI'm also assuming. But I would ask the opposite question: why store all that data if you'll have to scrape again anyway? You will have to scrape again because you want the next AI to get trained on updated data. And, even at the scale needed to train an LLM, storing all of the text on the entire known internet is a very non-trivial task!
- prmoustache 1y agoThat doesn't mean it isn't possible.
- 3abiton 1y agoThe only way this would work is with "leaks". But even then as we saw with everything on the internet, it just added another guardrail on content. Now I can't watch youtube videos without logging in, and nearly every website I need to solve some weird ash captchas. It's becoming easier to interact with this chatbots rather than search for a solution online. And I wonder with Veo 4 copy cats, it might be even easier to prompt for a video rather than search for one.
- ratamacue 1y agoMy brain was largely trained using immense amounts of copyrighted material as well. Some of it I can even regurgitate almost exactly. I could list the names of many of the copyrighted works I have read/watched/listened to. I suppose my brain isn't open source, although I don't think it would currently be illegal to take a snapshot of my brain and publish it if the technology existed and open-source that. Granted, this would only be "reproducible" from source if you define the "source" as "my brain" rather than all of the material I consumed to make that snapshot.
- overfeed 1y ago> Some of it I can even regurgitate almost exactly If you (or any human) violate copyright law, legal redress can be sought. The amount of damage you can do is limited because there's only one of you vs the marginal cost of duplicating AI instances. There are many other differences between humans and AI in terms of capabilities and motivations to f the legal persons making decisions.
- CamperBob2 1y agoThe amount of damage you can do is limited because there's only one of you vs the marginal cost of duplicating AI instances But enough about whether it should be legal to own a Xerox machine. It's what you do with the machine that matters.
- overfeed 1y ago> It's what you do with the machine that matters. The capabilities of a machine matter a lot under law. See current US gun legislation[1], or laws banning export of dual-use technology for examples of laws that have inherent capabilities - not just the use of the thing- as core considerations. 1. It's illegal to possess a new, automatic weapon with some grandfathering prior to 1986
- ben_w 1y agoWhile true, computers in general alreay had the ability to perfectly replicate data, hence blank media tax: https://en.wikipedia.org/wiki/Private_copying_levy https://en.wikipedia.org/wiki/Private_copying_levy I think the reason for all the current confusion is that we previously had two very distince groups of "mind" and "mindless"*, and that led to a lot of freedom for everyone to learn a completly different separation hyperplane between the categories, and AI is now far enough into the middle that for some of us it's on one side and for others of us it's on the other. * and various other pairs that are no longer synonyms but they used to be; so also "person" vs. "thing", though currently only very few actually think of AI as person-like