6 ms·
Companies have been syphoning this data for free for years to train their models. Why would the proprietor of the data not be able to sell it? Assuming of cours
by jerpint 3y ago
Companies have been syphoning this data for free for years to train their models. Why would the proprietor of the data not be able to sell it? Assuming of course their terms of use plainly and clearly state that the data belongs to them
- kobalsky 3y agothings changed, and they want a piece of the cake now. they get the chance to adapt to new stuff that's happening and profit from it, however the only ones left without a say are the users that generated that content. I think it would be fair to say that previously accepted terms and conditions shouldn't ethically count for this, and user data previously generated should be shared for AI training only on an opt in basis.
- zdragnar 3y ago> user data previously generated should be shared for AI training only on an opt in basis. Wouldn't all of the populate models and applications (chatgpt, copiolot, etc) have to be burned to the ground and start clean? My understanding was that pretty much every one of these were trained on "stolen" data, i.e. without prior agreement with the authors / content creators.
- takeda 3y agoI don't think that would be a concern if a proper law would be passed.
- deleted 3y ago[deleted]
- YetAnotherNick 3y agoI think the law was clear before that any data on public internet which doesn't require signing the TOS or login is publicly scrapable. People have been doing it for decades for things like market sentiment analysis and even reselling the data for the same. Why would AI wave suddenly change it.
- AnthonyMouse 3y ago> user data previously generated should be shared for AI training only on an opt in basis. This seems obviously backwards and a boon to enormous corporations. If you want to build a search engine, you have to index everything possible, or it won't be able to find that and then it's useless. Transacting with each individual person in the world would only be possible for megacorps -- it's already expensive enough to index everything if you don't have to do that. It's also perverse that someone could have a right to prevent someone else from providing true information about them. If you're standing there making a false claim, you shouldn't have a right to prevent someone else from proving you wrong just because the evidence is from your own past. But search engines have been ML since even before LLMs, and LLMs are fundamentally the same. How can someone have a default right to deny the public access to facts? Now, maybe there are some things you want to keep private, and then if you share them with someone you want to bind them to not sharing that information with others, like an NDA. But that's opt-out, not opt-in, and you can't really do that for things that are public.
- card_zero 3y agoThe EU disagrees, and have implemented the right to data erasure from search indexes. In the GDPR and applicable to Google since 2016, it says here. https://en.wikipedia.org/wiki/Right_to_be_forgotten https://en.wikipedia.org/wiki/Right_to_be_forgotten
- rpdillon 3y agoThat's opt-out, not opt-in.
- card_zero 3y agoIs it? It's not "like an NDA". You get the data erased (if it is agreed that your privacy is more important than the public interest) after the fact of the data being made public.
- AnthonyMouse 3y ago
- littlestymaar 3y ago> Why would the proprietor of the data not be able to sell it? Assuming of course their terms of use plainly and clearly state that the data belongs to them And assuming we collectively (as “we the people ”) grant them the ability to actually own this data. There's plenty of good reasons to consider that the data doesn't belong to reddit, that it's just public stuff that they happen to host in exchange for user traffic and advertising revenues.
- realusername 3y agoLegally maybe that all works, but ethically the data doesn't belong more to Reddit than Google. They should find a better economic model than selling user data.