7 ms·
Stack Overflow simply bans folks who don't want their advice used to train AI
- amelius 2y agoAnd let me guess, meanwhile they ask a big fee from AI companies for using the data.
- smilespray 2y agoNo need to guess — that's what the protests are about.
- lxgr 2y agoCan they even do that? I know that SO content is licensed under creative commons; is there an additional dual license that allows them to commercially use it under different (e.g. non-attribution) terms?
- lolinder 2y agoDiscussed the other day in a consolidated thread from a few duplicate submissions (311 comments total): https://news.ycombinator.com/item?id=40302792 https://news.ycombinator.com/item?id=40302792
- fer 2y agoContent contributed on StackOverflow (and all the StackExchange sites I believe) is CC, so it's really a losing battle.
- schlauerfox 2y agoHow is the use of CC-BY-4.0 by a LLM going to be attributed to satisfy the license conditions is my question. Every answer by the model contains the work as part of it's token weights and can regurgitate.
- trueismywork 2y agoYup, you need to cite every single poster ever.
- manuelabeledo 2y agoWouldn't a compliant solution be just a glorified search engine?
- jessriedel 2y agoCopyright covers the words not the ideas/concepts/facts. If I read a StackExchange post, I can absolutely use those ideas/concepts/facts to write a new document, which I can choose to copyright and sell, and there's no need for me to cite the SE post. Now, as a matter of academic courtesy it's nice to cite the SE post if it's feasible, but it's impossible for me to remember where I learned everything I've learned over the years. Likewise, an LLM shouldn't be forced to cite where it learned facts, which might be technically hard/impossible, but if there is a clear best source then as a matter of politeness it should. And LLMs mostly do this when they can!
- manuelabeledo 2y ago> Now, as a matter of academic courtesy it's nice to cite the SE post if it's feasible, but it's impossible for me to remember where I learned everything I've learned over the years. Likewise, an LLM shouldn't be forced to cite where it learned facts, which might be technically hard/impossible, but if there is a clear best source then as a matter of politeness it should. And LLMs mostly do this when they can! Scientific papers have included references for as long as I can remember. It's not about courtesy, but precision and acknowledgement. So, are we letting LLMs run free and do things that would be frowned upon if they were done by humans?
- jessriedel 2y agoWhen LLMs are used to write academic papers, then there is definitely a stronger requirement on the authors deploying them for identifying the source of ideas, as a matter of professional ethics (not law). But this isn't the issue being discussed, and it's not anything new with LLMs. Long before LLMs, academic authors occasionally got ideas from other places StackExchange without citing them properly in papers. (They frequently learn things from StackExchange, but most of the time they don't need to cite them because it's part of common knowledge. Cases where they got novel, cite-necessary insights is much rarer, as it will be for LLMs.)
- trueismywork 2y agoIts CC-BY-SA so openAI has to release the weights obtained by training.
- sethhochberg 2y ago> Stack Overflow has been banning users wholesale who have attempted to delete or deface their own posts on the site Its not just people who are upset getting banned for being upset, its people who are attempting to burn it all down on their way out the door in protest. I get where the protests are coming from, but the cardinal rule of online communities is "once you post it, its out there". Other people have reacted to it, replied to it, quoted it. You break not just your own content but entire discussions if you mass-delete your contributions. They stopped being exclusively yours to take back once you contributed them to a broader conversation.
- 7e 2y agoA right to one's own expression/speech is a basic human right.
- nathan_compton 2y agoIs this true? I don't disagree with the idea, but do you really have the _right_ to delete some previous thing you've said no matter where it is? Seems untenable.
- lupusreal 2y agoThe right to delete your own words after uttering them (and giving some company a license to reproduce them no less) is dubious. I guess it exists in the EU with their "right to be forgotten" and GDPR stuff, but I think not in most of the rest of the world.
- itsdrewmiller 2y agoThere is no basic human right of “takesies backsies” though - that seems more like the relevant question, along with IP ownership.
- lxgr 2y agoDeleting your past speech is not a basic human right, for good reasons: Imagine e.g. the effect on democracy of politicians being able to have all of their past utterances deleted from the public record because they don't feel like it represents their current branding anymore. For private individuals not "of public interest", the situation is a bit different, but e.g. the EU's "right to be forgotten" is arguably not a basic human right, but a relatively new and narrow concept for which there isn't any type of international consensus yet.
- lolc 2y agoIt's not clear to me what the deal is supposed to be about. Isn't it expected that Stackoverflow questions and answers can be used by anybody including to train models? If Stackoverflow is trying to make exclusive deals with Openai, that is against the collaborative spirit of the platform, and I will stop contributing. After all, Openai is charging people for service. If Openai are the only ones given access, Stackoverflow becomes a gatekeeper, peddling my contributions. It'll beget a fork.
- mort96 2y agoIt's shitty to take content under an attribution license like the CC-BY-SA license SO uses and pour it into a soup of linear algebra where the content remains but all attribution is lost.
- coldpie 2y ago> Isn't it expected that Stackoverflow questions and answers can be used by anybody including to train models? I don't think so. Content contributed to SO gets licensed under CC-BY-SA, no? How do these AI models respect the "BY" portion of that license? How do we enforce the "SA" portion of the license on people who use the output of the model? The answer is they don't, but the big companies decided violating copyright is OK so long as they do it all at once, because there's a lot of money to be made if you ignore copyright. Personally I'd be OK with a compromise where any entity that creates or uses an AI trained on copyrighted data forfeits copyright on all of their own works (not just those created by AI, all of their works).
- throwaway290 2y ago> Isn't it expected that Stackoverflow questions and answers can be used by anybody including to train models? CC ShareAlike requires attribution + derivative works to be shared under the same license. ClosedAI is absolutely in the wrong here.
- add-sub-mul-div 2y ago> It's not clear to me what the deal is supposed to be about. AI is unpopular and makes this an unpopular move, enough so to cause this drama. It's really not any more complicated than that.
- jessriedel 2y agoI don't get it. Like, it's one thing if a fact-collecting business like the NYTimes doesn't want its stories to train an MML. I think under current law they don't have much of a case, because facts aren't copyrightable, but there's a reasonable argument that the law should be updated somehow in light of technological change. But the work produced by all StackExchange users is explicitly released under a CC BY-SA license. The whole point is to collect and publish facts/ideas/understanding for anyone to see and use for any purpose, including running a business. Yes, the "SA" (share alike) part means if you want to use and modify the words then you need to release them under license that is at least as permissive, but LLMs aren't using the words; they are clearly digesting the facts and expressing them in their own words. And, unlike the NYTimes, there is no issue of "couldn't new tech undermine society's current method of economically incentivizing fact-collection?". The StackExchange users are not being paid, and the fact that the license is not NC (non-commercial) explicitly means that using their hard work to make money is allowed (and encouraged!).
- trueismywork 2y agoIt's not clear whether LLM actions are derivative works or not. The only way they would not he derivative works is if there's specific logic in the implementation to check for all copyrighted works used in training and prevent an infringing output from such a work.
- jessriedel 2y agoThere doesn't have to be specific logic.
- progval 2y ago> Yes, the "SA" (share alike) part means if you want to use and modify the words then you need to release them under license that is at least as permissive, but LLMs aren't using the words; they are clearly digesting the facts and expressing them in their own words The CC BY-SA text <https://creativecommons.org/licenses/by-sa/4.0/legalcode.en https://creativecommons.org/licenses/by-sa/4.0/legalcode.en> says nothing about the words. The words "word" and "words" do not even appear in the legal text. What it says, however, is: > In addition to the conditions in Section 3(a) , if You Share Adapted Material You produce, the following conditions also apply. > 1. The Adapter’s License You apply must be a Creative Commons license with the same License Elements, this version or later, or a BY-SA Compatible License. > 2. You must include the text of, or the URI or hyperlink to, the Adapter's License You apply. You may satisfy this condition in any reasonable manner based on the medium, means, and context in which You Share Adapted Material. where "Adapted Material" is defined as: > material subject to Copyright and Similar Rights that is derived from or based upon the Licensed Material and in which the Licensed Material is translated, altered, arranged, transformed, or otherwise modified in a manner requiring permission under the Copyright and Similar Rights held by the Licensor. For purposes of this Public License, where the Licensed Material is a musical work, performance, or sound recording, Adapted Material is always produced where the Licensed Material is synched in timed relation with a moving image. So if you want to argue that LLMs don't have to follow the SA clause, then you have to argue either that: 1. LLMs aren't an arrangement or transformation of their input, or 2. the license doesn't apply at all, eg. because it's fair use in whatever jurisdiction they fall under. OpenAI is using argument 2 <https://news.ycombinator.com/item?id=37780199 https://news.ycombinator.com/item?id=37780199>.
- ChrisArchitect 2y ago[dupe] Lots more discussion: https://news.ycombinator.com/item?id=40297027 https://news.ycombinator.com/item?id=40297027 https://news.ycombinator.com/item?id=40302792 https://news.ycombinator.com/item?id=40302792
- cosmin800 2y agoStackoverflow made millions, people got badges.
- lxgr 2y agoPeople also get a freely accessible database dump licensed under CC-BY-SA. That's more than almost every other content-oriented platform out there gives back.
- SirMaster 2y agoPeople also got answers to questions from other professionals that probably helped them personally financially and professionally in some way.
- renewiltord 2y agoFascinating. Perhaps Stallman's greatest innovation is copyleft. The idea of required reciprocation seems to tie deeply into people's views. For the little OSS I have I picked it by default but would gladly use BSD or some safe PD license on the other hand. The idea of Pillaging the Commons is interesting. I wonder when the mainstream opinion started shifting. It wasn't quite sudden but it seems to me that even ten years ago the dominant Internet visible position was that much of copyright was bogus: information wants to be free / if a pirate copies your stuff, you still have it But now there's a stronger sense of "this information is ours". Perhaps that subculture moved somewhere and this one came here or perhaps I moved from where the former was to where the latter is. I find it an interesting sociological phenomenon.
- throwaway290 2y agoUltimately they are digging their own grave (who needs SO if you can ask ChatGPT and shell out a few bucks to Microsoft eh?) If anyone else thinks it's a bad move, the most efficient way to boycott it is by adding senseless questions and answers and upvoting them. Us people is how they got big, us people is how they go down.
- elforce002 2y agoIf people stop posting, even chatGPT will be useless since it can't keep up with current bugs, new ways to solve them, etc...
- elforce002 2y agoStackOverflow is destroying its brand one step at a time. ChatGPT is good for boilerplate and generic errors. I still use SO for complex errors that are still getting updated with new ways to solve that said error, etc... Without the community posting and answering questions, ChatGPT won't work anymore. They (SO) have to block any scrapping, wait until GPU prices go down, or use open-source LLMs with their data and figure out how to monetize that service (maybe giving some percentage to devs, etc...) or even check if that approach (add a chatbot service) makes sense. They just caved to the FOMO mentality without considering that tech is always evolving and they need engineers, dev, etc... to keep finding out bugs, writing about their experiences on how they solved those errors, etc... We'll see if a different contender figures this out and comes out with that solution if SO doesn't change its course.