4 ms·
Content contributed on StackOverflow (and all the StackExchange sites I believe) is CC, so it's really a losing battle.
by fer 2y ago
Content contributed on StackOverflow (and all the StackExchange sites I believe) is CC, so it's really a losing battle.
- schlauerfox 2y agoHow is the use of CC-BY-4.0 by a LLM going to be attributed to satisfy the license conditions is my question. Every answer by the model contains the work as part of it's token weights and can regurgitate.
- trueismywork 2y agoYup, you need to cite every single poster ever.
- manuelabeledo 2y agoWouldn't a compliant solution be just a glorified search engine?
- jessriedel 2y agoCopyright covers the words not the ideas/concepts/facts. If I read a StackExchange post, I can absolutely use those ideas/concepts/facts to write a new document, which I can choose to copyright and sell, and there's no need for me to cite the SE post. Now, as a matter of academic courtesy it's nice to cite the SE post if it's feasible, but it's impossible for me to remember where I learned everything I've learned over the years. Likewise, an LLM shouldn't be forced to cite where it learned facts, which might be technically hard/impossible, but if there is a clear best source then as a matter of politeness it should. And LLMs mostly do this when they can!
- manuelabeledo 2y ago> Now, as a matter of academic courtesy it's nice to cite the SE post if it's feasible, but it's impossible for me to remember where I learned everything I've learned over the years. Likewise, an LLM shouldn't be forced to cite where it learned facts, which might be technically hard/impossible, but if there is a clear best source then as a matter of politeness it should. And LLMs mostly do this when they can! Scientific papers have included references for as long as I can remember. It's not about courtesy, but precision and acknowledgement. So, are we letting LLMs run free and do things that would be frowned upon if they were done by humans?
- jessriedel 2y agoWhen LLMs are used to write academic papers, then there is definitely a stronger requirement on the authors deploying them for identifying the source of ideas, as a matter of professional ethics (not law). But this isn't the issue being discussed, and it's not anything new with LLMs. Long before LLMs, academic authors occasionally got ideas from other places StackExchange without citing them properly in papers. (They frequently learn things from StackExchange, but most of the time they don't need to cite them because it's part of common knowledge. Cases where they got novel, cite-necessary insights is much rarer, as it will be for LLMs.)
- wang_li 2y agoLLMs don’t learn facts. They don’t think, cogitate, daydream, hallucinate, lie, or any other activity that requires intelligence, a mind, or a personality. Writing about them the way you did just advances the scam that is the current AI industry.
- jessriedel 2y agoWhen I say "my thermostat tries to keep the temperature at 72", this doesn't mean I'm asserting anything about the thermostat having "goals" in some deep philosophical sense. It just means that the behavior of the thermostat can be usefully understood, in the relevant context, as goal-directed. (Likewise, "legs are for walking" even though legs were designed through a evolutionary process, etc.) We use this sort of language all the time, it's very useful, and it's mere pedantry to demand people not to. Likewise here, a machine can absolutely "learn" facts in the sense that it extracts them from the copyright-able creative choices in the text written by a human author.
- wang_li 2y agoYou can't say I'm using the word learn as a metaphor for some purely mechanistic and deterministic process and not like when people say learn as in middle school kids learning algebra and then switch to using the second definition to absolve OpenAI from following the license. That'd be like saying my computer learned the film Oppenheimer when bittorrent copied it to my computer's storage system, transcoding it to a lower bitrate, and claim I'm free of copyright law because my computer just learned some facts. All of OpenAI's products are databases with highly lossy compression applied. Using language that implies a mind is pure scam.
- jessriedel 2y agoI’m not using anything about “learn” to justify what’s copyrightable and what’s covered by fair use. Those things are defined by the output, not the process that produced them. The point is that LLMs are very effective at extracting non-copyrightable facts from the words and other creative choices made by an author.
- cjpearson 2y agoIt's a losing battle in that the CC license and StackOverflow's ToS make it clear that StackOverflow is allowed to restore the post regardless of the original author's wishes. It's less clear that the license gives AI companies the right to train LLMs. They tend to take the view that any media can be used for training regardless of licensing because it's fair use. Many authors disagree and believe they must explicitly license their content for LLM training. AFAIK none of these legal disputes have been settled yet. In my view you shouldn't really expect to have much control over who uses the content you post to Stack Overflow. Maybe your answer will be used by a CS101 student. But it could also be used by a defense contractor, gambling site, North Korean hacker, or darknet drug marketplace. Yes, the license requires attribution, but if ChatGPT ignores that requirement it's behaving like 90% of programmers.
- trueismywork 2y agoIts CC-BY-SA so openAI has to release the weights obtained by training.