8 ms·
Another way to look at it is anyone can look at source available code to learn how to program without breaking a license.
by IncreasePosts 2y ago
Another way to look at it is anyone can look at source available code to learn how to program without breaking a license.
- deleted 2y ago[deleted]
- anileated 2y agoAnyone can, that’s orthogonal. This is about an automated tool that launders copyright at scale, generating revenue for its operator. (And if you seriously say that this tool is learning how to program, ask yourself if that tool’s operator is effectively a slave owner.)
- stale2002 2y ago> Anyone can, that’s orthogonal. Ok. So anyone "can" use a computer to do the same thing then. With the added part of "using a computer" it is now directly comparable and it is allowed. > And if you seriously say that this tool is learning how to program The tool is used by a person. The person is the one who takes the action, not the computer. So the point stands.
- anileated 2y agoIf you “use a computer” to watch a pirated film, does that make the practice legal? > The tool is used by a person. The person is the one who takes the action, not the computer. So the point stands. If watching that pirated film helps you learn something, does that make it legal? If the film was pirated not by you but by some for-profit company that charges you for watching it, does that make it legal?
- stale2002 2y ago> Can you “use a computer” to watch a pirated film? Sure. Is it legal? Nah. In many circumstances you can't mass distribute completely identical, non transformative, non fair use copies of large portions other people's copyrighted works, if thats what you meant. But there are many exceptions to that rule where you are allowed to use or distribute other people's works. And just like a human being is allowed to use other people's copyrighted works in those many exceptions, a human is also allowed to use a computer to take advantage of those legal exceptions. The only point here is that when you brought up that this uses a computer in your first post, thats not really a relevant detail. A person can use those exceptions that allow them to use other people's copyrighted works, and they can do that with or without a computer and it is legal in those exceptions either way. > If watching that pirated film helps you learn something, does that make it legal? > If the film was pirated not by you but by some for-profit company that charges you for watching it, does that make it legal? It depends on many factors. Yes there are many cases where yes it is legal to use other people's works. Edit: Evidence that I am right: you are right now commenting on a thread where a judge threw out all the copyright claims.
- anileated 2y ago> In most circumstances you can't mass distribute completely identical, non transformative, non fair use copies of large portions other people's copyrighted works That law was defined long before there was a capability to launder authorship at scale in the way being discussed. The law does not account for this novel capability. The law is intended to protect IP, which promotes innovation and creativity by creating relevant incentives. If that was the intention of the law, and it is not interpreted in that way, it ought to be revised for it to continue to serve those objectives. > Evidence that I am right: you are right now commenting on a thread where a judge threw out all the copyright claims. This only shows that you read the headline. It does not show that you (or the judge) are correct about the core issue.
- stale2002 2y ago> The law does not account for the new capability. Gotcha. Well, fortunately, you are commenting on a post right now where the judge threw out the copyright claims. So, apparently, I am correct that in this circumstance, that there is no illegal copyright infringement. > promotes innovation and creativity I'm this circumstance, it does seem to be promoting innovation and creativity because the AI stuff is allowed! Glad you agree.
- anileated 2y agoIt’s not a discussion about how the law is being interpreted by a particular court; that much is clear. It’s about how it ought to be interpreted.
- deleted 2y ago[deleted]
- danielmarkbruce 2y agoIt's more likely that relevant legislation needs to change if folks want the law to be different, rather than look to courts.
- drdeca 2y agoSaying that it “launders” only makes sense under the position you are claiming. So, it might be fine as a conclusion/claim, which I guess is how you’re using it, but it wouldn’t be good to use as part of an argument leading to your conclusion. (I didn’t phrase that well…) I generally don’t consider “learn” to apply only to entities which have the rights of a person, and of which ownership would amount to slavery. It is a common saying “You can’t teach an old dog new tricks.”. It is widely understood that, in contrast, one can often teach a young dog new tricks. The dog, in this case, learns the trick. We do not generally consider training an animal to do a task to be slavery. Well, some vegans might? But it is far from a typical view of the word “slavery”. So, am I saying that these language models are as rights-having and mind-having as a dog? No, much less so. Still, I have no objection to the word “learn” being used in this way.
- anileated 2y agoSome people say “it’s OK to ingest copyrighted material automatically at scale, since it’s for learning purposes”. They use two kinds of arguments for this. Argument A: A1. It’s a basic human right to be able to learn from things you see. You browse the Internet, you read some source code, you learn. Doesn’t matter what’s the license, you are free to do this. A2. It’s called “machine learning”, so the machine does the same. A3. Machine learning can use any content its operators can get a hold of. This is obviously wrong, because machine is being assigned human rights. We can argue about what exactly are pre-requisites for something to be granted human rights—it’s maybe not a specific physiology (some might say certain smart non-humanoid animals deserve it), but it’s pretty certainly sentience and consciousness. Meanwhile, the whole reason AI tech is big is that there is supposed to no sentient being who would understand (and therefore deserve any right to be treated well and get rewarded). If you take that away and grant AI human rights, then there is no point in this tech. So, either the machine has human-level sentience and is being forced to work (which humans famously tend to consider “slavery”), born and killed on demand, etc., or the machine is not learning in the sense under consideration because it’s an unthinking tool for its human operator. Which brings us to argument B: B1. It’s a basic human right to be able to learn from things you see. You browse the Internet, you read source code, you learn. Doesn’t matter what’s the license, you are free to do this. B2. If you use a computer or [insert technology] to learn, that’s OK. B3. An LLM is just another instance of that technology. You use LLM and you learn. This is wrong for slightly more subtle reasons, but on bright side there’s multiple of them. First, it’s not clear that someone learns while using Copilot to produce a work for them. If I asked Copilot to write me a Fibonacci number generator, have I learned how to write it? If I ask Midjourney to draw me 2055 Los Angeles skyline in the style of Picasso, did I learn how to draw? Second, and this is a crucial fallacy, making a computer famously does not require ingesting all of the copyrighted material you can subsequently access through that computer. Said computer can exist just fine without it; the LLMs, however, cannot. The inputs (knowledge and parts) required to produce the computer you’re using were largely obtained through ordinary ways (patents licensed, hardware paid for), whereas the inputs required to produce an LLM have been, some would say, effectively stolen.
- golergka 2y ago> Anyone can, that’s orthogonal. That's exactly what happens here. In this case anyone happens to be an LLM.
- anileated 2y agoLLM is not “anyone”, because LLM is a thing but “anyone” refers to people. If you consider LLMs people, then you should ask yourself whether they are suffering abuse from being treated the way they are by their operators.
- naasking 2y ago> And if you seriously say that this tool is learning how to program, ask yourself if that tool’s operator is effectively a slave owner. This doesn't follow. I don't see why knowledge and intelligence necessarily entail that it has a desire for autonomy, which is why slavery is really abhorrent.
- darby_nine 2y agoTo me, the term "learning" as opposed to "training" entails autonomy.
- anileated 2y agoAnd even still, these words are used in many, sometimes mutually exclusive, meanings (“learn” as in “machine learning” is a far cry from “learn” as in “live and learn”). I wonder how the courts could even properly consider all implications if these words don’t have precise legal definitions all the way down to what it means being a human.
- BobaFloutist 2y agoIf we could train the desire for autonomy out of humans, it wouldn't make human slavery any less abhorrent, even if they volunteered for the process and/or were well compensated.
- naasking 2y agoIt absolutely would make it less abhorrent. Maybe you think it would still be abhorrent, but this is debatable. People literally do consent to slavery-like roles in places like the BDSM community, and some people might find it distasteful but not illegal or morally abhorrent, because these people still have the autonomy to opt-out at any point. I also doubt training out the desire for autonomy is possible. Explore-exploit is fundamental to any kind of decision making, such as food foraging. That inclination goes deeper than higher brain functions.
- 2y ago
- deleted 2y ago[deleted]
- Spivak 2y agoThe significant step here is anything can do what you say. Because there's no human in the loop looking at source code and learning from it. You have an autonomous system that's ingesting copyrighted material, doing math on it, storing it, and producing outputs on user requests. There's no learning or analogy to humans, the court is ruling that this particular math is enough to wash away bit color. The ruling was based on the outputs and the reasonable intent of the people who created it and what they are trying to accomplish, not how it works internally. It's not the first, if you take copyrighted data and && 0x00 to all of it that certainly washes the bits too.
- RandallBrown 2y ago> You have an autonomous system that's ingesting copyrighted material, doing math on it, storing it, and producing outputs on user requests People are also autonomous systems that ingest copyrighted material, do "math" on it, store it, and produce outputs on user requests. The real difference is the scale at which a computer can ingest copyrighted material is MUCH greater than what a person can do. Does that make it illegal? Maybe, maybe not.
- Spivak 2y agoAm I in a bad sci-fi novel? People aren't machines! How is this a such a difficult concept? LLMs have as much thought as quicksort. I swear to god humans will anthropomorphize everything except ourselves. Do y'all's salaries depend on this or something? There is no rule that says "If a human can do something, a computer program instructed by a human can do the same thing." Hell that rule doesn't even exist for humans acting as stand-ins. I can't send someone I hire out of the country and have them use my passport. It's why you can watch a movie in a theater but an autonomous system working on your behalf, a camera, can't. Github made a tool, it's as alive as a hammer. It "learns" as much as your programmable pad lock. Whether or not the human employees of Github are allowed to use copyrighted material to make that tool, and whether the human employees of Github are performing a copyrighted work when users make use of the tool is the legal question. Y'all wouldn't survive the https://en.wikipedia.org/wiki/Philosophical_zombie https://en.wikipedia.org/wiki/Philosophical_zombie apocalypse.
- fsflover 2y agoIt's not so straightforward: https://en.wikipedia.org/wiki/Clean-room_design https://en.wikipedia.org/wiki/Clean-room_design
- dgfitz 2y ago> Another way to look at it is anyone can look at source available code to learn how to program without breaking a license. Yes, and exactly ZERO amount of money have exchanged hands in this scenario. Learning is dope, the more the better. The difference is, someone makes money off it, and not the persons(s) that wrote the code. This is not a valid argument
- threatofrain 2y agoLearning without making money does not shield you from copyright violation, otherwise public school teachers would just start saving money by copying whole texts. We don’t live in a society where it’s okay for an 8 year old to say “but I’m just trying to learn, I’m not a business!” And making new music which is a synthesis of your life experience with copyrighted music does not mean copyright violation, regardless if you’re making money or compensating all the authors who’ve inspired you.
- pc86 2y ago> We don’t live in a society where it’s okay for an 8 year old to say “but I’m just trying to learn, I’m not a business!” I mean we absolutely do if you're an 8 year old. Except in the most NIMBY HOA-driven areas of the culture nobody expects a kid setting up a lemonade stand to get a business license or submit to health code inspections.
- deleted 2y ago[deleted]
- darby_nine 2y ago> otherwise public school teachers would just start saving money by copying whole texts This is literally a thing today. It may be illegal but the idea of prosecuting this is insane.
- deleted 2y ago[deleted]
- pc86 2y agoIn this example a person is looking at code they can't legally copy, learning from it, and re-implementing the same functionality. Someone's definitely making money off of that. That person, that person's employer, clients and vendors, lots of people. People get upset about AI because 1) the scale is much bigger because no human can read and generally remember all the code on GitHub while a sufficiently large model can, 2) it's a lot easier to prompt an AI into giving you a passable MVP than it is to code one from scratch, ESPECIALLY as a junior or even mid level, 3) there are unlikeable billionaires making money now where there weren't before.
- saurik 2y agoAnd if you then write a program that is remarkably similar to the one you read, that's copyright infringement. As another reply noted--but without anywhere near enough verbosity--this is not without risk, and people who intend to work on similar systems often try to use a strategy where they burn one engineer by having them read the original code, have them document it carefully with a lawyer to remove all expressive aspects, and then have a separate engineer develop it from the clean documents.
- neilv 2y ago> strategy where they burn one engineer by having them read the original code, have them document it carefully with a lawyer to remove all expressive aspects, and then have a separate engineer develop it from the clean documents. Interesting. What kinds of situations is that strategy used for? (I'm familiar with cleanroom, which I understand means that you start with un-tainted engineers, who've credibly never been exposed to the proprietary IP, the work only from unencumbered public documentation and running the system as an opaque box. Then there's also validation, like with parallel systems and fuzzing. But I haven't thought through in what situations this might not work, so might require the tainted documenting approach.)
- TrueDuality 2y agoThis is the full or classic version of clean room reverse engineering. Using unencumbered public documentation is relatively new, that kind of detailed documentation wasn't widely available. Car manufacturers still protect their service manuals with an agreement that basically says they can't be used for this but I think a lot of service centers stopped making people sign them. The classic tech story that used this technique is the IBM BIOS and the resulting spread of "IBM PC-Compatible" machines. There is a little bit about it on the wikipedia page (https://en.wikipedia.org/wiki/IBM_PC%E2%80%93compatible https://en.wikipedia.org/wiki/IBM_PC%E2%80%93compatible). Random factoid, the Netflix Original "Halt and Catch Fire" has a depiction of doing this IBM clone reverse engineering and did a pretty good job at it.
- kelnos 2y agoThe strategy described by the GP is clean-room (reverse) engineering.
- ml-anon 2y agoFor the slow ones among us: Machine "learning" is not human learning. It is not similar to, analogous to, or in any way remotely comparable.
- shadowgovt 2y agoIt seems obvious how they're comparable, in the same way that you can compare a parrot talking to human speech. Black-box both systems and there's enough similarity to make a layperson go "Huh. Those look remarkably similar," even if the mathematicians among us know the underlying mechanisms, inputs, and outputs are quite different.
- ml-anon 2y agoif the inputs, mechanisms and outputs are different then they're...not comparable?
- croes 2y agoIsn't fascinating that the same isn't true for books and music. If it's too similar you get sued
- shadowgovt 2y agoI'm not exactly sure, but I think the underlying philosophy here is that code is a lot more like math than like music, and you can't copyright math. So to have any copyright protection at all for code, the Office had to carve a narrow trail where the standard for copying is higher, because there are plenty of circumstances where there is only one right (or most optimal) algorithm, and there's no protection for the algorithm itself.
- croes 2y agoMusic is pretty much math too.
- shadowgovt 2y agoAbsolutely, as Ada Lovelace correctly observed. But IIUC copyright law doesn't generally recognize that association in any deep way.
- kelnos 2y agoAnd anyone can also look at that source available code, write their own version, distribute it, be sued for copyright infringement, and lose in court, because their version is too similar to the original.