7 ms·
I am pretty sure this article is predicated on a misunderstanding of what a "clean room" implementation means. It does not mean "as long as you never read the o
by danlitt 7mo ago
I am pretty sure this article is predicated on a misunderstanding of what a "clean room" implementation means. It does not mean "as long as you never read the original code, whatever you write is yours". If you had a hermetically sealed code base that just happened to coincide line for line with the codebase for GCC, it would still be a copy. Traditionally, a human-driven clean room implementation would have a vanishingly small probability of matching the original codebase enough to be considered a copy. With LLMs, the probability is much higher (since in truth they are very much not a "clean room" at all).
The actual meaning of a "clean room implementation" is that it is derived from an API and not from an implementation (I am simplifying slightly). Whether the reimplementation is actually a "new implementation" is a subjective but empirical question that basically hinges on how similar the new codebase is to the old one. If it's too similar, it's a copy.
What the chardet maintainers have done here is legally very irresponsible. There is no easy way to guarantee that their code is actually MIT and not LGPL without auditing the entire codebase. Any downstream user of the library is at risk of the license switching from underneath them. Ideally, this would burn their reputation as responsible maintainers, and result in someone else taking over the project. In reality, probably it will remain MIT for a couple of years and then suddenly there will be a "supply chain issue" like there was for mimemagic a few years ago.
- dathinab 7mo agothe author speaks about code which is syntactically completely different but semantically does the same i.e. a re-implementation which can either - be still derived work, i.e. seen as you just obfuscating a copyright violation - be a new work doing the same nothing prevents an AI from producing a spec based on a API, API documentation and API usage/fuzzing and then resetting the AI and using that spec to produce a rewrite I mean "doing the same" is NOT copyright protection, you need patent law for that. Except even with patent law you need to innovations/concepts not the exact implementation details. Which means that even if there are software patents (theoretically,1) most things done in software wouldn't be patentable (as they are just implementation details, not inventions) (1): I say theoretically because there is a very long track record of a lot of patents being granted which really should never be granted. This combined with the high cost of invalidating patents has caused a ton of economical damages.
- jacquesm 7mo agoNo, that depends on whether or not the AI work product rests on key contributions to its training set without which it would not be able to the the work, see other comment. In that case it looks like 'a new work doing the same' but it still a derived work. Ted Nelson was years ahead of the future where we really needed his Xanadu to keep track of fractional copyright. Likely if we had such a mechanism, and AI authors respected it then we would be able to say that your work is derived from 3000 other original works and that you added 6 lines of new code.
- uyzstvqs 7mo agoNo, training and inference are two separate processes. Training data is never redistributed, only obtained and analyzed. What matters is what data is put into context during inference. This is controlled by the user. AI/ML is complex, so as a simpler analogy: If I watch The Simpsons, and I create an amusing infographic of how often Homer says "D'oh!" over time, my infographic would be an original work. AI training follows the same principle.
- jacquesm 7mo ago> my infographic would be an original work. > AI training follows the same principle. If you really believe that then we can't have a meaningful conversation about this, that's not even ELIF territory, that's just disconnected. You should be asking questions, not telling people how it works.
- ndriscoll 7mo agoHow exactly is it different? All the model itself is is a probability distribution for next token given input, fitted to a giant corpus. i.e. a description of statistical properties. On its own it doesn't even "do" anything, but even if you wrap that in a text generator and feed it literal gcc source code fragments as input context, it will quickly diverge. Because it's not a copy of gcc. It doesn't contain a copy of gcc. It's a description of what language is common in code in general. In fact we could make this concrete: use the model as the prediction stage in a compressor, and compress gcc with it. The residual is the extent to which it doesn't contain gcc.
- pmarreck 7mo ago> With LLMs, the probability is much higher (since in truth they are very much not a "clean room" at all). I beg to differ. Please examine any of my recent codebases on github (same username); I have cleanroom-reimplemented par2 (par2z), bzip2 (bzip2z), rar (rarz), 7zip (z7z), so maybe I am a good test case for this (I haven't announced this anywhere until now, right here, so here we go...) https://github.com/pmarreck?tab=repositories&type=source https://github.com/pmarreck?tab=repositories&type=source I was most particular about the 7zip reimplementation since it is the most likely to be contentious. Here is my repo with the full spec that was created by the "dirty team" and then worked off of by the LLM with zero access to the original source: https://github.com/pmarreck/7z-cleanroom-spec https://github.com/pmarreck/7z-cleanroom-spec Not only are they rewritten in a completely different language, but to my knowledge they are also completely different semantically except where they cannot be to comply with the specification. I invite you and anyone else to compare them to the original source and find overt similarities. With all of these, I included two-way interoperation tests with the original tooling to ensure compatibility with the spec.
- ostacke 7mo agoBu that's not really what danlitt said, right? They did not claim that it's impossible for an LLM to generate something different, merely that it's not a clean room implementation since the LLM, one must assume, is trained on the code it's re-implementing.
- galaxyLogic 7mo agoBUt LLM has seen millions (?) of other code-bases too. If you give it a functional spec it has no reason to prefer any one of those code-bases in particular. Except perhaps if it has seen the original spec (if such can be read from public sources) associated with the old implementation, and the new spec is a copy of the old spec.
- sarchertech 7mo agoYes if you are solving the exact problem that the original code solved and that original code was labeled as solving that exact problem then that’s very good reason for the LLM to produce that code. Researchers have shown that an LLM was able to reproduce the verbatim text of the first 4 Harry Potter books with 96% accuracy.
- zabzonk 7mo ago> It does not mean "as long as you never read the original code, whatever you write is yours" I think there is precedence that says exactly this - for example the BIOS rewrites for the IBM PC from people like Phoenix. And it would be trivial to instruct an LLM to prefer to use (say, in assembler) register C over register B wherever that was possible, resulting in different code.
- bandrami 7mo agoDifferent but still derivative
- zabzonk 7mo agoWell, I am not exactly a hotshot 8086 programmer (though I do alright) but if I was asked to reproduce the IBM BIOS (which I have seen) I think I would come up with something very similar but not identical - it is really not rocket science code, so the LLM replacing me would have rather few alternatives to choose from.
- fc417fc802 7mo agoI believe those are actually separate matters. A proper clean room implementation on the one hand, and the question of whether or not a particular outcome was a foregone conclusion on the other. I don't recall where I saw the latter but it might have come up during Google v Oracle?
- danlitt 7mo agoAs long as you never read the original code, it is very likely that whatever you write is yours. So I would not be surprised to read judges indicating in this direction. But I would be a little surprised to find out this was an actual part of the test, rather than an indication that the work was considered to have been copied. There are for instance lots of ways of reproducing copyrighted work without using a copy directly, but naive methods like generating random pieces of text are very time consuming, so there is not much precedence around them. LLMs are much more efficient at it!
- petercooper 7mo agoThe actual meaning of a "clean room implementation" is that it is derived from an API and not from an implementation I know you were simplifying, and not to take away from your well-made broader point, but an API-derived implementation can still result in problems, as in Google vs Oracle [1]. The Supreme Court found in favor of Google (6-2) along "fair use" lines, but the case dodged setting any precedent on the nature of API copyrightability. I'm unaware if future cases have set any precedent yet, but it just came to mind. [1]: https://en.wikipedia.org/wiki/Google_LLC_v._Oracle_America,_Inc https://en.wikipedia.org/wiki/Google_LLC_v._Oracle_America,_....
- lokar 7mo agoYeah, a cleanroom re-write, or even "just" a copy of the API spec is something to raise as a defense during a trial (along with all other evidence), it's not a categorical exemption from the law. Also, I find it important that here the API is really minimal (compared to the Java std lib), the real value of the library is in the internal detection logic.
- danlitt 7mo agoThis is exactly what I had in mind when I said I was simplifying :) it is a valid point.
- femto 7mo ago> If you had a hermetically sealed code base that just happened to coincide line for line with the codebase for GCC, it would still be a copy. That's not what the law says [1]. If two people happen to independently create the same thing they each have their own copyright. If it's highly improbable that two works are independent (eg. the gcc code base), the first author would probably go to court claiming copying, but their case would still fail if the second author could show that their work was independent, no matter how improbable. [1] https://lawhandbook.sa.gov.au/ch11s13.php?lscsa_prod%5Bpage%5D=33 https://lawhandbook.sa.gov.au/ch11s13.php?lscsa_prod%5Bpage%...
- jerf 7mo agoIt is true that if two people happen to independently create the same thing, they each have their own copyright. It is also true that in all the cases that I know about where that has occurred the courts have taken a very, very, very close look at the situation and taken extensive evidence to convince the court that there really wasn't any copying. It was anything but a "get out of jail free" card; it in fact was difficult and expensive, in proportion to the size of the works under question, to prove to the court's satisfaction that the two things really were independent. Moreover, in all the cases I know about, they weren't actually identical, just, really really close. No rational court could possibly ever come to that conclusion if someone claimed a line-by-line copy of gcc was written by them, they must have independently come up with it. The probably of that is one out of ten to the "doesn't even remotely fit in this universe so forget about it". The bar to overcoming that is simply impossibly high, unlike two songs that happen to have similar harmonies and melodies, given the exponentially more constrained space of "simple song" as compared to a compiler suite.
- brians 7mo agoI do not agree with your interpretation of copyright law. It does ban copies: there has to be information flow from the original to the copy for it to be a "copy." Spontaneous generation of the same content is often taken by the courts to be a sign that it's purely functional, derived from requirements by mathematical laws. Patent law is different and doesn't rely on information flow in the same way.
- BoredPositron 7mo agoWell discovery might be a fun exercise to see if the code is in the dataset of the llm.
- bjord 7mo agoif?
- kevin_thibedeau 7mo agoDerivative works can also run afoul of copyright. An LLM trained on a corpus of copyrighted code is creating derivative works no matter how obscure the process is.
- wareya 7mo agoThis actually isn't what legal precedent currently says. The precedent is currently looking at actual output, not models being tainted. If you think this is morally wrong, look into getting the laws changed (serious).
- Georgelemental 7mo agoWhat about a human trained on having 30 years of experience working with copyrighted codebases?
- mftrhu 7mo agoSaid human would likely not be able to create a clean-room implementation of any of the codebases they worked on.
- deleted 7mo ago[deleted]
- thousand_nights 7mo agothe whole concept of a "clean room" implementation sounds completely absurd. a bunch of people get together, rewrite something while making a pinky promise not to look at the original source code guaranteeing the premise is basically impossible, it sounds like some legal jester dance done to entertain the already absurd existing copyright laws
- Forgeties79 7mo agoHalt and Catch Fire did a pretty funny rendition of this song and dance
- dudeinhawaii 7mo agoIt usually refers to situations without access to the source code. I've always taken "clean room" to be the kind of manufacturing clean room (sealed/etc). You're given a device and told "make our version". You're allowed to look, poke, etc but you don't get the detailed plans/schematics/etc. In software, you get the app or API and you can choose how to re-implement. In open source, yes, it seems like a silly thing and hard to prove.
- myrmidon 7mo ago> it sounds like some legal jester dance done to entertain [...] copyright laws Clean room implementations are a jester dance around the judiciary. The whole point is to avoid legal ambiguity. You are not required to do this by law, you are doing this voluntarily to make potential legal arguments easier. The alternative is going over the whole codebase in question and arguing basically line by line whether things are derivative or not in front of a judge (which is a lot of work for everyone involved, subjective, and uncertain!).
- bandrami 7mo agoIn the archetypal example IBM (or whoever it was) had to make sure the two engineering teams were never in the cafeteria together at the same time
- foooorsyth 7mo ago>The actual meaning of a "clean room implementation" is that it is derived from an API and not from an implementation This is incorrect and thinking this can get you sued https://en.wikipedia.org/wiki/Structure,_sequence_and_organization https://en.wikipedia.org/wiki/Structure,_sequence_and_organi...
- umvi 7mo agoYou can be sued for any reason if a company feels threatened (see: Oracle v Google)
- amiga386 7mo agoWhether you get sued is more on the plaintiff than you. Per your link, the Supreme Court's thinking on "structure, sequence and organization" (Oracle's argument why Google shouldn't even be allowed to faithfully produce a clean-room implementation of an an API) has changed since the 1980s out of concern that using it to judge copyright infringement risks handing copyright holders a copyright-length monopoly over how to do a thing: > enthusiasm for protection of "structure, sequence and organization" peaked in the 1980s [..] This trend [away from "SS&O"] has been driven by fidelity to Section 102(b) and recognition of the danger of conferring a monopoly by copyright over what Congress expressly warned should be conferred only by patent The Supreme Court specifically recognised Google's need to copy the structure, sequence and organization of Java APIs in order to produce a cleanroom Android runtime library that implemented Java APIs so that that existing Java software could work correctly with it. Similarly, see Oracle v. Rimini Street (https://cdn.ca9.uscourts.gov/datastore/opinions/2024/12/16/23-16038.pdf https://cdn.ca9.uscourts.gov/datastore/opinions/2024/12/16/2...) where Rimini Street has been producing updates that work with Oracle's products, and Oracle claimed this made them derivative works. The Court of Appeals decided that no, the fact A is written to interoperate with B does not necessarily make A a derivative work of B.
- danlitt 7mo agoI did not expect people to take "API" so literally. This point is what I was referring to when I said "I am simplifying slightly". The point is that a clean room impl begins from a specification of what the software does, and that the new implementation is purported to be derived only from this. What I am trying to say is that "not looking at the implementation" is not exactly the point of the test - that is a rule of thumb, which works quite well for avoiding copyright infringement, but only when humans do it.
- wareya 7mo ago> If you had a hermetically sealed code base that just happened to coincide line for line with the codebase for GCC, it would still be a copy. If you somehow actually randomly produce the same code without a reference, it's not a copy and doesn't violate copyright. You're going to get sued and lose, but platonically, you're in the clear. If it's merely somewhat similar, then you're probably in the clear in practice too: it gets very easy very fast to argue that the similarities are structural consequences of the uncopyrightable parts of the functionality. > The actual meaning of a "clean room implementation" is that it is derived from an API and not from an implementation (I am simplifying slightly). This is almost the opposite of correct. A clean room implementation's dirty phase produces a specification that is allowed to include uncopyrightable implementation details. It is NOT defined as producing an API, and if you produce an API spec that matches the original too closely, you might have just dirtied your process by including copyrightable parts of the shape of the API in the spec. Google vs Oracle made this more annoying than it used to be. > Whether the reimplementation is actually a "new implementation" is a subjective but empirical question that basically hinges on how similar the new codebase is to the old one. If it's too similar, it's a copy. If you follow CRRE, it's not a copy, full stop, even if it's somehow 1:1 identical. It's going to be JUDGED as a copy, because substantial similarity for nontrivial amounts of code means that you almost certainly stepped outside of the clean room process and it no longer functions as a defense, but if you did follow CRRE, then it's platonically not a copy. > What the chardet maintainers have done here is legally very irresponsible. I agree with this, but it's probably not as dramatic as you think it is. There was an issue with a free Japanese font/typeface a decade or two ago that was accused of mechanically (rather than manually) copying the outlines of a commercial Japanese font. Typeface outlines aren't copyrightable in the US or Japan, but they are in some parts of Europe, and the exact structure of a given font is copyrightable everywhere (e.g. the vector data or bitmap field for a digital typeface, as opposed to the idea of its shape). What was the outcome of this problem? Distros stopped shipping the font and replaced it with something vaguely compatible. Was the font actually infringing? Probably not, but better safe than sorry.
- danlitt 7mo ago> If you somehow actually randomly produce the same code without a reference, it's not a copy and doesn't violate copyright. I don't believe this, and I doubt that the sense of copying in copyright law is so literal. For instance, if I generated the exact text of a novel by looking for hash collisions, or by producing random strings of letters, or by hammering the middle button on my phone's autosuggestion keyboard, I would still have produced a copy and I would not be safe to distribute it. There need not have been any copy anywhere near me for this to happen. Whether it is likely or not depends on the technique used - naive techniques make this very unlikely, but techniques can improve. It is also true that similarity does not imply copying - if you and I take an identical photograph of the same skyline, I have not copied you and you have not copied me, we have just fixed the same intangible scene into a medium. The true subjective test for copying is probably quite nuanced, I am not sure whether it is triggered in this case, but I don't think "clean room LLMs" are a panacea either. > dirty phase produces a specification ... it is NOT defined as producing an API This does not really sound like "the opposite of correct". APIs are usually not copyrightable, the truth is of course more complicated, if you are happy to replace "API" with "uncopyrightable specification" then we can probably agree and move on. > it's probably not as dramatic as you think it is In reality I am very cynical and think nothing will come of this, even if there are verbatim snippets in the produced code. People don't really care very much, and copyright cases that aren't predicated on millions of dollars do not survive the court system very long.
- jen20 7mo ago> What the chardet maintainers have done here is legally very irresponsible. Perhaps the maintainer wants to force the issue? > Any downstream user of the library is at risk of the license switching from underneath them. Checking the license of the transitive closure of your dependencies is table stakes for using them.
- danlitt 7mo ago> Perhaps the maintainer wants to force the issue? I doubt it, and I don't see any evidence that's what they're doing. There are probably better ways, if that's what they want. > Checking the license of the transitive closure of your dependencies is table stakes for using them. Checking the license of the transitive closure of your dependencies is only feasible when the library authors behave responsibly.
- fc417fc802 7mo agoThe problem is that the transitive closure isn't clear here. One of the entries is being claimed to be one thing but might in fact turn out to be another.
- j45 7mo agoThis reminds me of a full rewrite. When a developer reimplements a complete new version of code from scratch, with an understanding only, a new implementation generally should be an improvement on any source code not equal. In today’s world, letting LLMs replicate anything will generate average code as “good” and generally create equivalent or more bloat anyways unless well managed.
- StilesCrisis 7mo agoThe world is chock-full of rewrites that came out disastrously worse than the thing they intended to replace. One of Spolsky's most-quoted articles of all time was about this. https://www.joelonsoftware.com/2000/04/06/things-you-should-never-do-part-i/ https://www.joelonsoftware.com/2000/04/06/things-you-should-... > They did it by making the single worst strategic mistake that any software company can make: They decided to rewrite the code from scratch.
- j45 7mo agoOh, for sure, rewrites generally do fail especially if the incoming lessons from the existing version aren't clear. Finding a middle ground of building a roadmap to refactoring your way forward is often much better. Appreciate the Joel link, nice to see that kind of stuff again. With that being said if it's the same small team that built the first version, there can be a calculated risk to driving a refactor towards a rewrite with the right conditions. I says this because I have been able to do it in this conditions a few times, it still remains very risky. If it's a new or different team later on trying to rewrite, all bets are off anyways. We have to remember 70% of software projects fail at the best of times, independent of rewrites.
- aaron695 7mo ago[dead]