8 ms·
AI and the Ship of Theseus
- Devasta 7mo agoThis is awful news, but I don't know what can be done, is it possible to have a new GPL4 that deals with this? I doubt it.
- moralestapia 7mo ago[flagged]
- jimmaswell 7mo agoThis entirely misses the point. Re-implementing code based on API surface and compatibility is established fair use if done properly (Compaq v. IBM, Google v. Oracle). There's nothing wrong with doing that if you don't like a license. What's in question is doing this with AI that may or may not have been trained on the source. In the instance in the article where the result is very different, it's probably in the clear regardless. I'm sympathetic to the author as I generally don't like GPL either outside specific cases where it works well like the Linux kernel.
- blell 7mo agoThis reminds me of people crying over toybox https://en.wikipedia.org/wiki/Toybox#Controversy https://en.wikipedia.org/wiki/Toybox#Controversy
- trueismywork 7mo agoThe real test would be to see how much of generated code is similar to the old code. Because then it is still a copyright. Just becsuse you drew mickey mouse from memory doesnt above you if it looks close enough to original hickey mouse.
- the_mitsuhiko 7mo ago> The real test would be to see how much of generated code is similar to the old code. I have looked at the project earlier today there is effectively no resemblance other than the public API.
- kccqzy 7mo agoThat’s I believe woefully inadequate. There are some levels of code similarity: Level 0: the code is just copied Level 1: the code only has white space altered so the AST is the same Level 2: the code has minor refactoring such as changing variables names and function names (in a compiled language the object code would be highly similar; and this can easily be detected by tools like https://github.com/jplag/JPlag https://github.com/jplag/JPlag) Level 3: the code has had significant refactoring such as moving functionality around, manually extracting code to new functions and manually inlining functions Level 4: the code does the same conceptual steps as the old code but with different internal architecture At least in the United States you have to reach Level 4 because only concepts are not copyrightable. And I believe chardet has indeed reached level 4 in this rewrite.
- coldtea 7mo ago>Licenses exists for a reason Yes, and the choice of license for a project is made for a reason that not necessarily everybody agree with. And the people who don't agree, have every right to implement a similar, even file-format or API compatible, project and give it another license. Gnumeric vs Excel, for example, or forks like MariaDB and Valkey. But whether they do that alternative licensed project or not, it's perfectly rational, to not like the choice of license the original is in. They legally have to respect it, but that doesn't mean there's anything irational to disliking it or wishing it was changed. And it's not merely idle wishing: sometimes it can make the original author/vendor to reconsider and switch license. QT is a big example. Blender. Or even proprietary to open (Mozilla to MPL). "It's so disgusting to see people who are either malicious or non mentally capable enough to understand this"
- moralestapia 7mo agoHmm ... you don't have to ask for consent. You just slap the license you want to your code and that's it. It's not some sort of democracy, lol, it's a set of exclusive rights that are created the moment the work being copyrighted is produced. (For a quick intro I recommend: https://www.youtube.com/watch?v=bxVs7FCgOig https://www.youtube.com/watch?v=bxVs7FCgOig) In the case of the license in question (L/GPL), it's one of the most strict ones out there, it explicitly forbids relicensing code under a different non-compatible license, like MIT; let me says that again, L/GPL EXPLICITLY FORBIDS the thing that happened here from happening. I sympathize with the guy that spent 12 years of his life maintaining the code, thank you for your service or something, but that does not make a difference. The wording of the (L/GPL) license is clear and the original author and most of the other 50 or so contributors did not approve of this.
- coldtea 7mo ago[flagged]
- moralestapia 7mo agoHey, you can definitely rewrite your argument without resorting to bad language. Take a look at the guidelines that keep this place together: https://news.ycombinator.com/newsguidelines.html https://news.ycombinator.com/newsguidelines.html
- the_mitsuhiko 7mo ago> "But I wish that car was free", sure pal, but it's not. Are you like, 8 years old? Just because things are not as one wants, does not stop that desire to be there. > When the author of a project choose a specific license s/he is making a deliberate decision. Potentially, potentially not. I used to release software under GPL and LGPL but changed my mind a few years after that. I did so in part because of conversations I had with others that convinced me that my values are closer aligned with permissive licenses. So engaging in a friendly discourse with a maintainer to ask them to relicense is a perfectly fine thing to do and an issue has been with chardet for many, many years on the license.
- deleted 7mo ago[deleted]
- scuff3d 7mo agoThe solution to this whole situation seems pretty simple to me. LLMs were trained on a giant mix of code, and it's impossible to disentangle it, but a not insignificant portion of their capabilities comes from GPL licenced code. Therefore, any codebase that uses LLM code is now GPL. You have a proprietary product? Not anymore. Not saying there's a legal precedent for that right now, but it's the only thing that makes any sense to me. Either that or retain the models on only MIT/similarly licenced code or code you have explicit permission to train on.
- keithnz 7mo agoif you train yourself by looking at GPL code then go implement your own things, is that code GPL?
- AberrantJ 7mo agoOf course not, because everyone making these arguments wants people to have some magic sauce so they get to ignore all the rules placed on the "artificial" thing.
- estimator7292 7mo agoIf you copy and paste one line from a thousand different GPL projects, is the resulting program GPL? Let's be honest about what's happening here.
- 7mo ago
- nomdep 7mo agoIn this emerging reality, the whole spectrum of open-source licenses effectively collapses toward just two practical choices: release under something permissive like MIT (no real restrictions), or keep your software fully proprietary and closed. These are fascinating, if somewhat scary, times.
- measurablefunc 7mo agoIf you listen to the people who believe real AI is right around the corner then any software can be recreated from a detailed enough specification b/c whatever special sauce is hidden in the black box can be inferred from its outward behavior. Real AI is more brilliant than whatever algorithm you could ever think of so if the real AI can interact w/ your software then it can recreate a much better version of it w/o looking at the source code b/c it has access to whatever knowledge you had while writing the code & then some. I don't think real AI is around the corner but plenty of people believe it is & they also think they only need a few more data centers to make the fiction into a reality.
- HappyPanacea 7mo ago> b/c whatever special sauce is hidden in the black box can be inferred from its outward behavior. This is not always true, for an extreme example see Indistinguishability obfuscation.
- GaggiX 7mo ago>Real AI is more brilliant than whatever algorithm you could ever think of So with "Real AI" you actually mean artificial superintelligence.
- measurablefunc 7mo agoI wrote what I meant & meant what I wrote. You can take up your argument w/ the people who think they're working on AI by adding more data centers & more matrix multiplications to function graphs if you want to argue about marketing terms.
- thangalin 7mo agoTranslate an alternative? https://github.com/albfernandez/juniversalchardet https://github.com/albfernandez/juniversalchardet
- coldtea 7mo ago[flagged]
- the_mitsuhiko 7mo agoMaybe, but the LLM did not recite the chardet source code so that argument does not appear to apply here.
- irishcoffee 7mo agoThis whole "today" fascination with chardet is a classic example of manipulation. I suggest you disregard this term instead of defending it.
- 4star3star 7mo agoI agree. If we look to music, how can a musician unhear what they've heard? We celebrate musicians when they cite their influences. In the case of a software library, it is a tool, not a work of art. Its beauty is in accomplishing a specific, useful task. If we can accept musicians drawing inspiration from all the music they've ever listened to, we should be able to do the same for software, especially when its internal code is unrecognizable from a similar tool.
- coldtea 7mo ago>I agree. If we look to music, how can a musician unhear what they've heard? Unlike with music, in software traditionally a (human) programmer could be chosen who haven't "heard" (i.e. read the original code). That has traditionally called a "clean room" implementation (not to be confused with the software development process called "clean room").
- coldtea 7mo ago[flagged]
- logicprog 7mo agoAlso from that exact same study (why not cite the actual study? It's quite readable) the LLMs couldn't recite more than a small fraction of many other books, often ones just as well known[0] — in fact, from the bar charts shown in the exact news article you cited, it's pretty clear that Sonnet 3.7 was a massive outlier, and so was Harry Potter and the Sorcerer's Stone, so it really seems to me like that's an extremely unrepresentative example, and if all the other LLMs couldn't recite even a small fraction of all the other books except that one outlier pairing, despite them being widely reproduced classics, why would we expect LLMs to actually regurgitate regularly, especially a relatively unknown open source project that probably hasn't been separately reproduced that many times? Not to mention the fact that, as the other commenters mention, that appears to just... not have happened at all in this case, so it's a moot point. [0]: https://arxiv.org/pdf/2601.02671 https://arxiv.org/pdf/2601.02671
- erelong 7mo agohopefully this continues to show how awkward the idea of "intellectual property" (IP) is until people abandon it IP sounds good in theory but enables things like "patent trolling" by large corps and creating all kinds of goofy barriers and arbitrary questions like we're asking about if re-implementations of ideas are "really ours" (maybe they were never anyone's in the first place, outside of legally created mentalities) ideas seem to fundamentally not operate like physical things so asserting they can be considered "property" opens the door for all kinds of absurdities like as pondered in the OP
- AuthAuth 7mo agoI have no data to back this up but patent trolling seems to happen far less than companies that already own significant infra/talent ripping products from smaller companies and out competing them with their scale. I'd rather have patent trolling than have Amazon manufacturer everything i launch. The problem with IP laws and the US is that the big companies already do what IP is suppose to protect and the US refuses to legislate effectively against them.
- galaxyLogic 7mo agoAnd the reason for this is that there is no limit as to how much money corporations can pay for the election campaigns of politicians who make the laws. Right?
- moralestapia 7mo agoIs there anything you have created, spending considerable resources and time, that you ended up giving up for free? For the betterment of humanity? Let's see it!
- NewsaHackO 7mo agoUnfortunately, there are going to be people who push back on the virtue of this being a startup founder website.
- TZubiri 7mo agothe issue with this Stallmanian view on IP is that IP predates software and solves an actual issue. I don't think Stallman has a real proposal to how innovation can be incentivized and compensated. Take the example of medical innovations, sure big pharma is bad, but if they don't get to monetize their inventions, how will R&D get funded? If you destroy IP and allow everyone to clone whatever, you will have a great result in the short term, then no one will continue R&D
- cheesecompiler 7mo ago> I personally think all of this is exciting. I’m a strong supporter of putting things in the open with as little license enforcement as possible. I think society is better off when we share, and I consider the GPL to run against that spirit by restricting what can be done with it. I like sharing too but could permissive only licenses not backfire? GPL emerged in an era where proprietary software ruled and companies weren't incentivized to open source. GPL helped ensure software stayed open which helped it become competitive against the monopoly proprietary giants resting on their laurels. The restriction helped innovation, not the supposedly free market.
- jason_oster 7mo agoYou're putting a lot of responsibility on a license that has several permissive contemporaries. The original BSD license "Net/1" and GPL 1.0 were both published in 1989, while the MIT license has its roots set in "probably 1987" [1] with the release of X11. No doubt, GPL had some influence. But I would hardly single it out as the force that ensured software stayed open. Software stayed open because "information wants to be free" [2], not because some authors wield copyright law like a weapon to be used against corporations. [1]: https://opensource.com/article/19/4/history-mit-license https://opensource.com/article/19/4/history-mit-license [2]: A popular phase based on a fundamental idea that predates software.
- cheesecompiler 7mo agoI’m not saying it’s the only force. But if it wasn’t instrumental what’s your take on the cause of proprietary software dominating until relatively recently?
- randallsquared 7mo agoThe vast majority of running instances of operating systems are Linux or BSD. I don't think proprietary software has dominated for 15-20 years. The two places it has won out thus far is in retail and SaaS. The environment of 1980 when most important software was locked behind proprietary licenses is quite far behind us.
- cheesecompiler 7mo agoAfter cloning a test suite you're still left with ongoing maintenance and development, maintaining feature parity etc. There's a lot more than passing a test suite. If the rewrite is truly superior it deserves to become the new Ship of Theseus. But e.g. I doubt anyone's AI rewrites of SQLite will ever put a dent in its marketshare.
- 7777777phil 7mo agoThe legal question is a distraction. GPL was always enforced by economics: reimplementation had to cost more than compliance. At $1,100 for 94% API coverage, it doesn't. Copyleft was built for a world where clean-room rewrites were painful but they aren't anymore.
- badc0ffee 7mo agoI don't think it's been established that clean-room rewrites are no longer painful. We don't know if chardet could have been rewritten so easily if the original code wasn't in the training set.
- rzerowan 7mo agoStrange this with this whole incident apart from the rewrite/LLM part is the general misundrstanding of the licences. LGPL being a pretty permissive one going as far as allowing one to incorporate it in propriety code without the linking reciprocity clause [1] and MIT is even more permissive. Importantly these were meant to protect the USER of the code.Not the Dev , or the Company or the CLA holder - the USER is primary in the FreeSoftware world.Or at least was supposed to be , OSS muddied the waters and forgetting the old lessons learned when thing were basically bigcorp vs indie hacker trying to getthir electronic device to connect to what they want to connect to and do what they need is why were here. Bikeshedding to eventually come full circle to understand why those decisions were made. In a world where the large OEMs and bigcorps are increasinly locking down firmware , bootloaders , kernels and the internet. I would think a reappraisal of more enforcement that benefits the USER is paramount. Instead we have devs looking to tear down the few user protections FLOSS provides and usher in a locked down hacker unfiendly future. [1] https://licensecheck.io/blog/lgpl-dynamic-linking https://licensecheck.io/blog/lgpl-dynamic-linking
- the_mitsuhiko 7mo ago> Strange this with this whole incident apart from the rewrite/LLM part is the general misundrstanding of the licences. LGPL being a pretty permissive one going as far as allowing one to incorporate it in propriety code without the linking reciprocity clause The short version is that chardet is a dependency of requests which is very popular, and you cannot distribute PyInstaller/PyOxidizer builds with chardet due to how these systems bundle up dependencies. [1]: https://velovix.github.io/post/lgpl-gpl-license-compliance-with-pyinstaller https://velovix.github.io/post/lgpl-gpl-license-compliance-w... [2]: https://github.com/indygreg/PyOxidizer/issues/142 https://github.com/indygreg/PyOxidizer/issues/142
- rzerowan 7mo agoOk thanks for the background on that - again though this would be a painpoint on the packagers - but fully in line with the intentions of the GPL and with the LGPL to enpower the end user to be able to swap/update/tinker as they see fit. As i recall there were some similar situations in regards to licences for distro builders regarding graphicsdrivers and even mp3 decoders wherer there was a song and dance the end user had to go through to legally install them during/after setup. Or better yet to make a truly api compatible re-implementation to use with the license that they want to use, since what they have done i surmise would fall under a derivative work.So they havent really accomplised what they wanted - and instead introduced an unacceptable amount of risk to whoever uses the library going forward. Kinda reminds me of what the Inderner Archive did during the pandemic with the digital lending library.Pushing the boundaries to test them and establish precedence. in any case let see how it plays out.
- STARGA 7mo ago[dead]
- beloch 7mo agoPerhaps code licensing is going to become more similar to music. e.g. Somebody wrote a library, and then you had an LLM implement it in a new language. You didn't come up with the idea for whatever the library does, and you didn't "perform" the new implementation. You're neither writer nor performer, just the person who requested a new performance. You're basically a club owner who hired a band to cover some tunes. There's a lot involved in running a club, just like there's a fair bit involved in operating a LLM, but none of that gives you rights over the "composition". If you want to make money off of that performance, you need to pay the writer and/or satisfy whatever terms and conditions they've made the library available under. IANAL, so I don't even know what species of worms are inside this can I've opened up. It seems sensible, to me, that running somebody else's work through a LLM shouldn't give you something that you can then claim complete control over. --------- Edit: For the sake of this argument, let's pretend we're somewhere with sensible music copyright laws, and not the weird piano-roll derived lunacy that currently exists in the U.S..
- galaxyLogic 7mo agoIf a recordoing is made in a club, doesn't the party doing the recording have the copyright to that (live) recording, or is it the performers?
- jdndbdjsj 7mo agoSans contract? Probably like if I take a photo of you holding a copy of a recent book. I own copy right of the photo. The author still has copyright of the book.
- fpaf 7mo agoI find the music example very illuminating, thanks! Looking into US Copyright for songs there are two different kinds: - one for the composition, the musical idea, music, lyrics. -one for the recording, the music taking shape in a format that someone can listen to I don't think this is how software licenses work, as they cover the code itself, rather than the ideas (the specific recording rather than the composition, in the music example), but it's an interesting way to frame why using LLM this way is, if not illegal, at least unethical. source: https://www.copyright.gov/engage/musicians/ https://www.copyright.gov/engage/musicians/
- PaulDavisThe1st 7mo agoUS courts have ruled that machine generated code cannot be copyright. Ergo, it cannot be licensed (under any license; nobody owns the copyright, thus nobody can "license" it to anyone else). You cannot (*) use LLMs to generate code that you then license, whether that license is GPL, MIT or some proprietary mumbo-jumbo. (*) unless you just lie about this part.
- nl 7mo agoThis oversimplifies it. You can't copyright a work that is only generated by a machine: "In February 2022, the Copyright Office’s Review Board issued a final decision affirming the refusal to register a work claimed to be generated with no human involvement" But human direction of machine processes can be copyright: "A year later, the Office issued a registration for a comic book incorporating AI-generated material." and "In most cases, however, humans will be involved in the creation process, and the work will be copyrightable to the extent that their contributions qualify as authorship. It is axiomatic that ideas or facts themselves are not protectible by copyright law and the Supreme Court has made clear that originality is required, not just time and effort. In Feist Publications, Inc. v. Rural Telephone Service Co., the Court rejected the theory that “sweat of the brow” alone could be sufficient for copyright protection. “To be sure,” the Court further explained, “the requisite level of creativity is extremely low; even a slight amount will suffice." See https://www.copyright.gov/ai/Copyright-and-Artificial-Intelligence-Part-2-Copyrightability-Report.pdf https://www.copyright.gov/ai/Copyright-and-Artificial-Intell...
- galaxyLogic 7mo agoWould writing a prompt, or few, for an LLM qualify as "the requisite level of creativity is extremely low; even a slight amount will suffice"
- nl 7mo agoRead the linked report - it discusses this. The short answer is that it's possible if the prompt has sufficient control but only the parts controlled by the human are eligible for copyright. Using AI doesn't automatically disqualify from copyright protection though.
- __mharrison__ 7mo agoLicensing is done. Reimplementation will be to easy...
- infinitewars 7mo agoThe ship never existed, only the idea of a ship.
- fouc 7mo ago> But this all causes some interesting new developments we are not necessarily ready for. Vercel, for instance, happily re-implemented bash with Clankers but got visibly upset when someone re-implemented Next.js in the same way. Kinda surprised nobody commented on this
- benob 7mo agoIt's funny that real value is now in test suites. Or maybe it's always been...
- mellosouls 7mo agoNote the Ship of Theseus, while a nice comparison for the title, is not - as the author eventually points out - an appropriate analogy here. A fundamental contribution to the idea of whether the identity of the entity persists or not is the continuity between intermediate states. In the example given and discussed here the last couple of days there seems to be a process more akin to having an AI create a cast of the pre-existing work and fill it for the new one.
- senko 7mo agoMaybe, just maybe, this whole AI thing could result in us collectively waking up and realizing copyright is entirely unsuitable for software.
- philipwhiuk 7mo agoOr maybe that AI is committing copyright theft?
- cubefox 7mo ago> Unlike the Ship of Theseus, though, this seems more clear-cut: if you throw away all code and start from scratch, even if the end result behaves the same, it’s a new ship. That's not how copyright works. It doesn't require exact copies. You also can't just rephrase an existing book from scratch when the ideas expressed are essentially the same. Same with music.
- StephenHerlihyy 7mo agoAt what point does the cost of reimplementation shrink below the benefits of obfuscation? Consider a new CVE in Linux. Well maybe my Linux is not the same as the public one. Maybe I just set a swarm of AI agents on making me a drop in replacement that is different but with an identical interface. Same-same but different. Right now writing your own OS to replace the entirety of Linux would be costly and error prone. Foolish. But will it always? What happens when Claude Code Infinute Opus can 1-shot a perfect reimagining in 24 hours? Or 30 minutes? Do all my servers have the same copy or are they all slightly different implementations of the same thing? I dunno.
- fergie 7mo ago> A court still might rule that all AI-generated code is in the public domain, because there was not enough human input in it. That’s quite possible, though probably not very likely. Its not only likely, it is in fact the current position, at least in the US.
- marcus_holmes 7mo ago> For me personally, what is more interesting is that we might not even be able to copyright these creations at all. A court still might rule that all AI-generated code is in the public domain, because there was not enough human input in it. That’s quite possible, though probably not very likely. As I understand it, the US Supreme Court has just this week ruled exactly this. LLM output cannot be copyrighted, so the only part of any piece of software that can be copyrighted is that part that was created by a human. If you vibe-code the entire thing, it's not copyrightable. And if it can't be copyrighted that means it is in the public domain from the instant it was created and can't be licensed.
- BerislavLopac 7mo agoCode is one thing, but what about writing? There is no 100% foolproof way to identify content written by LLMs, and human writing routinely gets incorrectly flagged as such. If I write a book, and a checker says that it's written by LLM, is it automatically in the public domain?
- marcus_holmes 7mo agoReally good question. My understanding is that only human creativity can be copyrighted. So if you sketched out the plot and got the LLM to write all the words, then only the plot is copyrightable. So someone else can copy all the words, as long as they don't copy your plot. However, as you point out, someone has to determine which bits the LLM created and which bits you created. If you wrote the whole book, and a tool incorrectly flags your writing as LLM writing, and then someone copies chunks of your book because they believed the tool and assumed they could (and assuming you filed a DMCA claim and they denied it using the tool's output as proof) then there's going to have to be a court case. I suspect there's going to be a few court cases about this.
- BerislavLopac 7mo ago> only the plot is copyrightable But the plot can't be copyrightable, as the copyright applies only to a tangible representation of an idea (e.g. written text), and not to an idea itself.
- jneen 7mo agoI mean, it has to be asked... was the source of chardet not in the training set...?
- Towaway69 7mo ago> There is an obvious moral question here, but that isn’t necessarily what I’m interested in. Interestingly that‘s also the exact same spot I stopped reading. The dilution of morals weakens societies. We ignore them at our own peril, the planet and most certainly any god figure doesn’t care.
- vbarrielle 7mo agoThe test suite was also licensed under the LGPL. The reimplementation can be seen as a derivative work of the test suite, and thus should fall under the LGPL. This does not even mention the fact that the coding agent, AND the user steering it, both had ample exposure to chardet's source code, making it hard to argue that the reimplementation is a new ship.
- radarsat1 7mo agoThis is interesting because I've been considering a similar project. I maintain a package for a scientific simulation codebase, it's all in Fortran and C++ with too much template code, which takes ages to build and is very error prone, and frankly a pain to maintain with its monstrous CMake spaghetti build system. Furthermore the whole thing would benefit with a rewrite around GPU-based execution, and generally a better separation between the API for specifying the simulation and the execution engine. So I've been thinking of rewriting it in Jax and did an initial experiment to port a few of the main classes to Python using Gemini. It did a fairly good job. I want to continue with it, but I'm also a bit hesitant because this is software that the upstream developers have been working on for 20+ years. The idea of just saying to them "hey look I rewrote this with AI and it's way better now" is not something I would do without giving myself pause for thought. In this case it's not about the license, they already use a permissive one, but just the general principle of suggesting a "replacement" for their work.. if I was doing it by hand it might be different, I don't know, they might appreciate that more, but I have no interest in spending that much time on it. Probably what I will do is just present the PoC and ask if they think it's worth attempting to auto-convert everything, they might be open to it. But yeah, the possibilities of auto-transpiling huge amounts of software for modernization purposes is a really interesting application of AI, amazing to think of all the possibilities. But I'm happy to have read the article because I certainly didn't think about the copyright implications.
- duskdozer 7mo agoIf you really want to do that, the sensible thing is to keep it separate from the original and respect the original license. There would have been no outcry if that happened with chardet. If the different package is genuinely better, it will be used.
- ChrisMarshallNY 7mo ago> slopforks Good term. For myself, I tend to have a similar view as the author (I publish MIT on most of my work), but it’s not really something I’m zealous about, and I’m not really into “slopforking” the work of others. I tend to prefer reinventing the wheel.
- rmoriz 7mo agoI know it's a bit off-topic, but https://www.youtube.com/watch?v=DTYnzLbHUHA https://www.youtube.com/watch?v=DTYnzLbHUHA
- deleted 7mo ago[deleted]
- philipwhiuk 7mo agoMeanwhile elsewhere: https://www.theguardian.com/technology/2026/mar/06/uk-arts-must-not-be-sacrificed-for-speculative-ai-gains-peers-say https://www.theguardian.com/technology/2026/mar/06/uk-arts-m...
- LucasAegis 7mo agoAI is merely a sophisticated tool. If your original thoughts achieve a tangible result through this tool, the ownership should reside with the thinker. Reverse-engineering, in this context, shouldn't be seen merely as an infringement on AI-generated code, but as a violation of the human intellect and systemic design that orchestrated that code. We need to move past protecting 'lines of code' and start protecting the 'intent and architecture' behind them.
- bayindirh 7mo agoWhat if the tool needs an amalgam of everything on the internet to barely function and some of this everything has a big red label saying that adding said thing to this amalgam is forbidden for a reason or another? Further, what if this tool can reproduce these forbidden things almost or completely verbatim and the user of the tool has no way to verify it?
- LucasAegis 7mo agoYou are focusing on the 'bricks' (the literal lines of code), but your argument overlooks the fundamental reality of Architectural Interdependency. In the era of AI-driven synthesis, we must shift our perspective from linguistic expression to systemic logic. Think of software development as finding a structural path from point A to point D. 1.The Foundational Gateway (A → B): You are correct that AI tools are an amalgam of existing data. This foundational layer (A-B) represents the "Prior Art" or the existing IP that serves as a necessary gateway for any further development. If the path starts here, the rights of the original creators must be respected through the established legal framework of Intellectual Property Offices. 2.The Innovative Branch (F → D): However, if an orchestrator uses a tool to forge a new path via a distinct architecture (F) to reach the destination (D), that specific "delta" is a unique intellectual asset. Even if the tool "borrows" the bricks, the topological map of the new architecture belongs to the thinker who directed it. 3.The Necessity of Cross-Licensing: This is where the true core of IP exists. If the owner of the foundation (A-B) wishes to utilize the superior, optimized results of the new path (ABFD), they must respect the IP of the FD architecture. Conversely, the FD creator must acknowledge the base. We aren't just talking about 'verbatim reproduction' of code; we are talking about the Systemic Design that justifies the existence of IP offices worldwide. The future isn't about "cleaning" licenses through AI, but about a more sophisticated world of Cross-Licensing where the foundational layer and the innovative layer recognize each other's functional logic.
- davidcollantes 7mo ago> Right now I would argue that unless some evidence of the contrary could be provided, this can be seen as a new implementation from ground up. Not ship of Theseus, but a "new implementation from ground up. Evidently, the author prefers MIT (https://github.com/chardet/chardet/issues/327#issuecomment-4003822108 https://github.com/chardet/chardet/issues/327#issuecomment-4...), and seems OK with slop-coding.
- latexr 7mo ago> There is an obvious moral question here, but that isn’t necessarily what I’m interested in. And thus we arrive at the absolute shit state the world is in. We keep putting morality aside for something “more interesting” then forget to consider it back in when making the final point. “Have you tried: “kill all the poor?”” https://youtube.com/watch?v=s_4J4uor3JE https://youtube.com/watch?v=s_4J4uor3JE
- globular-toast 7mo ago> The motivation: enabling relicensing from LGPL to MIT. Good heavens, that's incredibly unethical. I suppose I should expect nothing more from a profession that has shied away from ethics essentially since its conception. > I think society is better off when we share Me too. > and I consider the GPL to run against that spirit by restricting what can be done with it. The GPL explicitly allows anyone to do anything with it, apart from not sharing it. You want me to share with you, but you don't want to share with me.
- emporas 7mo agoPorting code from one programming language to another will be one of the most important tasks of code gen A.I. Imagine doing the same with vehicle engines. Less fuel consumption, less pollution, less weight and who knows how many more benefits. Just letting the A.I. do it by itself is sloppy though. The real benefit is derived only when the resulting port is of equal or better quality than the original. It needs a more systematic approach, with a human in the loop and good tools to index and select data from both codebases, the original and the ported one. The tools are not invented yet but we will get there.
- Splinelinus 7mo agoI'm waiting for AGPL to become AIGPL: If you train a model with some or all of the licensed work, you agree that the weights of that model constitute a derivative work, and further for the weights, as well as any inference output produced as a result of those weights to be bound by the terms of the license. If you run a model with the licensed work in part or in full as input, you agree that any output from the model is bound by the terms of the license.
- rzmmm 7mo agoBingo. I can see this is a possible future, and probably desirable scenario for anyone with preference for free software.
- sigmar 7mo agoYou can't change the law with a license agreement and redefine what constitutes a derivative work. If that was possible, people could have done it pre-LLMs. also how would you prove it was in the training set? re: your last sentence, the licensed work wasn't in the input in the chardet example ("no access to the old source tree")
- ncruces 7mo agoAgree. But then, the test suite was the input (chardet). So, is the test suite creative or functional in nature? And does the concept of fair use apply globally?
- glkindlmann 7mo agoSure, a license can't create new legal understanding of "derived work", but I think the intent of what Splinelinus said still works: a license outlines the terms under which a licensee can use the licensed Work. The license can say "if you train a model on the Work, then here are the terms that apply to model or what the model generates". If you accept the license, those terms apply, even if the phrase "derived work" never came up. I hope there are more licenses that include terms explicitly dealing with models trained on the Work. Also, for comparison, both GPL and LGPL, when applied to software libraries (in the C sense of the word), assert that creating an application by linking with the library creates a derived work (derived from the library), and then they both give the terms that govern that "derived work" (which are reciprocal for GPL but not for LGPL). IANAL but I believe those terms are enforceable, even if the thing made by linking with the library does not meet a legal threshold for being a derived work.
- andsoitis 7mo ago> I’m a strong supporter of putting things in the open with as little license enforcement as possible. > © Copyright 2026 by Armin Ronacher. Oooohkaaaay?
- falcor84 7mo agoLicensing, and particularly copyleft is based on copyright - you cannot offer a license, if you don't have a copyright on the thing. You can put it in the public domain, but that is very different.
- andsoitis 7mo agoI understand that. It was just curious to me why, if one holds the position that information ought to be as open as possible, the author still chooses to copyright their won writing. It seems to me, that the ideal is putting it in the public domain, i.e. no copyright. But maybe I'm missing something.
- mannanj 7mo agoI think at the core this is a problem of abuse of the commons and parasitic and extractive behavior being tolerated as a norm. How would I defend myself against hostile entities and societal norms that make it OK to steal from me and my effort without compensation? I will close my doors, put up walls, and distrust more often. That's clearly the trend the world is going towards and I don't see that changing until we find some a way to make it cheaper to detect deception and parasitic behavior along with holding said entities accountable. Since our world leaders have had a history of unaccountable leadership and they are whom model this behavior, I have difficulty seeing the norms change without drastic worldwide leadership change.
- jFriedensreich 7mo agoNon-permissive licenses, open core and proprietary software will just not survive. There is no reality in which I or anyone in my community would use something like eg. raycast or the saas email clients that someone locks down and does rent extraction and top down decisions on. The experience of being able to change anything about the software i use with a prompt while using it is impossible to come back from to all the glitches, limitations and stupidities. we have to come to terms with infinite software.
- StacyRawls 7mo ago[dead]
- bloppe 7mo ago> I’m a strong supporter of putting things in the open with as little license enforcement as possible. I think society is better off when we share, and I consider the GPL to run against that spirit by restricting what can be done with it. This is a head-spinning argument. The whole point of GPL is to force more things out into the open. You'd think someone who espouses open source would cheer the GPL. The only practical difference between MIT and GPL is that the former allows more closed-source code. This feels analogous to the paradox of freedom. Truly unlimited freedom would include the freedom to oppress others, so "freedom maximalism" is an unsound philosophy (unless applied solipsistically). When I publish, I tend to do so under MIT. I also write plenty of closed-source code. And I do generally believe in open source. But I don't use that as a justification for preferring MIT. If anything, I like MIT despite believing in open source, not because. Mainly because I want people to actually use what I wrote.
- casey2 7mo agoPretty simple, if the model was trained on GPL or any copyleft then the output is copyleft (in whole or in part!) you just have a really long preprocessing step before hitting compile.
- just6979 7mo agoI think the reimplementation in question rubs people the wrong way because of the intentions of parties on both ends and the ignoring of one of them by the other (erasure of, from some POV). The original author of the code obviously chose the license they did intentionally (copyleft "keep it open" reasons, seemingly). And the the rewrite author has their intentions as well (unknown beyond "less restrictions on derivative"). The problem comes when those intentions conflict, and in this case the rewrite author basically just ignored the usual convention to resolve the conflict, which is forking or just starting a new project. Claiming "I've maintained it for a while so I can do whatever I want" is kinda gross because is just completely overrides the original authors' intention with their own. They're basically saying "my intentions as maintainer are more important than the creator's", and that doesn't feel even. The "is it a real clean-room" due to prior exposure due to LLM training and working on the codebase is always going to be contentious. But "should I override erase someone else intentions?" question is easy to answer. No. Especially since we have come up with so many ways to make it easy not to (forking is practically free, the abstraction of APIs is powerful, etc). It also just feels a little nefarious. There isn't much reason to change between those licenses in question beyond to allow it to be more tightly integrated into something commercial and closed-source. In which case, having an LLM write a compatible rewrite _in a new project_ seems reasonable at the current moment in time. It's this intentional overriding of the original intentions, seemingly _for profit_ as well, that is the grossest part, because the alternatives are just so easy and common.
- just6979 7mo agoIf Theseus recreated the ship from the original plans but all new parts, created new plans, and then burned the original plans and original parts, it is the same ship? If yes, what if they (with some ship building magic) converted to the second one to have a completely open floor plan inside? Still the same ship?