4 ms·
It is enforceable, I think you mean to say that it cannot be prevented since people can attempt to hide their usage? Most rules and laws are like that, you pros
by BlackFly 7mo ago
It is enforceable, I think you mean to say that it cannot be prevented since people can attempt to hide their usage? Most rules and laws are like that, you proscribe some behavior but that doesn't prevent people from doing it. Therefore you typically need to also define punishments:
> This policy is not open to discussion, any content submitted that is clearly labelled as LLM-generated (including issues, merge requests, and merge request descriptions) will be immediately closed, and any attempt to bypass this policy will result in a ban from the project.
- hparadiz 7mo agoWhat happens when the PR is clear, reasonable, short, checked by a human, and clearly fixes, implements, or otherwise improves the code base and has no alternative implementation that is reasonably different from the initially presented version?
- pm215 7mo agoIf you're going to set a firm "no AI" policy, then my inclination would be to treat that kind of PR in the same way the US legal system does evidence obtained illegally: you say "sorry, no, we told you the rules and so you've wasted effort -- we will not take this even if it is good and perhaps the only sensible implementation". Perhaps somebody else will eventually re-implement it later without looking at the AI PR.
- hparadiz 7mo agoHow funny would it be if the path to actually implement that thing is then cut off because of a PR that was submitted with the exact same patch. I'm honestly sitting here grinning at the absurdity demonstrated here. Some things can only be done a certain way. Especially when you're working with 3rd party libraries and APIs. The name of the function is the name of the function. There's no walking around it.
- pm215 7mo agoThat's why I said "somebody else, without looking at it". Clean-room reimplementation, if you like. The functionality is not forever unimplementable, it is only not implementable by merging this AI-generated PR. It's similar to how I can't implement a feature by copying-and-pasting the obvious code from some commercially licensed project. But somebody else could write basically the same thing independently without knowing about the proprietary-license code, and that would be fine.
- ranger_danger 7mo agoThe trick is getting people to believe you.
- joaohaas 7mo agoIt follows the same reasoning as when someone purposefully copies code from a codebase into another where the license doesn't allow. Yes it might be the only viable solution, and most likely no one will ever know you copied it, but if you get found out most maintainers will not merge your PR.
- pmarreck 7mo agoYou not realizing how ridiculous this is, is exactly why half of all devs are about to get left behind. Like, this should be enshrined as the quintessential “they simply, obstinately, perilously, refused to get it” moment. Shortly, no one is going to care about anyone’s bespoke manual keyboard entry of code if it takes 10 times as long to produce the same functionality with imperceptibly less error.
- bigstrat2003 7mo ago> Shortly, no one is going to care about anyone’s bespoke manual keyboard entry of code if it takes 10 times as long to produce the same functionality with imperceptibly less error. Well that day doesn't appear to be coming any time soon. Even after years of supposed improvements, LLMs make mistakes so frequently that you can't trust anything they put out, which completely negates any time savings from not writing the code.
- pmarreck 7mo agoSorry, but this is user error. 1) Most people still don't use TDD, which absolutely solves much of this. 2) Most poople end up leaning too heavily on the LLM, which, well, blows up in their face. 3) Most people don't follow best practices or designs, which the LLM absolutely does NOT know about NOR does it default to. 4) Most people ask it to do too much and then get disappointed when it screws up. Perfect example: > you can't trust anything they put out Yeah, that screams "missing TDD that you vetted" to me. I have yet to see it not try to pass a test correctly that I've vetted (at least in the past 2 months) Learn how to be a good dev first.
- notpachet 7mo ago> no one is going to care about anyone’s bespoke manual keyboard entry of code if it takes 10 times as long to produce the same functionality with imperceptibly less error. No one is going to care about anyone’s painstaking avoidance of chlorofluorocarbons if it takes ten times as long to style your hair with imperceptibly less ozone hole damage.
- 7mo ago
- pmarreck 7mo agoThis is where most reasonable people would say “OK, fine” CLEARLY, a lot of developers are not reasonable
- eschaton 7mo agoIt is entirely reasonable for a project to require you to attest that the thing you are contributing is your own work. The unreasonable ones are the ones with the oppositional-defiant “You can’t tell me I can’t use an LLM!” reaction.
- pmarreck 7mo agoIt IS their own work. The simplest refutation of your point of view is, who or what is responsible if the work submission is wrong? It will always be the person’s, never the computer’s. Conveniently, AI always acts as if it has no skin in the game… because it literally and figuratively doesn’t… so for people to treat it like it does, should be penalized
- eschaton 7mo agoIf it’s the output of an LLM, it’s not their own work.
- pmarreck 7mo agoWho prompted the LLM? Who vetted the output? Who ensured there was adequate test coverage? Who insisted on a certain design? Who is to blame if it's bad code? That is the same entity that is responsible, and the same entity that "did it" tl;dr your stance is full of poop, my dude
- eschaton 7mo ago“I looked up the topic on Wikipedia and I highlighted the text and I selected copy and I selected paste so I don’t see how this is plagiarism.” That’s what you sound like.
- ralferoo 7mo agoThe problem is that even if the code is clear and easy to understand AND it fixes a problem, it still might not be suitable as a pull request. Perhaps it changes the code in a way that would complicate other work in progress or planned and wouldn't just be a simple merge. Perhaps it creates a vulnerability somewhere else or additional cognitive load to understand the change. Perhaps it adds a feature the project maintainer specifically doesn't want to add. Perhaps it just simply takes up too much of their time to look at. There are plenty of good reasons why somebody might not want your PR, independent of how good or useful to you your change is.
- pjc50 7mo agoHow would you tell that it's LLM-generated in that case? If the submitter is prepared to explain the code and vouch for its quality then that might reasonably fall under "don't ask, don't tell". However, if LLM output is either (a) uncopyrightable or (b) considered a derivative work of the source that was used to train the model, then you have a legal problem. And the legal system does care about invisible "bit colour".
- hparadiz 7mo agoIt's (c) copyright of the operator. For one simple reason. Intention. Here's some code for example: https://i.imgur.com/dp0QHBp.png https://i.imgur.com/dp0QHBp.png Both sides written by an LLM. Both sides written based on my explicit prompts explaining exactly how I want it to behave, then testing, retesting, and generally doing all the normal software eng due diligence necessary for basic QA. Sometimes the prompts are explicitly "change this variable name" and it ends up changing 2 lines of code no different from a find/replace. Also I'm watching it reason in real time by running terminal commands to probe runtime data and extrapolate the right code. I've already seen it fix basic bugs because an RFC wasn't adhered to perfectly. Even leaving a nice comment explaining why we're ignoring the RFC in that one spot. Eventually these arguments are kinda exhausting. People will use it to build stuff and the stuff they build ends up retraining it so we're already hundreds of generations deep on the retraining already and talking about licenses at this point feels absurd to me.
- rswail 7mo agoI think you need to read the report from the US Copyright office that specifically says that it's *not* (c) copyright of the operator. It doesn't matter if the "change this variable name" instruction ends up with the same result as a human operator using a text editor. There is a big difference between "change this variable name" and "refactor this code base to extract a singleton".
- hparadiz 7mo agoYou may as well be the MPAA right now throwing threats around sharing MP3s. We're past the point of caring and the laws will catch up with reality eventually. The US copyright office says things that get turned over in court all the time.
- BlackFly 7mo agoYes, what happens when the murder looks like a heart attack? This isn't hypothetical, some assassinations occur like this. That doesn't make murder laws unenforceable. Lots of people try to get away with perfect crimes and sometimes do. That doesn't make the rule unenforceable, it just highlights the limits of human knowledge in the face of a dishonest person. Hence the escalations for trying to destroy evidence of crimes or in this case to work around the AI policy. Here, instead of just closing your PR, they ban you if you try to hide it.
- repelsteeltje 7mo agoI think the bigger point about enforcement is not whether you're able to detect "content submitted that is clearly labelled as LLM-generated", but that banning presumes you can identify the origin. Ie.: any individual contributor must be known to have (at most) one identity. Once identity is guaranteed, privileges basically come down to reputation — which in this case is a binary "you're okay until we detect content that is clearly labelled as LLM-generated". [Added] Note that identity (especially avoiding duplicate identity) is not easily solved.
- khalic 7mo agoUnenforceable means they can't actually enforce it since they can't discriminate high quality LLM code from hand typed
- eschaton 7mo ago[flagged]
- khalic 7mo agoKeep wishing, in the meantime some people have to deal with the real world and plan accordingly
- BlackFly 7mo agoWell, unenforceable isn't a synonym for undetectable or awkward. Their policy indicates that they are aware of this difficulty: if you admit to using AI then they close your pull request, if you do not admit to using AI but evidence later surfaces that you did then they ban you. They can enforce this. The hope here is the same hope as most laws: that lies eventually catch up to people. That truth comes to light. But sure, in the meanwhile, there are always dishonest people around trying to flout rules to varying degrees of success. Some are caught right away, some live their entire lives without it catching up to them. That doesn't make the rule unenforceable, that just highlights the limits of rules: it requires evidence that can be hard to come by.
- hrmtst93837 7mo ago[flagged]
- eschaton 7mo ago[flagged]