5 ms·
OpenJDK Interim Policy on Generative AI
- KoleSeise1277 2mo ago[flagged]
- SomeHacker44 2mo agoSeems short sighted for an interim policy, to me. Why not just require clear and detailed disclosure of the use of said tools, to generate data for which to make a properly informed non interim decision?
- pron 2mo agoBecause, as the FAQ section clearly states, the technical and management aspects are not the only ones. There are also legal issues, and until they are clarified or resolved it's best not to let code in that has the potential of causing problems later. We're talking about one of the biggest and most critical open source projects in the world, and unless there's some immediate urgency, it's okay to wait before taking on some risk. I should add that the potential loss here isn't big. For any fix, enhancement, or feature in the JDK, the cost of writing the code is rarely more than 10% of the total cost (which is easy to see if you compare the number of people involved in the project to the rate of code being added or changed), and the interim policy explicitly does allow the use of AI where it helps the most.
- skwirl 2mo agoOracle also has a vested interest in defending an expansive view of intellectual property rights. They lean very heavily on litigation as part of their business model and will want to reserve the right to sue anyone generating software that competes with theirs when the model may have been trained on anything they hold the copyright to. If they held here that it is OK to use LLM generated code for OpenJDK it could kill these cases later as they will have established that they believe it is fair use.
- ToucanLoucan 2mo agoI suspect one of the principle reasons for the blanket decline is specifically that AI generated PRs are already responsible for a substantial uptick in the amount of work required to be done by maintainers. Engaging in discussions about those PRs may be more fulfilling and useful for the developer but it also then increases that work yet again manifold, so that would likely be highly counter-productive. What I don't understand is: if you, for example, are a heavily AI-using dev, who spots a problem in OpenJDK, and you use Claude to figure it out: ok, why can't you then write the solution and PR yourself? Write your own code to replicate the functionality that Claude did, in your own style. Have Claude explain why this fixes the issue (it probably already did) and test it. If it works, write your own explanation and submit it as a PR. I've done this a handful of times for smaller projects, in languages I'm unfamiliar with. Nobody complained because I didn't just copy the shit out of Claude, paste, copy the explanation, paste, and make a PR. I used Claude to increase my understanding, if even a bit and in a necessarily incomplete way, and did it myself.
- grodriguez100 2mo ago> why can't you then write the solution and PR yourself? You can. The policy explicitly allows you to use AI for understanding the codebase, debugging, etc. It just says “no AI generated content”. So if you use Claude to increase your understanding but then write your own code, that’s fine.
- ToucanLoucan 2mo agoMy actual question got a bit lost in my own reply. What I'm really driving at is: Why do people chafe against this?
- wildzzz 2mo agoIt sounds like you did exactly what OpenJDK would have wanted you to do, be the human in the loop.
- pron 2mo agoYou can, and indeed that's what we JDK developers do. But even aside from the fact that writing the code is only a very small portion of the effort in this particular project (you can see that the volume of code making its way into the JDK is very small compared to the number of people involved), I'm often amazed by the gap between how well a frontier model (GPT 5.6 Sol in my case) can comprehend code and investigate a bug, and how badly it writes code and documentation, even when it understands things well. So not only is writing the code not a large portion of the effort in this project, it also happens to be one of the things current models don't do as well as other things.
- CBLT 2mo agoThe interim policy you're proposing seems to have far different priors. The clearest indication of their stance is that they don't even want bug reports that have AI generate part of the text. You don't take a stand like that (imo) unless you've experienced some serious AI-related burnout.
- streetfighter64 2mo ago> short sighted for an interim policy Bit of an oxymoron, no? > just require clear and detailed disclosure of the use of said tools Do you mean disclosure of the use of tools used "privately to help comprehend, debug, and review" contributions? Or do you mean that they should not prohibit AI-generated contributions? If the former I wonder how useful the data would be, and if the latter, they are probably not (yet) willing to allow such contributions because of the mentioned risks: "to reviewer burden, to safety and security, and to intellectual property."
- bpicolo 2mo agoThe copyright issue is an important one. Under the current law, you don’t own copyright on the output of GenAI.
- hajile 2mo agoAll "interim" AI contributions to OpenJDK can and will be sued to make big money off of Oracle if those court cases don't go a certain way. This is ironic given how Oracle has bet almost everything on AI succeeding.
- ofjcihen 2mo agoPast everything. Even the CEO is leveraged up to his neck in personal debt at this point.
- pizlonator 2mo agoHow do you enforce this?
- asdfsa32 2mo agoBlockchain.
- talon8635 2mo agoAlmost spit my coffee out
- pizlonator 2mo agoGenius.
- stonogo 2mo agoAs best as you can. I don't know why people keep asking this question, when the answer is obvious and exactly like all other contribution policies. In this case, this is coming from Oracle, which holds the OpenJDK purse strings, so being found in violation is also likely to have financial consequences for the violator.
- Marha01 2mo ago> I don't know why people keep asking this question Asking how do you enforce a policy that is practically unenforceable is a legitimate question.
- Volundr 2mo agoHow do you enforce the law against murder? Mostly after the fact when people get caught. It's hard to enforce is rarely a good argument against a rule.
- Marha01 2mo agoCatching a murderer is much, *much* easier than catching someone using a modern AI model to generate code, especially if they actually read it and fix occasional LLMisms, and not just vibecode from the hip. The societal damage is also infinitely more serious in the case of murder.
- kdavis 2mo agoWhile I understand the caution, the current policy seems too draconian. It states in part: > Until that policy is in place, the Governing Board has approved this interim policy: > Contributions in the OpenJDK Community must not include content generated, in part or in full, by large language models... Note, this would exclude most spell checkers, as they often are LLM based. That said, they do soften this with the addition: > Q: Is it okay to continue using the spell-checking, grammar-checking, auto-completion, and refactoring features in my editor or IDE? > A: Yes, so long as they are not based on large language models or similar deep-learning systems. This addition will likely not help as most spell checkers are OS/IDE services and their implementation details are opaque to users.
- malfist 2mo ago> this would exclude most spell checkers, as they often are LLM based Who is using an LLM as a spell checker?
- bpicolo 2mo agoYeah - spell checking has been on device for literal decades.
- kdavis 2mo agoMy point is you might be when you wrote your comment or I might be in writing this comment. However, I don't know if I am as I haven't seen, and will likely never see, the implementation behind the text box that is suggesting the next letter, word, or spelling correction as I type.
- yjftsjthsd-h 2mo agoCan you point at a specific spell checker that is known to use an LLM?
- kdavis 2mo ago"A Simple yet Effective Training-free Prompt-free Approach to Chinese Spelling Correction Based on Large Language Models"[1] [1] https://aclanthology.org/2024.emnlp-main.966/ https://aclanthology.org/2024.emnlp-main.966/
- effnorwood 2mo agoOracle determining how to add more sypware while not violating their own TOS
- pron 2mo agoWhat spyware is there in the JDK? It's open-source, so you can look for yourself.
- abc42 2mo agoMy prophecy is that in 3 years we'll see a complete reversal of this. Using GenAI to code will be the default and we'll see policies that put limits on human/artisan development. Possibly even projects that outright ban non-LLM development.
- ACCount37 2mo agoAt this point, I wouldn't be too surprised if LLMs just get better at "write understandable, maintainable code" than most developers are. For now, LLMs are still bad at it, but they're already at "competitive with humans" tier of bad. I expect them to get better.
- realusername 2mo agoI don't think they ever will on the "maintainable" part until their absolutely atrocious memory gets improved at least tenfold. (and no, the tricks we deployed until now aren't enough) They already are better than most humans at writing code for sure but even the best SOTA model is worse than your average intern at memorization.
- ACCount37 2mo agoThat's the entire point of the "maintainable" part: someone with no "memory" of what was written where and why should be able to walk into the codebase, understand how it works, and be able to make a sensible change to it without setting things on fire. Someone with "memory" should be able to do the same faster, but having that memory should not be a requirement. Whether that "someone" is a freshly onboarded mid-level SWE or a mid-tier coding LLM is largely irrelevant.
- realusername 2mo agoDepends what you mean by your comment then, if you want to say that you can direct LLMs to write maintainable code yes sure you can through very strict guardrails, if they can do it themselves, no they can't and need help because of their bad memory.
- Viliam1234 2mo agoThis makes sense. AI contribution is basically just "prompt + AI work". Even if you are okay with AI work per se, you should accept prompts (after reviewing them) and let your own AI generate the code (and then also review the code)... rather then accept an output of someone else's AI with an unknown prompt, that may or may not include an instruction to create a vulnerability. In the age of AI, the prompt is becoming the actual source code. Accepting AI-generated code would be like accepting binary code from unknown source.
- this_user 2mo agoNot really though, since the result of the prompt is not deterministic. It greatly depends on the model, the version, the harness, even time of day if the provider's infrastructure is currently overloaded and is silently degrading performance. Some things also require multi-turn interactions.
- kylemart 2mo ago100% agree with this. Was about to write the same thing.
- doctorpangloss 2mo agothis is the unintentional HN humor i come here for
- lelanthran 2mo ago> Not really though, since the result of the prompt is not deterministic. It greatly depends on the model, the version, the harness, even time of day if the provider's infrastructure is currently overloaded and is silently degrading performance. Some things also require multi-turn interactions. All of that are even more reasons to reject AI-generated code: if it cannot be trusted to produce same (or even similar) output just from the prompts, why accept it at all?
- rat9988 2mo ago
- 123OnenO123 2mo ago[dead]
- mrkeen 2mo agoAm I not even allowed to use https://www.oracle.com/artificial-intelligence/enterprise-ai/ https://www.oracle.com/artificial-intelligence/enterprise-ai... ?
- piinbinary 2mo ago> Most generative AI tools, however, are trained on copyrighted and licensed content, and their output can include content that infringes those copyrights and licenses To some extent, it feels like the genie is out of the bottle on this. There's so much LLM-generated code out there, and I'm sure plenty of it could be argued to infringe a copyright or license (though I think the legal bar for counting as infringement is set too low), that there's no way to go back and undo it. That said, OpenJDK might be afraid that someone will decide to make an example of them because they are a high-profile target.