9 ms·
Constraint Decay: The Fragility of LLM Agents in Back End Code Generation
- ElenaDaibunny 4mo agoFragility compounds fast when you add visual grounding to the loop. Code agents at least get structured feedback from the compiler.
- volume_tech 4mo ago[flagged]
- maxbond 4mo agoReminds me of the recent paper about delegating document editing tasks to LLMs across different disciplines [1]. That paper found that programming was the only discipline most LLMs can perform long horizon tasks on without accumulating errors & corrupting the document. I've only read the abstract of this one so far but it seems like this paper has zoomed in on programming with greater fidelity and shown a similar phenomenon. But not about long horizon tasks, more like "long style horizons" of larger sets of structural constraints. [1] https://arxiv.org/abs/2604.15597 https://arxiv.org/abs/2604.15597 Discussion: https://news.ycombinator.com/item?id=48073246 https://news.ycombinator.com/item?id=48073246
- emp17344 4mo agoIf it’s not easily verifiable, LLMs aren’t good at it.
- jeremyjh 4mo agoI think that’s mostly because they get so much more of that reinforcement learning - since it is so economical. I dont know if there is any evidence of a fundamental reason they can’t be just as good at other tasks, but it might be economically infeasible for awhile yet.
- mjburgess 4mo agoNo one is curating vast amounts of data for them in other domains. Programmers send programs with fixes
- knollimar 4mo agoThere's no diff of my excel lambdas being fixed? :(
- jeremyjh 4mo agoIts more about how costly it is to verify work in reinforcement learning. It is cheap in Mathematics and coding because it can be automated. It is expensive in other domains because while you can capture certain datasets to do pre-training on, you ultimately need humans in the loop to judge the quality of work.
- emp17344 4mo agoRLVR doesn’t work for unverifiable tasks, so they won’t be able to effectively use tools to boost reliability for those tasks.
- jeremyjh 4mo agoRight, so you have to use RLHF. That is the economics problem I was referring to.
- dominotw 4mo agobut what does it mean to be good at something that cant be verified. how do you know that they are not good at it, you are obviously using some measure. sounds like an oxymoron of a claim.
- maxbond 4mo agoIt means having taste. People say Picasso was a great painter, but that cannot be verified (at least, not in the sense of a verified reward).
- dominotw 4mo ago"people say picasso was a great painter" is definitely not hard to verify . lol.
- deleted 4mo ago[deleted]
- maxbond 4mo agoI don't know if you're being factitious or not but that was not what I meant. Picasso being a great painter is an example of "having taste"; "create an artistic image generation model with Picasso-level performance" is a valid problem statement we could attack with RLHF, but not with RLVR, because "taste" is not amenable to modeling with a reward function. "Write this code in a way that is readable and maintainable" is another example.
- dominotw 4mo agohttps://futurism.com/artificial-intelligence/real-monet-ai-chaos https://futurism.com/artificial-intelligence/real-monet-ai-c...
- deleted 4mo ago[deleted]
- 4mo ago
- jdlshore 4mo ago“Our systematic study exposes a phenomenon of constraint decay in LLM-based coding agents. While current models excel at unconstrained generation, their performance drops when forced to navigate explicit architectural rules. For end-users, this dichotomy implies that agents are reliable for rapid prototyping but remain unreliable for production-grade backend development.” One major weakness of this study is that they didn’t fully test frontier models for cost reasons, so the specific performance results should be taken with a grain of salt. But the overall conclusion that models degrade when both behavior and architecture must be correct is interesting, and something to keep an eye on.
- qsort 4mo agoI think it's downstream of "you can't optimize for two different objectives". If you only have functional requirements, then in effect you're doing some form of program synthesis, and RL can optimize that very hard. If you have a mixture of functional and non-functional requirements, you are basically giving the model an incomplete specification, and it must in some way guess at the user's intent to fill in the blanks. This is also why adding to the prompt examples of the style of code you want (hats off to antirez for this particular tip ;)) is phenomenally powerful.
- apsurd 4mo agoWould you mind sharing antirez' suggestion?
- qsort 4mo agoI am obviously paraphrasing, but the general idea is that trying to synthesize style from a codebase into e.g. a markdown guide generally doesn't work very well. What achieves style transfer is providing the model with a lot of examples of the style, conventions, patterns you want. To put it in practice: if you point claude/codex to a repository and you ask it to implement feature X using style guide Y, the code will probably work, but you can usually get better results by saying "do it in the style of this file, it was done well there".
- gkfasdfasdf 4mo agoOdd they used GPT-5.2 and not GPT-5.2-codex. i.e. the one optimized for coding agent tasks.
- maleldil 4mo agoConsidering this is from academia, there's a chance there were limitations on the available models. My research group accesses OpenAI models via Azure, and until recently (last week) the latest model was GPT 5. We just got 5.4.
- beering 4mo agoThat’s wild. Are you at a university that bans using the OpenAI APIs directly?
- deleted 4mo ago[deleted]
- maleldil 4mo agoThe university doesn't « ban » using the OpenAI APIs directly. It's a question of funding. If you want to use OpenAI, you usually use your own account and ask the university for a refund later, where you justify your usage. It's easier for the university if you use their pre-approved Azure endpoint instead, though you'll still need approval if you're going to spend a significant amount of money.
- yomismoaqui 4mo agoAlso they used languages with dynamic typing like Python & JS. In my experience a statically typed codebase is easier to maintain for humans so maybe it is also for agents. When using Codex/Claude Code with Go code I cannot count the times the agent does some change, runs a build to check for errors, find some and fix them.
- acbart 4mo agoIt's crazy to me that people think of Python as dynamically typed by default. Strong static typing has been an option in Python for years now, and it should just be the default.
- epgui 4mo agoThe python type hints are useful for static analysis (and yes, should be the default) but it’s a joke compared to the utility of types in a language like Haskell.
- shepherdjerred 4mo agoIf you're comparing type systems against Haskell you're excluding all mainstream languages except maybe Scala and Rust
- epgui 4mo agoYes.
- mrob 4mo ago>Strong static typing has been an option in Python for years now, and it should just be the default. https://docs.python.org/3/library/typing.html https://docs.python.org/3/library/typing.html "The Python runtime does not enforce function and variable type annotations. They can be used by third party tools such as type checkers, IDEs, linters, etc." Which third-party enforcement mechanism do you propose become the default?
- leecommamichael 4mo agoThese things don’t think. We’re going to have to reiterate this for a long time, I fear.
- emp17344 4mo agoThere is now a trillion-dollar industry bent to the task of convincing people these things can think. It’s gonna cause some damage.
- suprfnk 4mo agoI don't think they think. I still use them a lot despite that, because they are very powerful parameterised code generators.
- sheeshkebab 4mo ago…but they reason well enough given enough context (using their matmuls).
- noosphr 4mo agoTo this day frontier models think that A and not B means A and B when the sentence gets pushed far enough back in their context window. The context length that model can reason over without obvious errors is much smaller than the advertised context. Between a 1/4th to a 1/20th what is advertised on the tin.
- Npovview 4mo agoDo you also happen to remember what you ate last thrusday?
- leecommamichael 4mo agoIs that the same gap as what you’re responding to? To me, it seems his critique is about advertised capability and logical statements, and your rhetorical(?) question is about memory.
- p0w3n3d 4mo agotasks spanning eight web frameworks Does anyone else have this experience that LLM create better pure html+CSS+js than work with existing frameworks?
- bob1029 4mo agoI think web frameworks have been "in trouble" as of gpt-5.4. I can't imagine using something like React anymore. The most incredible combo I've seen lately is progressive enhancement of Razor Pages with javascript. With this arrangement the newest models tend to make a really good call on if something should happen server-side (cshtml) or on the client (js).
- p0w3n3d 4mo agoI've recently vibecoded pure html+css+js frontend for WWTBM-alike game: https://github.com/pawel-jaworski-loftyworks/mili-game https://github.com/pawel-jaworski-loftyworks/mili-game - it consists of one file and is blazingly fast. Previous attempts to vibecode something with a vite-framework something were more harsh and clumsy.
- rbbydotdev 4mo agoThis is interesting, anecdotally I have felt like I was having better luck with raw sqlite than using an ORM in a recent typescript project, using raw sqlite queries vs drizzle
- bob1029 4mo ago> Our findings reveal a phenomenon of constraint decay: as structural requirements accumulate, agent performance exhibits a substantial decline. I have exactly the inverse findings on my end. The bigger and more legacy the codebase, the more accurate the patches become. The harness itself seems to be the most important part. I use a recursive loop that primes the root context based on the user prompt each time. My agent will often make over 100 tool calls to sql and git before it finally decides to apply a patch. If I was greenfield, there would be nothing to query or constrain against.
- richardlblair 4mo agoI find the same. We have abstractions with multiple concrete implementations, examples of patterns and examples of anti patterns. I usually find I can achieve 90% of the outcome I'm trying to achieve. I use sonnet for planning, qwen for coding, sonnet for review.
- xcjsam 4mo agoThe harness mattering more than the model lines up with my experience too. What this paper measures is within-turn constraint decay. The version that bites in multi-agent setups is across-session — the architectural rules an agent wrote down on Monday don't reach the agent making the next change on Tuesday.
- haeseong 4mo ago[flagged]
- dwa3592 4mo agoThis sounds like another version of "As a chat becomes longer, the guardrails seem to become fuzzy". You can't use all of the context window bc at the end, the output would not respect the constraints (or guardrails) but to reliably produce production grade code you want the model to have expansive awareness which fills up the context window pretty quickly. It's like saying "Keep everything in mind from these 6 directories - and make this <insert ticket> change" - but keeping everything in mind already fills it's context window which makes it lose it's ability to follow the constraints (or guardrails).
- whatever1 4mo agoThis is not a new problem though. This is why we started writing modular code, strict interfaces etc
- lanstin 4mo agoAnd doing incremental dev, so once a feature is done you can mostly ignore it.
- Silhouette 4mo agoIf there is one good thing that the generative AI tools have shown beyond any doubt it's that the classic "good programming" practices are still useful and effective. Self-documenting code. Modular design. Clearly defined architecture. Incremental development. Coding standards. Automated tests. Automated everything. If there's a second thing the generative AI tools have shown beyond any doubt it's that many of the more modern (relatively speaking) "best practices" that have always been over-hyped and questionably-evidenced really do tend to produce worse results. LLMs take these methods to their logical conclusions and show us the end result much sooner. You can't just iterate your way to a solution when you don't even know what problem you're trying to solve. If you don't have a clear spec then you don't know what a correct product looks like. You need to invest time in reviewing code properly. If you don't keep the big picture in mind then the big picture becomes a mess. Maybe one day the LLMs will leave me out of a job but at least I'll feel validated first!
- phrotoma 4mo ago"constraint decay" isn't this just another name for the (already well understood) idea of "context rot"?
- oulipo2 4mo agoExactly why you can't remove humans in the loop to assess that the solution is not only correct (which LLMs are quite bad at, once concurrency, logic, etc are involved), but also elegant, maintainable, etc
- spacedoutman 4mo agoThis research is useless and nearly all other LLM research is too. gpt 5.2 is the strongest model they tested, a nearly 6 month old model. Traditional research can not keep up.
- acgourley 4mo agoI disagree, their findings should generalize to the frontier. Even if the latest can deal with the extra complexity, it stands to reason it will take more tokens to do less. This could be a useful insight into the next generation of evals.
- abujazar 4mo agoAgreed. As Simon Willison points out, November 2025 was a a critical months because that's pretty much when coding agents became «good enough», eliminating most of the problems pointed out in this study.
- sanxiyn 4mo agoGPT-5.2 was released after November 2025.
- anygivnthursday 4mo agoI regularly see Claude Opus 4.7 dropping constraints from an otherwise small CLAUDE.md at merely 20% context use. I have to keep reminding it, and it has all info ready in its context, still time to time decides to ignore parts.
- vishvananda 4mo agoI've been experimenting quite a bit with long-horizion agentic coding[1] and I have also noticed that agents seem to perform worse when forced into certain architectural patterns. I have found that is a bit better when including the constraints along the way instead of adding them after the fact. There seems to be a side-effect I have been calling "calcification", where a pattern starts appearing in the codebase and the agent follows the pattern to the point where it dominates the context and becomes self-reinforcing. This could potentially be a strength or a weakness for existing code bases depending the codebase quality. I will have more insights on this soon as more from-scratch runs conclude that include architectural guidance from the beginning. [1]: https://medium.com/@vishvananda/i-spent-2-billion-tokens-writing-a-c-compiler-so-you-dont-have-to-d3e4eec4781e https://medium.com/@vishvananda/i-spent-2-billion-tokens-wri...
- jumploops 4mo ago> agents seem to perform worse when forced into certain architectural patterns. FWIW I've noticed this too. I've found that the agents/models have their own style, which is mostly summed up as overly verbose. Additionally, the models are OK at modularization when given space to "plan" their implementation, but rarely decide that abstracting something would be helpful after the fact (i.e. after many iterations on a greenfield codebase or when being dropped into a legacy codebase). This often leads to "god files" which, when pointed to by the user/architect, causes the models to correctly critique (humorously when they're the ones that wrote the code in the first place).
- rrook 4mo agoAs a codebase grows, divergent structural emergence from incidental(lang and lib) details results in prolonged complexity costs. I'm working on a language that enforces structure for agents: https://github.com/hale-lang/hale https://github.com/hale-lang/hale
- wetpaws 4mo ago[dead]
- pianopatrick 4mo agoI think someone is going to figure out a framework for using LLMs for coding. A framework would use static code checking tools to force an architecture on to LLMs instead of trying to do so in markdown. I don't know exactly what it will look like but for example I could imagine a Java Framework where the LLM could only create subclasses of certain classes.
- cheevly 4mo agoA lot of us have been doing this for over a year now.
- deleted 4mo ago[deleted]
- AmazingTurtle 4mo agoSo my finding is: planning is worth it. For a little complex changes, I always run codex (5.5-high) in planning mode first. I have linked various docs/{ARCHITECTURE,BACKEND-GUIDELINES,NESTJS-DI,..}.md etc. from AGENTS.md so they can quickly discover relevant docs at planning time, only if they are needed. No need to know react specific stuff when it's dealing with a backend problem for example. I typically blindly approve plans made by the agent with a fresh context, because that's as if I had prompted it. Works the best for me. Using /goal however, it's really just constantly compacting and doing it's thing, of course it gets sloppy. If only there was a state machine that would transform tickets into a Planning Mode Prompt, then use, idk. guardian approvals (somehow a "Product Management Perspective Lens" approving or making changes to the plan) and then letting a less capable or less reasoning agent execute the plan, I think that would work the best.
- Geezus_42 4mo agoSounds like you want something like this. https://github.com/tremtec/maestro https://github.com/tremtec/maestro
- siliconc0w 4mo agoI recommend spending some time getting a few parts of the codebase idiomatic and then @-ing those files as exemplars. This works a lot better than trying to steer it with markdown. This works reasonably well for like FastAPI but JavaScript seems to be the worst, even with guidance and exemplars it'll prefer in-lining a bunch of garbage rather than use the APIs as directed.
- KronisLV 4mo ago> For end-users, this dichotomy implies that agents are reliable for rapid prototyping but remain unreliable for production-grade backend development. Time to start writing linting tools that check the architecture and spoon feed the LLM what exactly it's doing wrong. I reckon something like this would be good for every project out there: https://www.archunit.org/getting-started https://www.archunit.org/getting-started They expand a bit more on the reasoning behind it: https://www.archunit.org/motivation https://www.archunit.org/motivation (I also wrote a simple linter for architecture/code checks that aren't well encapsulated by ones that just focus on individual files, that uses Go + goja to write rules in ECMAScript and parallelize the read only ones and also allow ones that change files as necessary, in addition to something like Ruff / Oxlint / Oxfmt / whatever is present in each stack; though it's is still in development and not as good of a focused example as ArchUnit is) If we write software specification docs, bother describing how it evolves with ADRs, enforce code style automatically and require certain test coverage automatically (or at least should), why couldn't we go a step further, formalize those specs and ensure that any new code is also up to snuff? I don't think that's any more of a job for an LLM, than telling it how it should format code is. Also, I'm in the camp that believes that at least many of your ORM mappings and similar stuff should be the output of codegen, since you've already gone through the trouble of describing the schema/migrations to get there. I don't think this would be only good for LLMs, though - I've seen projects that have like 3 different audit systems built in, not because of some fancy business requirement, but rather cause the devs either didn't know about the previous one(s) or just didn't feel like following what should have been the pre-established conventions, even when there were docs in place (nobody read those).
- try-working 4mo ago[flagged]
- dalemhurley 4mo agoThis is why we as an industry have spent so much effort optimising the code generation process with things like skills, rules, tests, reviews, lints, agentic loops with feedback and sub-agents, and the code-runners. It is not just LLMs building code, it is an eco-system collaborating together. I would agree too that as the codebase grows the LLM struggles more and more with generating code. It is probably misaligned incentives, it wants to complete the isolated task without too much context consumed, at the POC it can consume most of the app, by about 30K lines of code it is quite complex code base to navigate.
- deleted 4mo ago[deleted]
- hottrends 4mo ago[flagged]
- alasano 4mo agoI've been building https://engine.build https://engine.build to introduce a proper structured external agent orchestrator that's used to build with clear constraints and make sure the end result is what you wrote in your spec or requirements. Without having to babysit and micromanage the models. Implementation phases very often go through 5-10 review and fix rounds to actually get the implementation to match the spec. It takes longer but that's what's necessary to get actually good results on long horizon tasks with detailed requirements. I'll be open sourcing it fully soon.
- zane_shu 4mo ago[flagged]
- guhcampos 4mo agoI'm a convert. I was 100% skeptical about LLM code generation, now over 80% of the professional code I write is generated. That said, the limitations are kind of obvious and are starting to show in some of my projects, and this article seems to confirm my suspicions. If it's just confirmation bias or not, I can't say yet. In my experience, for anything complex enough, I have to start adding more and more constraints, style guides, corner cases, error handling, optimization guidelines and all this good stuff to my Markdown specifications, rules and skills. At some point this starts to look like we're all just moving complexity from the more formal and deterministic world of programming languages to the informal and non-deterministic world of natural language. The writing speed gains are enormous, yeah, and business sees this as productivity gains, of course - and we do it because the pressure for increased productivity is there, as it's always been; yet the trade off seems to be clear and a lot of people are just ignoring it.
- runhelm 4mo ago[flagged]
- dominotw 4mo ago[flagged]
- apsec112 4mo agoLLMs recently solved a major, famous open mathematical problem in combinatorial geometry: https://www.reddit.com/r/math/comments/1tj534d/openais_internal_model_disproves_unit_distance/ https://www.reddit.com/r/math/comments/1tj534d/openais_inter...
- loeg 4mo agoThere is nothing new under the sun.
- son_of_gloin 4mo agoBut there are other suns :)
- alexwwang 4mo agoI am trying to avoid this by building a plugin based on my memory management project Aristotle. I add a status machine to monitor the activities of LLM while it does jobs following my tdd-pipeline skills, which begins with requirements clarification and ends up with delivery. These two projects are on GitHub, you may search alexwwang/aristotle and alexwwang/tdd-pipeline to dive into the details or just ask your LLM to scan them to tell you the points you are interested in.
- dundunUp 4mo ago[flagged]
- brentrhodes 4mo ago[flagged]
- Developer_H 4mo ago[flagged]
- codepack 4mo ago[flagged]
- launchseed 4mo ago[flagged]
- luodaint 4mo agoNot those carefully designed constraints that I set up from the beginning, but short-term ones that I came up with after an agent failed in some way: "Validate JWT at the route level, not the component." "Call workspace provisioning on each user creation." Both because of things the agent had done incorrectly. Aspiration vs. consequence, in other words. An aspiration constraint describes a desired outcome for the system; a consequence constraint maps to a problem already encountered. And the agent ignores the former when faced with the path of least resistance while obeying the latter because it is brief, unambiguous, and precise about preventing that particular failure mode. Which is key rather than the harness in determining survival through session rotation.
- zenai666 4mo ago[flagged]
- jixter_apps 4mo ago[flagged]
- lemax 4mo agoWould love to see this benchmark tested on more perceivably LLM friendly frameworks/ORM (e.g. is NestJS or Drizzle / Kysely more performant than their choice of Sequelize) and more frontier model vs just GPT 5.2. Anyone read whether these tests include any validation loops? What happens if the models get back test failures, for instance? Understanding how many turns to hit full passing behavior suite would also be interesting. Great methodology in the study though.
- GhostGains 4mo ago[flagged]
- huaiorg 4mo ago[flagged]
- fredcallagan 4mo agoVery interesting paper and I must say that I totally agree with it. But also that is something that is not new. I would say that the initial expectation is a bit off. I never expected that picking up any agentic coding solution, drop it in a project and fire at it a list of tasks would just magically work and follow a project pre-defined constraints. I do not believe that any agentic coding stack comes out of the box capable of this. Agents still need proper mechanics to understand the context, constraints and objectives reliably and that's still a work in progress as we can see by the constant updates on tools and skills and processes from Leading AI labs. They are now trying to fill that additional layer, which by the way could be much more profitable then bare model and token consumption. I would also argue that current OS models, like the ones tested, if properly driven can already produce production code following the desired constraints. What has been you experience? What has your production code looked like in recent months?
- pron 4mo agoThe situation is worse. Not only do agents have more difficulty under "structural constraints", but structural constraints may need to change, and agents are even worse at that. When designing a system or a component we have ideas that form invariants. Sometimes the invariant is big, like a certain grand architecture, and sometimes it’s small, like the selection of a data structure. Except, eventually, you’ll want to add a feature that clashes with that invariant. At that point there are usually three choices: - Don’t add the feature. The invariant is a useful simplifying principle and it’s more important than the feature; it will pay dividends in other ways. - Add the feature inelegantly or inefficiently on top of the invariant. Hey, not every feature has to be elegant or efficient. - Go back and change the invariant. You’ve just learnt something new that you hadn’t considered and puts things in a new light, and it turns out there’s a better approach. Often, only one of these is right. Often, at least one of these is very, very wrong, and with bad consequences. Even when they are able to follow constraints, agents are terrible at identifying when the constraints need to change.
- abalashov 4mo agoDespite my very limited enthusiasm for agentic coding, I have some experience with it, and my experience matches what you say perfectly. This is one of the seams that runs between pattern recognition and reasoning, and--despite the marketing claims around chain of thought--LLMs do not reason, not in the slightest. All attempts to make them appear to reason are basically recursive confinement efforts by the harness, to try to get the lightning into the bottle.
- MultiAgt 4mo ago[flagged]
- jessyt 4mo ago[flagged]
- danborn26 4mo ago[dead]
- sspoisk 4mo ago[flagged]
- AIorNot 4mo agoWhy the heck does this have to be a scientific study? since when did SWEs publish archive style science peices.. lol a good blog post would have been better. LLMs write working code, but have trouble following the script. They are slot machines of code. Human oversight is under pressure to deliver faster but code takes time to comprehend and analyze. Also in LLM coding, we end up with lots of natural language based spec files to manage and code we don't have an intuitive feel for unless we commit to the rigor of deep code review..(which no human really does anyway)
- dimitrismrtzs 4mo ago[dead]
- ucekmez 4mo agoThe paper frames it as a model thing but I believe a lot of it is the interface. HTML, undocumented JSON, APIs move frequently. that's why the agent is re-guessing constraints every call and that sure compounds. Static typing works because the constraint is in the compiler (not in the model). Same for schemas, manifests, signed stuff. Often the 'agent forgot the rule' just means nothing in the stack ever carried the rule.