5 ms·
> The Claude C Compiler illustrates the other side: it optimizes for > passing tests, not for correctness. It hard-codes values to satisfy > the test suite. I
by roadbuster 7mo ago
> The Claude C Compiler illustrates the other side: it optimizes for
> passing tests, not for correctness. It hard-codes values to satisfy
> the test suite. It will not generalize.
This is one of the pain points I am suffering at work: workers ask coding agents to generate some code, and then to generate test coverage for the code. The LLM happily churns out unit tests which are simply reinforcing the existing behaviour of the code. At no point does anyone stop and ask whether the generated code implements the desired functional behaviour for the system ("business logic").
The icing on the cake is that LLMs are producing so much code that humans are just rubber stamping all of it. Off to merge and build it goes.
I have no constructive recommendations; I feel the industry will keep their foot on the pedal until something catastrophic happens.
- Herring 7mo ago> At no point does anyone stop and ask whether the generated code implements the desired functional behaviour for the system ("business logic"). Obvious question: why not? Let’s say you have competent devs, fair assumption. Maybe it’s because they don’t have enough time for solid QA? Lots of places are feature factories. In my personal projects I have more lines of code doing testing than implementation.
- sarchertech 7mo agoIt’s because people will do what they’re incentivized to do. And if no one cares about anything but whether the next feature goes out the door, that’s what programmers will focus on. Honestly I think the other thing that is happening is that a lot of people who know better are keeping their mouths shut and waiting for things to blow up. We’re at the very peak of the hype cycle right now, so it’s very hard to push back and tell people that maybe they should slow down and make sure they understand what the system is actually doing and what it should be doing.
- shigawire 7mo agoOr if you say we should slow down your competence is questioned by others who are going very fast (and likely making mistakes we won't find until later). And there is an element of uncertainty. Am I just bad at using these new tools? To some degree probably, but does that mean I'm totally wrong and we should be going this fast?
- catlifeonmars 7mo agoThere is a saying: slow is smooth and smooth is fast. I have personally outpaced some of my more impatient colleagues by spending extra time up front setting up test harnesses, reading specifications, etcetera. When done judiciously it pays off in time scales of weeks or less.
- citizenpaul 7mo agooh yeah, let them dig a hole and charge sweet consultant rates to fix it. the the healing can begin
- harimau777 7mo agoDevelopers aren't given time to test and aren't rewarded if they do, but management will rain down hellfire upon their heads if they don't churn out code quickly enough.
- ojo-rojo 7mo agoHow about a subsequent review where a separate agent analyzes the original issue and resultant code and approves it if the code meets the intent of the issue. The principle being to keep an eye out for manual work that you can describe well enough to offload. Depending on your success rate with agents, you can have one that validates multiple criteria or separate agents for different review criteria.
- g947o 7mo agoYou are fighting nondeterministic behavior with more nondeterministic behavior, or in other words, fighting probability with probability. That doesn't necessarily make things any better.
- tbossanova 7mo agoProbably true (Sorry.)
- pyridines 7mo agoIn my experience, an agent with "fresh eyes", i.e., without the context of being told what to write and writing it, does have a different perspective and is able to be more critical. Chatbots tend to take the entire previous conversational history as a sort of canonical truth, so removing it seems to get rid of any bias the agent has towards the decisions that were made while writing the code. I know I'm psychologizing the agent. I can't explain it in a different way.
- citizenpaul 7mo agoI think of it as they are additive biased. ie "dont think about the pink elephant ". Not only does this not help llms avoid pink elphants instead it guarantees that pink elephant information is now being considered in its inference when it was not before. I fear thinking about problem solving in this manner to make llms work is damaging to critical thinking skills.
- Foobar8568 7mo ago
- MartyMcBot 7mo ago[flagged]
- tbossanova 7mo agoOnce upon a time people advocated writing tests first…
- codegladiator 7mo agoonce upon a time 'engineering' in software had some meaning attached to it... no other engineering profession would accept the standards(or rather their lack of) on which software engineering is running.
- hulitu 7mo ago> no other engineering profession would accept the standards(or rather their lack of) on which software engineering is running. I have bad news for you: they are pushing those "standards" (Agile, ASPICE) also in hardware and mechanical engineering. The results can be seen already. Testing is expensive and this is the field where most savings can be implemented.
- philipallstar 7mo agoAgile isn't a coding standard or approach.
- 8note 7mo agoi dont think that would help. the agent would hard code the test details into the code.
- tbossanova 7mo agoAh, so another way they’re like humans haha
- scotty79 7mo ago
- bentobean 7mo agoThis hits hard. I’m getting hit with so much slop at work that I’ve quietly stopped being all that careful with reviews.
- SoftTalker 7mo agoUm, you're supposed to write the tests first. The agents can't do this?
- daliusd 7mo agoThey can, but should be explicitly told to do that. Otherwise they just everything in batches. Anyway pure TDD or not but tests catches only what you tell AI to write. AI does not now what is right, it does what you told it to do. The above problem wouldn’t be solved by pure TDD.
- alexsmirnov 7mo agoActually, they extremely bad at that. All training data contains cod + tests, even if tests where created first. So far, all models that I tried failed to implement tests for interfaces, without access to actual code.
- WhyNotHugo 7mo agoThis is why you write the tests first and then the code. Especially when fixing bugs, since you can be sure that the test properly fails when the bug is present.
- usefulcat 7mo agoAgreed 1000%. But that can be a lot of work; creating a good set of tests is nearly as much or often even more effort than implementing the thing being tested. When LLMs can assist with writing useful tests before having seen any implementation, then I’ll be properly impressed.
- byzantinegene 7mo agofrom experience, AI is bad at TDD. they can infer tests based on written code, but are bad at writing generalised test unless a clear requirement is given, so you the engineer is doing most of the work anyway.
- 9rx 7mo agoMy day job has me working on code that is split between two different programming languages. I'd say LLMs are pretty good at TDD in one of those languages and a hot mess in the other. Which, funny enough, is a pretty good reflection of how I thought of the people writing in those languages before LLMs: One considers testing a complete afterthought and in the wild it is rare to find tests at all, and when they are present they often aren't good. Whereas the other brings testing as a first-class feature and most codebases I've seen generally contain fairly decent tests. No doubt LLM training has picked up on that.
- pmontra 7mo agoWhen fixing bugs, yes. When designing an app not so much because you realize many unexpected things while writing the code and seeing how it behaves. Often the original test code would test something that is never built. It's obvious for integration tests but it happens for tests of API calls and even for unit tests. One could start writing unit tests for a module or class and eventually realize that it must be implemented in a totally different way. I prefer experimenting with the implementation and write tests only when it settles down on something that I'm confident it will go to production.
- IAmGraydon 7mo agoYeah this is the exact kind of ridiculousness I've noticed as well - everything that comes out of an LLM is optimized to give you what you want to hear, not what's correct.
- 8note 7mo ago> At no point does anyone stop and ask whether the generated code implements the desired functional behaviour for the system ("business logic"). its fun having LLMs because it makes it quite clear that a lot of testing has been cargo-culting. did people ever check often that the tests check for anything meaningful?
- taatparya 7mo agoProperty testing could've helped
- Foobar8568 7mo ago15years ago, I had tester writing "UI tests" / "User tests" that matched what the software was cranking out. At that time I just joined to continue at the client side so I didn't really worked on anything yet. I had a fun discussion when the client tried to change values... Why is it still 0? Didn't you test? And that was at that time I had to dive into the code base and cry.
- mattacular 7mo agoTest automation is kind of like a religion. It is comforting to believe that the solution to code is more code.
- porphyra 7mo agoAt my job we have a requirement for 100% test coverage. So everyone just uses AI to generate 10,000 line files of unit tests and nobody can verify anything.
- harimau777 7mo agoExactly! It's frustrating how much developers get blamed for the outcomes of incompetent management.
- philipallstar 7mo ago> everyone just uses AI to generate 10,000 line files of unit tests and nobody can verify anything This is not a guaranteed outcome of requiring 100% coverage. Not that that's a good requirement, but responding badly to a bad requirement is just as bad.
- DeathArrow 7mo ago>LLM happily churns out unit tests which are simply reinforcing the existing behaviour of the code. At no point does anyone stop and ask whether the generated code implements the desired functional behaviour for the system ("business logic"). You can use spec driven development and TDD. Write the tests first. Write failing code. Modify the code to pass the tests.
- scotty79 7mo ago> The LLM happily churns out unit tests which are simply reinforcing the existing behaviour of the code. I always felt like that's the main issue with unit testing. That's why I used it very rarely. Maybe keeping tests in the separate module and not letting th Agent see the source during writing tests and not letting agent see the tests while writing implemntation would help? They could just share the API and the spec. And in case of tests failing another agent with full context could decide if the fix should be delegated to coding agent or to testing agent.
- ZaoLahma 7mo agoI think it boils down to how companies view LLMs and their engineers. Some companies will do as you say - have (mostly clueless) engineers feed high level "wishes" to (entirely clueless) LLMs, and hope that everyone kind of gets it. And everyone will kind of get it. And everyone will kind of get it wrong. Other companies will have their engineers explicitly treat the LLMs as collaborators / pair programmers, not independent developers. As an engineer in such a company, YOU are still the author of the code even if you "prompted" it instead of typing it. You can't just "fix this high level thing for me brah" and get away with it, but instead need to continuously interact with the LLM as you define and it implements the detailed wanted behaviors. That forces you to know _exactly_ what you want and ask for _exactly_ what you want without ambiguity, like in any other kind of programming. The difference is that the LLM is a heck of a lot quicker at typing code than you are.
- mrighele 7mo ago> The LLM happily churns out unit tests which are simply reinforcing the existing behaviour of the code This is true for humans too. Tests should not be written or performed by the same person that writes the code
- nly 7mo agoThat's a complete fantasy world where companies have twice the engineers they actually need instead of half.
- missingdays 7mo ago> [Reviews] should not be written or performed by the same person that writes the code > That's a complete fantasy world where companies have twice the engineers they actually need instead of half.
- harimau777 7mo agoAgreed, but then companies shouldn't complain about the consequences of understaffing their teams.
- yanis_t 7mo agoHow long till the industry discover TDD?
- bluefirebrand 7mo ago> I have no constructive recommendations; I feel the industry will keep their foot on the pedal until something catastrophic happens I can't wait. Maybe when shitty vibe coded software starts to cause real pain for people we can return to some sensible software engineering I'm not holding my breath though
- littlestymaar 7mo agoLong time ago in France the mainstream view by computer people was that code or compute weren't what's important when dealing with computers, it is information that matters and how you process it in a sensible way (hence the name of computer science in French: informatique. And also the name for computer: “ordinateur”, literally: what sets things into order). As a result, computer students were talked a lot (too much for most people's taste, it seems) about data modeling and not too much about code itself, which was viewed as mundane and uninteresting until the US hacker culture finally took over in the late 2000th. Turns out that the French were just right too early, like with the Minitel.
- msh 7mo ago"Computer science is no more about computers than astronomy is about telescopes." -Dijkstra
- harimau777 7mo agoHonestly, unit tests (at least on the front-end) are largely wasted time in the current state of software development. Taking the time that would have been spent on writing unit tests and instead using it to write functionally pure, immutable code would do much more to prevent bugs. There's also the problem that when stack rank time comes around each year no one cares about your unit tests. So using AI to write unit tests gives me time to work on things that will actually help me avoid getting arbitrarily fired. I wish that software engineers were given the time to write both clean code and unit tests, and I wish software engineers weren't arbitrarily judged by out of touch leadership. However, that's not the world we live in so I let AI write my unit tests in order to survive.
- msh 7mo agoI like unit tests when I have to modify code that someone made years ago, as a basic sanity check.
- DiscourseFan 7mo agoYou are overvaluing “clean code.” Code is code, it either works within spec or it doesn’t; or, it does but there are errors, more or less catastrophic, waiting to show themselves at any moment. But even in that latter case, no single individual can know for certain, no matter how much work they put in, that their code is perfect. But they can know its useable, and someone else can check to make sure it doesn’t blow something else up, and that is the most important thing.
- Illniyar 7mo agoBuilding a C compiler should not have this problem. There is probably a million test suites coming from outside the LLM that it can sue verify correctness.
- cmrdporcupine 7mo agoMy only hope is that all of this push leads in the end to the adoption of more formal verification languages and tools. If people are having to specify things in TLA+ etc -- even with the help of an LLM to write that spec -- they will then have something they can point the LLM at in order for it to verify its output and assumptions.
- HWR_14 7mo ago> The icing on the cake is that LLMs are producing so much code that humans are just rubber stamping all of it. I don't understand the value of that much code. What features are worth that much more than stability?
- salawat 7mo agoMwahahahahaha! Suffer, devs, SUFFER! KNOW MY PAIN! Ah hem... Welcome to the wonderful world of Quality Assurance, software developing audience. That part of the job, after you yeet your code over the fence, where the job is to bridge the gap between your madness, and the madness of the rest of the business. Here you will find: frustration, an ever present sense the rest of the world is just out to make your life more difficult, a creeping sense of despair, a hot ice pick in the back of your mind every time the language model does something syntactically valid, but completely nonsensical in the real world, the development of an ever increasing time horizon over which you can accurately predict the future, but no one will believe you anyway, a smoldering hatred of the overly confident executive with an over developed capacity for risk tolerance; a desire to run away and start a farm, and finally, a fundamental distrust of everything software, and all the people who write it. Don't forget your complimentary test framework and swag bag on your way out, and remember, you're here forever. You can try to check out, but you can never leave.
- rcpffm 7mo agoThx. you hit the nail