8 ms·
No. The way to code going forward with AI is Test Driven Development. The code itself no longer matters. You give the AI a set of requirements, ie. tests that
by reenorap 1y ago
No.
The way to code going forward with AI is Test Driven Development. The code itself no longer matters. You give the AI a set of requirements, ie. tests that need to pass, and then let it code whatever way it needs to in order to fulfill those requirements. That's it. The new reality us programmers need to face is that code itself has an exact value of $0. That's because AI can generate it, and with every new iteration of the AI, the internal code will get better. What matters now are the prompts.
I always thought TDD was garbage, but now with AI it's the only thing that makes sense. The code itself doesn't matter at all, the only thing that matters is the tests that will prove to the AI that their code is good enough. It can be dogshit code but if it passes all the tests, then it's "good enough". Then, just wait a few months and then rerun the code generation with a new version of the AI and the code will be better. The humans don't need to know what the code actually is. If they find a bug, write a new test and force the AI to rewrite the code to include the new test.
I think TDD has really found its future now that AI coding is here to stay. Human code doesn't matter anymore and in fact I would wager that modifying AI generated code is as bad and a burden. We will need to make sure the test cases are accurate and describe what the AI needs to generate, but that's it.
- pcarolan 1y agoI mostly agree, but why stop at tests? Shouldn’t it be spec driven development? Then neither the code or the language matter. Wouldn’t user stories and requirements à la bdd (see cucumber) be the right abstraction?
- reenorap 1y agoI don't think you're wrong but I feel like there's a big bridge between the spec and the code. I think the tests are the part that will be able to give the AI enough context to "get it right" quicker. It's sort of like a director telling an AI the high level plot of a movie, vs giving an AI the actual storyboards. The storyboards will better capture the vision of the director vs just a high level plot description, in my opinion.
- __MatrixMan__ 1y agoMaybe one day. I find myself doing plenty of course correction at the test level. Safely zooming out doesn't feel imminent.
- gmd63 1y agoWhy stop there? Whichever shareholders flood the datacenter with the most electrical signals get the most profits.
- int_19h 1y agoNatural language is too ambiguous for this, which makes it impossible to automatically verify What you need is indeed spec-driven development, but specs need to be written in some kind of language that allows for more formal verification. Something like https://en.wikipedia.org/wiki/Design_by_contract https://en.wikipedia.org/wiki/Design_by_contract, basically. It is extremely ironic that, instead, the two languages that LLMs are the most proficient in - and thus the ones most heavily used for AI coding - are JavaScript and Python...
- blibble 1y agoyou will end up with something that passes all your tests then smashes into the back of the lorry the moment it sees anything unexpected writing comprehensive tests is harder than writing the code
- reenorap 1y agoThen you write another test. That's the whole point of TDD. As you keep writing more tests, the closer it gets to its final form.
- blibble 1y agoright, and by the time I have 2^googolplex tests then the "AI" will finally be able to produce a correctly operating hello world oh no! another bug!
- 3vidence 1y agoI've definitely seen a number of files where the implementation is maybe like 500 LOC and the test file is 10000+ LOC. I agree rigidly defining exactly what the code does through tests is harder than people think.
- topaz0 1y agoThe idea of TDD is that you should have the tests before you have the code. If your code is failing in real life before you have the tests, that's no longer TDD.
- BobbyTables2 1y agoHave you ever seen someone carve the inverse of a statue from a solid block of stone? If so, they are doing TDD. Yeah, me neither…
- throwaway7783 1y agoAI can help here too, by exploding the spec into a series of questions to clarify behavior. Today, it just does something and when corrected it says "You are right!....".
- HellDunkel 1y agoNo. The reason AI code generation works so well is a) it is text based- the training data is huge and b) the output is not the final result but a human readable blueprint (source code), ready to be made fit by a human who can form an abstract idea of the whole in his head. The final product is the compiled machine code, we use compilers to do that, not LLMs. Ai genereted code is not suitable to be directly transferred to the final product awaiting validation by TDD, it would simply be very inefficient to do so.
- rvz 1y ago> We will need to make sure the test cases are accurate and describe what the AI needs to generate, but that's it. Yes. The first thing I always check in every project (an especially vibe-coded projects) is whether if: A. Does it have tests? B. Is the coverage over 70%? C. Do the tests actually test for the behaviour of the code (good) or just its implementation (bad.) If any of those requirements are missing, then that is a red flag for the project. While TDD is absolutely valuable for clean code, focusing too much on it can be the death of a startup. As you said the code itself is $0, then the first product is still worth $10 and the finished product is worth $1M+ once it makes money, which is what matters.
- sarchertech 1y ago> Then, just wait a few months and then rerun the code generation with a new version of the AI and the code will be better. How many times have you seen a code change that “passed all the tests” take down production or break an important customer’s workflow? Usually that was just a relatively small change. Now imagine that you regenerated literally all the code. The code is the spec. Any other spec comprehensive enough to cover all possible functionality has to be at least as complex as the code.
- wonnage 1y agoTDD is testing in production in disguise. After all, bugs are unexpected and you can’t write tests for a bug you don’t expect. Then the bug crops up in production and you update the test suite.
- embedding-shape 1y agoTDD has always been about two things for me; be able to move forward faster because I have something easy to execute that compares it against the known wanted state, and in the future preventing unwanted regressions. I'm not sure I've ever thought of unit testing as "prevent potential future bugs", mostly up front design prevents that, or I'd use property testing, but neither of those are inside the whole "write test then write code" flow.
- Retric 1y agoThe intended workflow of TDD is to write a set of tests before some code. The only reason that makes sense conceptually is to prevent possible future bugs from going undetected. Put another way if your TDD always pass then there’s no point in writing them, and there’s no known bugs before you have any code. So discovering future bugs that didn’t exist when you’re writing those tests is the point.
- black_knight 1y agoI don’t really understand how to write tests before the code… When I write code, the hard part is writing the code which establishes the language to solve the problem in, which is the same language the tests will be written in. Also, once I have written the code I have a much better understanding as the problem, and I am in a way better position to write the correct tests.
- baq 1y ago> The new reality us programmers need to face is that code itself has an exact value of $0. This is not new at all. Code has always been a liability. It having $0 value would be a great improvement IMHO. The value was always in the product regardless of the amount of code in it and regardless of its quality. Customers don’t buy code. (Except of course when the code is the product, which is very unusual nowadays.)
- epicureanideal 1y agoTDD doesn’t ensure the code is maintainable, extendable, follows best practices, etc, and while AI might write some code that can pass tests while the code is relatively small, I would expect in the long run it will find it extremely difficult to just “rewrite everything based on this set of new requirements” and then do that again, and again, and again, each time potentially choosing entirely different architectures for the solution.
- lkjdsklf 1y ago> TDD doesn’t ensure the code is maintainable, extendable, follows best practices, etc, and while AI might write some code None of that matters of its not a person writing the code
- sarchertech 1y agoAI has a hard time working with code that humans would consider hard to maintain and hard to extend. If you give AI a set of tests to pass and turn it loose with no oversight, it will happily spit out 500k LOC when 500 would do. And then it will have a very hard time when you ask it to add some functionality. AI routinely writes code that is beyond its ability to maintain and extend. It can’t just one shot large code bases either, so any attempt to “regenerate the code” is going to run into these same issues.
- jjav 1y ago> If you give AI a set of tests to pass and turn it loose with no oversight, it will happily spit out 500k LOC when 500 would do. And then it will have a very hard time when you ask it to add some functionality. I've been playing around with getting the AI to write a program, where I pretend I don't know anything about coding, only giving it scenarios that need to work in a specific way. The program is about financial planning and tax computations. I recently discovered AI had implemented four different tax predictions to meet different scenarios. All of them incompatible and all incorrect but able to pass the specific test scenarios because it hardcoded which one to use for which test. This is the kind of mess I'm seeing in the code when AI is left alone to just meet requirements without any oversight on the code itself.
- heavyset_go 1y agoThe code always matters. Black box coding like this leads to systems you can't explain, and that's your whole damn job: to understand the system you're building. Anything less is negligence.
- BobbyTables2 1y agoThe problem is — nobody commits code that fails tests. The bugs occur because the initial tests didn’t fully capture the desired and undesired behaviors. I’ve never seen a formal list of software requirements state that a product cannot take more than an hour to do a (trivial) operation. Nobody writes that out because it’s implicitly understood. Imagine writing a “life for dummies” textbook on how to grow from a 5yr old to 10yr old. It’s impossible to fully cover.
- rob_c 1y ago> The problem is — nobody commits code that fails tests. Hah, if that were true the industry would be a better place. Or a worse place. Or a slower place but exactly the same. I should build a test for that... I've worked on many projects where tests get disabled as nobody can tell why it's failing (or why it was even written in some cases). I've rewritten test systems from scratch in the past to drag projects out of the dumpster fire by getting them into a state of passing simple startup/shutdown safely routines and then watched as I pass the project onto others how it rots until some "genius" young coder comes along and "removes the slow test-suite because it takes 2hr+ to run on my way out of spec laptop".
- DanHulton 1y agoThis is incorrect for a lot of reasons, many of which have already been explored, but also: > with every new iteration of the AI, the internal code will get better This is a claim that requires proof; it cannot just be asserted as fact. Especially because there's a silent "appreciably" hidden in there between "get" and "better" which has been less and less apparent with each new model. In fact, it more and more looks like "Moore's law for AI" is dead or dying, and we're approaching an upper limit where we'll need to find ways to be properly productive with models only effectively as good as what we already have! Additionally, there's a relevant adage in computer science: "Debugging is twice as hard as writing the code in the first place. Therefore, if you write the code as cleverly as possible, you are, by definition, not smart enough to debug it." If the code being written is already at the frontier capabilities of these models, how the hell are they supposed to fix the bugs that crop up, especially if we can't rely on them getting twice as smart? ("They won't write the bugs in the first place" is not a realistic answer, btw.)
- lugu 1y agoThe argument they are making is that if a bug is discovered, the agent will not debug it, instead a new test case is created, and the code is regenerated (I suppose if a quick fix isn't found). That is why they don't need debugging agent twice as capable as coding agent. I don't know if this works in practice, as in my experience, tests are intertwined with the code base.
- CuriouslyC 1y agoJust because you're not writing code where you can see that the new models are appreciably better doesn't mean they aren't. LLM progress now isn't in making it magically appear smarter at the top end (that's in diminishing returns as you imply), but at filling in weak points in knowledge, holes in capability, improving default process, etc. That's relevant because it turns out most of the time the LLM doesn't fail at coding because it's not a general super genius, but because it just had a hole in its capabilities that caused it to be dumb in a specific scenario. Additionally, while the intelligence floor is shooting up and the intelligence ceiling is very slowly rising, the models are also getting better at following directions, writing cleaner prose, and their context length support is increasing so they can handle larger systems. The progress is still going strong, it just isn't well represented by top line "IQ" style tests. LLMs and humans are good at dealing with different kinds of complexity. Humans can deal with messy imperative systems more easily assuming they have some real world intuition about it, whereas LLMs handily beat most humans when working with pure functions. It just so happens that messy imperative systems are bad for a number of reasons, so the fact that LLMs are really good at accelerating functional systems gives them an advantage. Since functional systems are harder to write but easier to reason about and test, this directly addresses the issue of comprehending code.
- Yhippa 1y agoI remember talking about this with a friend a long time ago. Basically, you'd write up tests and there was a magic engine that would generate code that would self-assemble and pass tests. There was no guarantee that the code would look good or be efficient--just that it passed the tests. We had no clue that this could actually happen one day in the form of gen AI. I want to agree with you just to prove that I was right! This is going to bring up a huge issue though: nailing requirements. Because of the nature of this, you're going to have to spec out everything in great detail to avoid edge cases. At that point, will the juice be worth the squeeze? Maybe. It feels like good businesses are thorough with those kinds of requirements.
- jcgrillo 1y agoHow would you handle production incidents in such a codebase? The primary focus of a software engineer is to make the codebase easy (or at least possible) to understand. To tame complexity while achieving some business objectives. If we're going to just throw that part out the window you need to have a plan for how to operate the resultant mess in production.
- lifeformed 1y agoNot everything can be tested by a computer.
- zeroq 1y ago> code itself has an exact value of $0. That's because AI can generate it That's only true for problems that has been solved and well documented before. AI can't solve novel problems. I have ton of examples I use from time to time when new models come out. I've tried to ride the hype train, and I've been frustrated working with people before, but I've never been so frustrated as trying to make AI follow simple set of rules and getting: "Oh yes, my bad, I get that now. Black is white and white is black. Let me rewrite the code..." My favorite example is tasked AI with a rudimentary task and it gave me a working answer but it was fishy, so I googled the answer and lo and behold I landed on stackoverflow page with exact same answer being top voted answer to question very similar to my task. But that answer also had a ton of comments explaining why you never should do it that way. I've been told many times that "you know, kubernetes is so complicated, but I tell AI what I want and it gives me a command I simply paste in my terminal". Fuck no. AI is great for scaffolding projects, working with typical web apps where you have repeatable, well documented scenarios, etc. But it's not a silver bullet.
- tartoran 1y agoIf the code doesn't matter anymore, in order of it to be of any quality the test should be as detailed as was the code in the first place, you'd end up writing the code in tests more or less.
- tcmart14 1y agoWhile TDD can have some merits, I think this is being way to generous to the value of tests. As Dijkstra said once, "Testing shows the presence, not the absence of bugs." I'm not a devout follower of Uncle Bob, but I was just thumbing through Clean Architecture today and he has a whole section to this point (including the above quote). Right after that quote he writes, "a program can be proven incorrect by a test, but it can not be proven correct." Which is largely true. The only garuntee of TDD is you can show a set of behaviors your program doesn't do, it never proves what the program actually does. To extrapolate to here, all TDD does it put up guardrails for the the AI should not generate.
- astahlx 1y agoIt depends on how you define testing now: Property-based testing would test sets of behaviors. The main idea is: Formalize your goal before implementing. So specification driven development would be the thing to aim for. And at some point we might be able to model check (proof) the code that has been generated. Then we are the good old idea of code synthesis.
- AstralStorm 1y agoDon't worry, you're going to be searching for logic vs requirements mismatches instead if the thing provides proofs. That means, you have to understand if it is even proving the properties you require for the software to work. It's very easy to write a proof akin to a test that does not test anything useful...
- practal 1y agoNo, that misunderstands what a proof is. It is very easy to write a SPEC that does not specify anything useful. A proof does exactly what it is supposed to do.
- svieira 1y agoNo, a proof proves what it proves. It does not prove what the designer of the proof intended it to prove unless the intention and the proof align. Proving that is outside of the realm of software.
- android521 1y agoThis is wrong in so many ways. Have you even tried what you believe? If you have tried, you would find out it is nonsense quickly.
- heavyset_go 1y agoThe irony is that I tried this with a project I've been meaning to bang out for years, and I think the OP's idea a natural thought to have when working with LLMs: "what if TTD but with LLMs" When I tried it, it "worked", I admittedly felt really good about it, but I stepped away for a few weeks because of life and now I can't tell you how it works beyond the high level concepts I fed into the LLM. When there's bugs, I basically have to derive from first principles where/how/why the bug happens instead of having good intuition on where the problem lies because I read/wrote/reviewed/integrated with the code myself. I've tried this method of development with various levels of involvement in implementation itself and the conclusion I came to is if I didn't write the code, it isn't "mine" in every sense of the term, not just in terms of legal or moral ownership, but also in the sense of having a full mental model of the code in a way I can intellectually and intuitively own it. Really digging into the tests and code, there are fundamental misunderstandings that are very, very hard to discern when doing the whole agent interfacing loop. I believe they're the types of errors you'd only pick up on if you wrote the code yourself, you have to be in that headspace to see the problem. Also, I'd be embarrassed to put my name on the project, given my lack of implementation, understanding and the overall quality of the code, tests, architecture, etc. It isn't honest and it's clearly AI slop. It did make me feel really productive and clever while doing it, though.
- svieira 1y ago> It did make me feel really productive and clever while doing it, though. And that's the greatest trap of this whole thing. That the _feels_ are so quickly diverged from the actual.
- thesz 1y ago> You give the AI a set of requirements, ie. tests that need to pass, and then let it code whatever way it needs to in order to fulfill those requirements. SQLite has tests-lines-to-code-lines ratio above 1000 (yes, 1000 lines of tests for single line of code) and still has bugs. AMD, at the time it decided to apply ACL2 to its FPU, had 29 million tests (not lines of code, but test inputs and outputs). ACL2 verification found several bugs in the FPU. Just to make a couple of points for someone to draw a line.
- deleted 1y ago[deleted]
- huflungdung 1y ago[dead]
- pjmlp 1y agoTry to do TDD with graphics programming. I never bought into TDD because it is only usefull for business logic, plain algorithms and data structures, it is no accident that is what 99% of conference talks and books focus on. There isn't a single TDD talk about shader programming for GPGPU, and validating that what the shader algorithms produce via automated tests, the reason being the amount of enginneering effort only to make it work, and still lacks human sensitivity for what gets rendered.
- MoreQARespect 1y agoI have. I call it snapshot test driven development. You put the preconditions in, generate and record the graphics as an artefact at runtime and when it looks right, freeze it.
- pjmlp 1y agoBut that isn't TDD, no line of code should be written without broken tests.
- MoreQARespect 1y agoYes it is. Until the artefact which has been visually validated is locked in it is still a broken test. You can argue semantics until you're blue in the face it still follows red-green-refactor and it confers the same benefits as TDD.
- brazukadev 1y agoYour nickname tells me you are not talking bs.
- mettamage 1y agoRelatable. Every time I read something about testing it seems backend web dev related. Of course, that’s great but what about the rest?
- rob_c 1y ago> The code itself no longer matters. Good luck explaining that when you get hacked out of oblivion. This is like saying the fine-print of contracts don't matter so I get "AI" to regurgitate them all for me as a lawyer. It's so wrong as to be beyond laughable. Put the coffee down and go for a walk, preferably to a library, and LEARN SOMETHING.
- fainpul 1y agoNo. TDD combined with vibe-coding can create code that has unwanted side-effects, because your tests only check the result. It can also have various security vulnerabilities, which you don't test for, because how would you know what to test. It can also lead to massive duplication and code bloat, while tests still pass. It can lead to software which wastes a lot of resources (memory, cpu, inefficient network requests and the like) due to bad algorithms. If you try to keep that in check by writing performance tests, how do you know what acceptable performance is, if you have no idea how your program works?
- CuriouslyC 1y agoTDD doesn't solve those problems for human code either. That's why every org has several security scanners that most engineers ignore unless you hard gate them, linting, code duplication detection, etc. Also, you can give AI a SLO for code and fail stress tests that don't meet it. AI will happily respond to a failing stress test with profiling and well thought out optimizations in many cases.
- skywhopper 1y agoWho is arguing that TDD solves those problems with human coders?
- MarcelOlsz 1y agoI've been experimenting with various TDD methods with AI and it cannot do frontend work. Frontend has too many ancient illogical incantations and ways of doing things that it has no clarity on, you have to handhold it every step of the way. When I let AI go off the rails and build a frontend it's an absolute mess and it frequently chooses the hardest and dumbest way to do things. Stellar for low surface-area work though. Once AI has cheap real-time eyes it might get slightly better, but all the logs and browser MCP tools and yadda yadda in the world will not get it to produce anything remotely efficient.
- mettamage 1y agoYou can have eyes by pasting in screenshots. So you could write tests that create a screenshot and send it to an llm if it doesn’t match the output.
- MarcelOlsz 1y agoBeen there done that lol. It needs real-time extremely badly. If I wanted to write English instead of code I'd have been a writer instead. It will nudge pixels but it will not take in the myriad of reasons that button is the way that it is and solve it in any meaningful way. Decent for MVPing with stuff like shadcn/tailwind but falls apart with anything else.
- player1234 1y ago[flagged]
- tmoertel 1y agoShow me the TDD tests you would use to show that your AI-generated code isn't creating security vulnerabilities.
- jcgrillo 1y agoIME devs actually do precisely the opposite. They write code and then ask the LLM to do the "boring" part and write the tests for them.
- dboreham 1y agoTests are also code and can be buggy, incomplete etc.
- lisbbb 1y agoWhy did you think TDD was garbage? Formalizing a specification is all that test first is. It's just that most devs I know had big egos and believing writing tests was somehow below them. I prefer the "build a little, test a little" approach, personally, but there's nothing inherently wrong with TDD. My prediction is that in the future, a lot of desperate companies are going to need living, breathing reverse software engineers to aid them because they have lost the ability to understand their own codebases. Oh, and why is code worth $0? A lot of code is throwaway, but I still got paid to produce it and much of it makes money for the company or saves them money.
- halfcat 1y agoI could not disagree more strongly with everything you’ve said in this comment. > The way to code going forward with AI is Test Driven Development. No. TDD already collapses under its own weight as a project grows. > The code itself no longer matters. No. Definitely no. That’s absurd. You can’t box in a correct solution with guard rails. Especially since, even if you could get something close to that, you would also lose the ability to understand the tests. > You give the AI a set of requirements, ie. tests that need to pass, and then let it code whatever way it needs to in order to fulfill those requirements. That's it. The new reality us programmers need to face is that code itself has an exact value of $0. No. The opposite. When code is cheap, understanding and control become expensive. Code a human can understand will be the most valuable going forward. > That's because AI can generate it, and with every new iteration of the AI, the internal code will get better. No. All code is technical debt. AI produces code faster. Therefore AI produces bugs faster. ”Debugging is twice as hard as writing the code in the first place. Therefore, if you write the code as cleverly as possible, you are, by definition, not smart enough to debug it” -Brian Kernighan This is literally where we’re at. AI writes code just beyond its ability to fix. > What matters now are the prompts. No. This is such a dead end. It’s a roll of the dice, and so we have examples of people who seem to get it to build something faster. That’s like saying there are people who win the lottery. It’s true, and it also says nothing of your ability to repeat their process. Confirmation bias of the wins. But in building something reliable, we care more about the floor (minimum quality) than the ceiling (the peak it can reach sometimes).
- RayVR 1y agoI feel terrible for anyone relying on anything you produce as a proompt engineer