4 ms·
My points though are 1) the development isn't actually using red/green TDD, and 2) the result doesn't show "really good results", including not following a ve
by eesmith 8mo ago
My points though are
1) the development isn't actually using red/green TDD, and
2) the result doesn't show "really good results", including not following a very well-defined specification
so doesn't work as a concrete example of your description of what the second chapter is supposed to be about.
Perhaps you could show the process of refining it more, so it actually is spec compliant and tests all the implemented features?
What's the outcome difference between this approach vs. something which isn't TDD, likes test-after with full branch coverage or mutation testing? Those at least are more automatable than manual inspection, so a better fit to agentic coding, yes?
(Of course regular branch coverage doesn't test all the regexp branches, which makes regexp use tricky to test.)
- simonw 8mo agoYeah I'm going to ditch those examples and find better ones. I was hoping to illustrate the idea as simply as possible but they're not up to scratch.
- eesmith 8mo agoI think the problem-to-solve is a good one. The Google Markdown spec is very clear, with plenty of examples, and I think the problem is well-defined. I've seen entirely too many examples of how to use TDD which give under-specified toy problems, where the solution is annoyingly incomplete for something more realistic. And I've seen TDD projects which didn't follow the spec, but instead implemented the developers' misconceptions about the spec. That's exactly what we see here with Markdown, where there's a spec, along with a lot of non-conformant examples in the training set by people who didn't read the spec but instead based it on their experiences in using Markdown. The code generated by ChatGPT is almost correct. Seeing the process of how to get from that to a valid and well-tested solution would make for a good demonstration of the full process. I'll again add that showing how to integrate something like branch coverage or hypothesis testing for automatic test suite generation would be really useful.
- eesmith 8mo agoWill you be updating the text at https://simonwillison.net/guides/agentic-engineering-patterns/red-green-tdd/ https://simonwillison.net/guides/agentic-engineering-pattern...? As it currently says: > A significant risk with coding agents is that they might write code that doesn't work, or build code that is unnecessary and never gets used, or both. > Test-first development helps protect against both of these common mistakes, and also ensures a robust automated test suite that protects against future regressions. while the ChatGPT generated code contains bugs, contains unnecessary code which never gets used, and the ChatGPT generated test suite is not robust. (As an example of unnecessary code which never gets used, _FENCE_RE contains "(?P<info>.*)$" but neither the group name nor the group are used, and the pattern is unneeded -- and all of the tests pass without it.) Your writings are widely read and influential. I think it's important that you let readers know the results produced in your experiment are not actually a complete example of a "fantastic fit" of Red/Green TDD for coding agents, and to highlight their limitations.
- simonw 8mo agoI'll be replacing the examples with ones that better illustrate the technique. I dashed off those off in a hurry using the wrong tools (I used ChatGPT and Claude directly, not the Coding agent harnesses Claude Code and Codex) and that was a mistake.
- eesmith 8mo agoYou didn't think they were the wrong tools when you wrote it. You said "this example is simple enough that both Claude and ChatGPT can implement it using their default code environments". From what I gather, a lot of people are using these code assistance tools because they too are in a hurry, under pressure from management forcing them to go faster with AI, and with limited ability to push back. You have significantly more experience than most of your readership. Will you be providing guidelines about which tools to avoid for which problems, based on your experience? Will you use this or something similar as an example of the negative consequence of being in a hurry, hopefully leading to a worked-out example of one might better audit or inspect tool-generated code, and the effort involved? That would be invaluable for people dealing with overly-optimistic management pressure. My personal belief is that one of the reasons for TDD's success is as a way for programmers to respond to ill-advised pressure to skimp on testing found in some test-after shops. That disappears if managers believe instructing an agentic code generator to "use Red/Green TDD" easily ensures a robust automated test suite. My apologies if you have already done this. I have not followed your work. My interest in this thread is from my views of TDD as a development approach, and the difficulty in generating a test suite which is robust, minimal, understandable, and maintainable.