3 ms·
I believe the ChatGPT code has a bug, in that it accepts three spaces or tabs before a code fence, while the Google Markdown spec says up to three spaces, and d
by eesmith 8mo ago
I believe the ChatGPT code has a bug, in that it accepts three spaces or tabs before a code fence, while the Google Markdown spec says up to three spaces, and does not allow a tab there.
I also see that the tests generated by ChatGPT are far too few for the code features implemented. The cannot be the result of actual red/green TDD where the test comes before the feature is added.
For examples, 1) the code allows "~~~" but only tests for "```", 2) there are no tests when len(fence) < fence_len nor when len(fence) > fence_len, and 3) there are no tests for leading spaces.
There's also duplicate code. The function _strip_closing_hashes is used once, in the line:
text = _strip_closing_hashes(m.group("text")).strip()
The function is:
def _strip_closing_hashes(s: str) -> str:
s = s.rstrip()
# remove trailing " ###" style closers
s = re.sub(r"[ \t]+#+\s*$", "", s).rstrip()
return s
The ".rstrip()" is unneeded as the ".strip()" does both lstrip and rstrip.
I think that rstrip() should be replaced with a strip(), the function renamed to "_get_inline_content", and used as "text = _get_inline_content(m.group("text")).
Also, the Google spec also says "A sequence of # characters with anything but spaces following it is not a closing sequence, but counts as part of the contents of the heading:" so is it really correct to use "\s*" in that regex, instead of "[ ]*"? And does it matter, since the input was rstrip'ped already?
So perhaps:
def _get_inline_content(s: str) -> str:
s = s.rstrip(" ") # remove trailing spaces
s = s.rstrip("#") # removing "#" style closers
return s.strip() # remove leading and trailing whitespace
would be more correct, readable, and maintainable?
- simonw 8mo ago100%. That's why if you want good code you need to pay attention to what it's writing and testing and throw feedback like that at it.
- eesmith 8mo agoMy points though are 1) the development isn't actually using red/green TDD, and 2) the result doesn't show "really good results", including not following a very well-defined specification so doesn't work as a concrete example of your description of what the second chapter is supposed to be about. Perhaps you could show the process of refining it more, so it actually is spec compliant and tests all the implemented features? What's the outcome difference between this approach vs. something which isn't TDD, likes test-after with full branch coverage or mutation testing? Those at least are more automatable than manual inspection, so a better fit to agentic coding, yes? (Of course regular branch coverage doesn't test all the regexp branches, which makes regexp use tricky to test.)
- simonw 8mo agoYeah I'm going to ditch those examples and find better ones. I was hoping to illustrate the idea as simply as possible but they're not up to scratch.
- eesmith 8mo agoI think the problem-to-solve is a good one. The Google Markdown spec is very clear, with plenty of examples, and I think the problem is well-defined. I've seen entirely too many examples of how to use TDD which give under-specified toy problems, where the solution is annoyingly incomplete for something more realistic. And I've seen TDD projects which didn't follow the spec, but instead implemented the developers' misconceptions about the spec. That's exactly what we see here with Markdown, where there's a spec, along with a lot of non-conformant examples in the training set by people who didn't read the spec but instead based it on their experiences in using Markdown. The code generated by ChatGPT is almost correct. Seeing the process of how to get from that to a valid and well-tested solution would make for a good demonstration of the full process. I'll again add that showing how to integrate something like branch coverage or hypothesis testing for automatic test suite generation would be really useful.
- eesmith 8mo agoWill you be updating the text at https://simonwillison.net/guides/agentic-engineering-patterns/red-green-tdd/ https://simonwillison.net/guides/agentic-engineering-pattern...? As it currently says: > A significant risk with coding agents is that they might write code that doesn't work, or build code that is unnecessary and never gets used, or both. > Test-first development helps protect against both of these common mistakes, and also ensures a robust automated test suite that protects against future regressions. while the ChatGPT generated code contains bugs, contains unnecessary code which never gets used, and the ChatGPT generated test suite is not robust. (As an example of unnecessary code which never gets used, _FENCE_RE contains "(?P<info>.*)$" but neither the group name nor the group are used, and the pattern is unneeded -- and all of the tests pass without it.) Your writings are widely read and influential. I think it's important that you let readers know the results produced in your experiment are not actually a complete example of a "fantastic fit" of Red/Green TDD for coding agents, and to highlight their limitations.