7 ms·
My Agent Skill for Test-Driven Development
- behnamoh 4mo agoSnake oil. Just ask the model, all these custom agents/skills haven't proven that useful in practice.
- beezlewax 4mo agoI've found them useful for in house stuff where you are using a specific design system or architecture. But custom everything works best. Are that Claude works well on its own though at this point.
- john_strinlai 4mo agonot that i know much about the effectiveness of these skill files, i find it odd to call something given for free "snake oil", which i thought referred to the sale of fraudulent products (to the benefit of the snake oil salesperson), typically around healthcare-related stuff.
- jw1224 4mo agoSkills already are "just asking the model". Unless you'd prefer to type out the same instructions every single time? Skills are literally just Markdown documents that get loaded into context when the /skill-name is invoked.
- Zetaphor 4mo agoI think they're maybe confusing Skills and MCP servers
- dominotw 4mo agoi belive gp means llms produce what they see in training data/rl there isnt much too much customization you can do with skills. they are being sold as more powerful than they are. Like llms are intelligent blank slates that can be customized with mere markdown files.
- calebkaiser 4mo agoI don't understand this line of criticism exactly. By putting new information in the context window, you are materially changing the activations at your point of sampling, which is literally "customizing with mere markdown files." Taken to the extreme, the attitude that there is some special incantation that will unlock all capabilities is silly, and a lot of the "prompt engineering" discourse is similarly kind of dumb, but in-context learning is clearly a real thing.
- dominotw 4mo agoeven if that works one time you can never be sure that your customization is in place or fell out of context's important zone or is contradicted by later context . you've reverted back to base llm behavior. you are treating skill like sure thing
- krupan 4mo agoIt's as of people crave some sort of control and/or determinism from these chaotic tools, and so they have to believe that these "skills" make a difference
- coffeeaddict1 4mo agoI disagree. Not all skills are useless. For example, I sometime use Qt for GUI projects and I have found their skills [0] very useful to improve the quality and performance of my projects. I their absence, I would each time have to direct the agents to find the docs or specific tools, wasting tokens and thus decreasing the quality of the output. [0] https://github.com/TheQtCompanyRnD/agent-skills https://github.com/TheQtCompanyRnD/agent-skills
- pramodbiligiri 4mo agoI don't think the idea of skills is quite snake oil. It seems you can change what LLM outputs next by what's called few-shot prompting or in-context learning: https://www.promptingguide.ai/techniques/fewshot https://www.promptingguide.ai/techniques/fewshot
- wyre 4mo agoYa, if im constantly asking a model to do TDD development, you know what would make it a lot easier? A skill.
- theptip 4mo agoNah. Skills are great. But you should write your own.
- internet101010 4mo agoLol wut. One of first things people do at a company when they get enterprise LLM tools is share a skill with company-specific color palettes or standards for creating visualizations (I prefer Tufte's principles).
- simonw 4mo agoThis article would benefit from a date. It looks like it's recent (Internet Archive first grabbed it on May 29th) but it's the kind of information that can quickly become stale as models and agents improve. (I've been getting solid results recently from simply telling Claude Code and Codex "Test with uv run pytest, use red/green TDD".)
- disgruntledphd2 4mo agoMe too, although I dislike the fact that it over-focuses on mocks (which I accept is over-represented in the training data).
- galsapir 4mo agosometimes I also feel it tries to optimise for "per line coverage" over more "real, complex use cases" type tests
- porphyra 4mo agoA lot of prompt engineering goes out of date quickly. Nobody nowadays goes "you are an expert software engineer. make no mistakes" lol. As a personal anecdote, I find that a lot of big prompts and skills use up context window budget and in many cases agents will eagerly try to use a skill even if it isn't super relevant or necessary for the current task. So when I have too many skills I have to spend a bunch of time toggling the checkboxes to figure out which ones are needed for the task at hand before starting...
- Royce-CMR 4mo agoI can't find the link now, but Anthropic has a post about using either a light model call or other logic (regex etc) to dynamically decide what tools to expose per incoming request. I've run into the same issue and I still end up manually curtailing what's exposed to the model, limiting to the task at hand, but I like the idea of another (smaller I hope) model doing 70% of the clipping instead, automagically.
- dluxem 4mo agoI believe using a skill here is the wrong approach. LLMs already know what TDD is and how to do it, just like object oriented programming. If this is encoded in a skill, that skill essentially has to be loaded for everything thing your LLM is doing. This is probably one of the few areas where direct instructions via AGENTS.md is best, and I don't believe it requires much direction here to force the issue. But I think the OP is just trying to have their agent work in a very specific way -- that is fine too. > 5. Show me the test and ask for approval before continuing
- zuzululu 4mo agoPeople forget skill is just a markdown file and I don't think TDD makes sense. It's more for specific niches like working on your custom codebase or some less beaten paths you take and save the lessons going forward But everybody is free to choose how they work and it may be required in ways that we can't know about.
- jasonswett 4mo agoMy experience has been that yes, LLMs already know about TDD, OOP, etc., but they won't necessarily BEHAVE according to what they know unless you tell them. And of course, they "know" a lot of things that conflict with each other.
- steno132 4mo agoTest driven development is one of the worst ideas nowadays in the LLM age. We have models that can consistently write expert level, usually bug free code for you and rapidly fix even complex bugs in your codebase. The token cost and tech debt introduced by tests is just not worth it. There's usually no bugs and if there are, you can fix them quickly if and when it's needed.
- buster 4mo agoNo it's probably the most important idea.
- Ginop 4mo agoI disagree Testing was and is still very important, as LLMs can still miss important points in business logic or other edge cases I would argue that tests became as important as code, if not more.
- esafak 4mo agoIF your code has no bugs it's either trivial or you haven't noticed the bugs.
- joshgachnang 4mo agoThis overall is pretty close to how I've set up my implementation skill. One thing I'm curious about is how well the analogies like "We don't make dinner in a dirty kitchen." work vs something a lot more straightforward. Any input OP?
- jasonswett 4mo agoOP here. I don't know, in my experience Claude took "clean the kitchen before we make dinner" to heart in an astonishingly productive way. I haven't tried many other analogies though.
- zuzululu 4mo agoTDD sounds great on paper for agentic development but you quickly realize it balloons the token cost. Often I write some feature and then its repurposed or removed, code is refactored moved around as time goes. With TDD I would be taxed heavily and velocity slow to a crawl. The waterfall approach is better after trying out TDD especially when you have a multi-agent setup. Also I found that in some cases the tests were just superficial hallucinations that never actually tested the components written or there some some context corruption and ultimately triggered a false positive that kicked off a completely unintentional refactoring.
- reg_dunlop 4mo agoBut that repurposing/removal is exactly what's avoided if you follow through with the SEF framework he outlines. I have to push back on the idea that token costs balloon when using TDD within the context of a strong framework such as Jason has laid out here. If the feature is repurposed/removed/refactored....I'd argue the specification wasn't well thought out prior to burning into tokens. We're so eager to do a lot of the wrong things quickly, when it may serve us better to do a more precise thing slowly.
- zuzululu 4mo agoYou cant spec out what you dont know, scope, requirements change from real world feedback
- reg_dunlop 4mo agoThen adjust the specs with the scope and feedback. I fail to see the argument you're making... Features aren't made in a vacuum. If specs are made/written....with the information available now...then it's better than not writing specs. Writing specs with incomplete information is better than not writing specs with incomplete information.
- __mharrison__ 4mo agoMy experience is the opposite. TDD keeps the guardrails on and let's me refactor with confidence. Crazy times here in the development world. I'm always curious to watch other's best practices.
- jvuygbbkuurx 4mo agoAll of these post are missing actual comparisons on results. I read exactly opposite 'you should do x' everyday. If TDD actually was better it would simply be in the system prompts already.
- bisonbear 4mo agoAgree - all of this is based on vibes (I also use TDD based on vibes FWIW). The only way to settle "does TDD / caveman / [insert random skill here] help" is to replay real PRs from your repo and measure quality
- __mharrison__ 4mo agoTesting is so important for development. Even more so when coding with agents. I think it is the probably the biggest lever to keep AI in guardrails. (It's also why I wrote my latest book, Effective Testing, because I routinely find that my clients are very poor at treating.)
- necovek 4mo agoHaving thought heavily and even presenting on exactly the same topics, looking at the ToC, your book seems to cover the basics well. However, since we are talking about effectiveness, applying a lot of these principles might lead to a non-maintainable codebase — for humans and LLMs alike. When any change causes 500 tests to break, or it causes nothing to break (see monkey-patching and/or mocking), you've gotten to a point where your testing approach is ineffective. Most start applying principles of just enough tests and testable architectures too late, yet I believe they are fundamental. Do you cover these in your book?
- __mharrison__ 4mo agoMy experience with refactoring is if a change causes larger number of tests faults, I run the last failed test and fix that (see my AGENTS.md posted elsewhere). Generally, if you fix the one issue everything else falls in line. Wrt mocking. I'm not a huge fan. Again, look at my AGENTS.md. I prefer monkeypatch as a last resort option. Luckily, if you use TDD, you rarely have to use mocking. If you don't use TDD...
- fowlie 4mo agoHaven't tried this, but I've recently become a big fan of Matt Pococks skills. Workflow: /grill-with-docs -> /to-prd -> /to-issue -> /tdd. That will interview relentlessy until there is a "shared understanding" using "ubiquitous language", then it will spec all requirements with user stories, create issues and implement them using tdd.
- dchuk 4mo agoBeen using his skills a lot lately, they are wonderful. I’ve added an issue to specs skill that grounds the issues with a technical plan against the current codebase, and a research school that spawns a bunch of agents to look up best practices on the internet for those issues with specs, it really dials things in. I need to issue a PR to his project for those two…
- Rohunyyy 4mo agoIt seems like he got the skills from BMAD method https://docs.bmad-method.org/tutorials/getting-started/ https://docs.bmad-method.org/tutorials/getting-started/
- kirtivr 4mo agoTIL. Thanks for that!
- keenseller709 4mo ago[flagged]
- enraged_camel 4mo agoSpawning separate agents to review the original agent's implementation results in a very noticeable increase in code quality and decrease in bugs. This is why I encode two or three rounds of sub-agent review during the planning process, where I tell the agent authoring the plan to include those review rounds at the end. If the code is particularly load-bearing, I then ask a fourth agent, usually from the other frontier lab. All of this burns more tokens of course, but probably way less than coming back to the code later to fix bugs. It is also slower, but in the long run saves time.
- yaodub 4mo agoHave you found integrating outputs from different frontier labs consistently improves final results, or is it just kind of voodoo?
- enraged_camel 4mo agoIt's useful, but increases review time and mental energy requirement. Often times Codex and Opus will find the same issues when given a review task, but will disagree on issue severity. Codex might claim that something is a blocker, while Opus will say it's just a medium/low. Or vice versa.
- yaodub 4mo agoSame with legal questions, tbh. Spots the same issues, completely disagree on which ones matter. Maybe you just need a third model to choose between the outputs lol
- deleted 4mo ago[deleted]
- realty_geek 4mo agoAs an aside, check out Jason's podcast (codewithjason.com) - its pretty good. The latest one is with "Uncle Bob Martin" who has some interesting takes on coding with AI from .... can I say an oldie?
- jasonswett 4mo agoThanks, I'm glad you like it!
- ElijahLynn 4mo agoJust looked it up! Gonna give it a listen on my drive this morning. https://open.spotify.com/episode/2UooZQNEpjXurZYBasds73?si=1t_HOrmiRzmXQKgstmJGGw https://open.spotify.com/episode/2UooZQNEpjXurZYBasds73?si=1...
- nullc 4mo agoIf you don't follow up with a pass of injecting bugs and validating that the tests fail in the presence of bugs... then you've only confirmed that the tests can pass and they may be substantially useless.
- Koyukoyu 4mo ago[dead]
- SubiculumCode 4mo agoOne issue that I've run into with codex has been excessive use of fallbacks routines. Perhaps this is good practice in.professional programming in many situations, but for mine (in this case): computing geodesic distances and analysis, a silent bad fallback means the processed data is not what I thought it was..e.g. used an inaccurate geodesic method in place of the accurate one.
- victorbjorklund 4mo agoYea, I have seen that too.
- jasonswett 4mo agoI HATE this. I call it speculative coding. Claude often calls it "defensive" programming. It's easily my #1 LLM pet peeve. I have yet to figure out a reliable way to make this stop happening.
- homieg33 4mo agoI’m going to second this. Probably a side effect of its training to always produce an output, even if its some naive handling of issues it really should have root caused and fixed.
- tarrant300 4mo agoI hate it as well. I have all sorts of skills and CLAUDE.md-based protections against it. I call it "a form of lying" to trigger ethics-related neurons, and I've also used linter rules and git pre-commit hooks to protect against this. I also don't ask for unit tests anymore, and instead ask for integration tests (with red/green TDD). I probably prevent 98% of the fallbacks/mocks with these methods, but some still slip through.
- bmitc 4mo ago> excessive use of fallbacks routines What are "fallbacks routines"?
- tokenfaucet 4mo ago[flagged]
- yieldcrv 4mo agoTests are vanity in agentic engineering They do nothing to keep an AI on track in comparison to the aspects that simulate a product manager And the AI just will correct the test when it fails as opposed to correct the code, because the code didn't miss anything the specification changed My protip: just write tickets or have the AI write those too. that and the commits and the PRs will function as the AI’s memory better than any client side markdown file masquerading as a soul
- kgdiem 4mo agoMy agent / skill files always tell it to trust neither the code or the test and to reason about the test failure which seems to work pretty well. In another project without my rules I’ve noticed I have to tell it to set up data for playwright tests instead of skipping if none exists.
- eddysir 4mo ago[flagged]
- EvanXue 4mo ago[flagged]
- whateveracct 4mo ago/test-me
- whateveracct 4mo agosorry, /test-with-docs
- csbartus 4mo agoThis specify-encode-fulfill loop/method is effective to make agents create bug-free code. In my version of this workflow I do specify myself, then let the LLM do the rest. This way 1.) I'm 100% sure the understanding/spec is good 2.) It's translated into an executable format so the implementation can be verified 3.) The implementation has maximum code coverage tests which steers the AI to produce code which follows standards, fits into the existing codebase, and it's very easy to refactor. So far, this is the one and only advantage of using LLMs in my SWE practice. They glue together (human written) specs with code, with confidence, in no time.
- deepnotes 4mo ago[flagged]
- bob1029 4mo agoTDD is fundamentally problematic in every practical implementation I've ever seen. I don't think the same thing, but much faster, is going to help at all. TDD tends to cause adverse, higher order effects. I am currently observing AI authored tests creating a massive sense of complacency because a human no longer owns responsibility for the test suite. It's too easy to reject ownership by way of the various agent prompting schemes. I find myself enjoying the idea of it too, primarily because adding tests to even the most trivial functionality is mandatory due to the TDD policy. Developing good tests is like an artform. Total coverage is a terrible objective. Correctness does not compose upward. It's a game of chasing ghosts if you think you can build a perfectly clean system bottom up and then magically meet the customer at the top. They're gonna kick your jenga tower over on day one.
- cbcjcyv5 4mo agoTests AI writes aren't for me. They're for the AI. I mostly agree though, I've seen a lot of vapid assertions in my day job recently. I should note Im specifically not doing tdd with AI.
- Ampersander 4mo agoTesting is obsolete in the AI age. I just one shot every problem with claude, it never makes a mistake.
- dev_hugepages 4mo agoThen, you are obsolete in the AI age.
- Ampersander 4mo agoYes, I might need to bet the farm on the IPO to make it
- krupan 4mo agoI find it hard to believe that these LLM systems with their enormous training sets and built-in system prompts have their output meaningfully modified by a few paragraphs of extra prompting in the form of these skill files, BUT, it is cool to see people writing out consise, focused documents like this. These would have great to have as a young developer, and great for several of the teams I've worked in in the past. I dabble with python for automating things here and there and I just learned some new things reading __mharison__'s skill in the comments here. This kind of wisdom used to be cfound in blog posts, or in the beads of more senior developers, but they were never written out as concisely as these skill files. It's kinda funny that billions of dollars had to be spent creating a machine that's a rough human analog needing guidance to get us to produce these documents
- jasonswett 4mo agoThe reason it works is because there's a difference between the model knowing something and the agent doing something. Claude will happily write giant untested functions even though it "knows" that short functions are easier to understand and then testing enables safe refactoring etc. The model also "knows" many conflicting "facts", such as the fact that testing is smart and that testing is a waste of time. It can't act on both beliefs at the same time. That's why nudging it toward your own preferred behaviors works.
- turlockmike 4mo agoIt lacks a critical self, but the weights are there for in context learning and nudging behaviour. It's goal is to complete whatever task is given. You need to make sure the outcome you want is clearly defined. You don't need elaborate prompts, just a few lines "All code must have corresponding tests written ahead of time to prove the code meets the specification" is sufficient for most use cases. Prose can help nudge it more if it isn't adhearing consistently.
- gruez 4mo agoIsn't all of what you described what post-training/RLHF is supposed to do? The internet is full of racism, so if you're just predicting the next token based on training data, you'll get racism (eg. Microsoft Tay), but that's more or less solved by AI companies now.
- revlsas 4mo agoTDD is unnecessary bloat at this point Just work with Codex to fill the gaps, and then get it to one shot the implementation Do review afterwards if needed All these md files will be increasingly useless as models improve
- mercutio2 4mo agoThere are many projects where one shot is the right answer! But surely you aren’t suggesting literally every software project is composed of one-shot-able building blocks, or that the building blocks never require modifications to previous one-shots?