6 ms·
Micro-agent: make an AI write code until it passes an unit test
- nrabulinski 2y agoEnterprise developer from hell (https://fsharpforfunandprofit.com/posts/property-based-testing/ https://fsharpforfunandprofit.com/posts/property-based-testi...) but as a CLI tool
- mrdevlar 2y agoNeat, thanks for this, I enjoyed this viewpoint as it might help systems like this actually be able to build something reasonable.
- BiteCode_dev 2y agoIndeed, and it will force you to write good tests :)
- amatic 2y agoThis sounds amazing! Are there any metrics on how often different models pass tests? Has someone used a similar process to finetune an LLM?
- benve 2y agoA very sad future awaits us if a developer's only job is to write tests
- _flux 2y agoMaybe in the future developers will be able to write just specifications.
- jacamera 2y agoThat's exactly what we do now.
- _flux 2y agoArguably we write instructions: instead of writing out the problem and what the solution looks like, we describe a set of steps we go through—and if those steps are incorrect, there's nothing to compare against, because that was what we called the "specification". Whether there's a difference there is in the eye of the beholder, but it does look like that specification languages such as TLA+/PlusCal/Squint or Alloy, or theorem proving languages like Coq (to be renamed Rocq) or Lean look a lot different from the likes of C, JavaScript or even Haskell.
- benve 2y agothis already happens, in many companies mid-level managers write the specifications (ambiguously) and other low-cost people implement them. And when things don't work they call people like me, to try to understand the performance problems of something poorly defined and worse written.
- deleted 2y ago[deleted]
- tiborsaas 2y agoIt's worse, it starts by writing the test so you have to verify if the tests work :)
- philote 2y agoThat's what I don't get about this. Instead of writing code that may or may not be correct, it's writing tests that may or may not be correct.
- TuringNYC 2y agoPerhaps a minority opinion, but i LOVE writing tests. I write them before I write the code, it is like playing chess with yourself.
- benve 2y agoI appreciate that your workflow is so linear. I often write tests, then the implementation, then I realize that the tests need to be corrected, then I change the implementation, then I change the tests, then I add other tests etc... etc... I don't really like maintaining tests, it's often a lot of code that needs to be understood and changed carefully
- colechristensen 2y agoWhy? It will evolve into a slightly higher level language where the compiler is an ML model. Was it a tragedy when developers mostly didn’t have to write assembly any more?
- selcuka 2y ago> Was it a tragedy when developers mostly didn’t have to write assembly any more? It wasn't, but for starters compilers have always been generally deterministic. I'm not saying that this is completely useless (I personally think code completion tools such as GitHub CoPilot are fantastic), but it is still early to compare it to a compiler.
- benve 2y agoI think it's different... I like high level languages, but this is not a programming language, this is a technique for writing tests in an existing language and leaving the implementation to the AI. I like programming for problem solving, I don't really like writing tests, but that's personal taste, a lot of people like to just use PowerPoint and Jira and tell others what they need to implement, but these people are not software developers.
- root_axis 2y agoI've tried these LLM "code from test" things (and vice-versa) dozens of times over the last couple of years... they're not even close to approaching being practical.
- awwaiid 2y agoI knew declarative languages would eventually win!
- autonomousErwin 2y agoReally it's just validator code instead of feature code. I think this is the only realistic way forward for production level code written by AI, don't ask it to write code - ask it to pass your validation tests. Essentially, everyone becomes a red team member trying to think of clever ways they can outwit the AI's code which I for one think this is going to be a lot of fun in the future - though we're still quite a way from there yet!
- shakna 2y agoMost LLMs struggle to even do a null check. Could this check for those kinds of glaring security holes?
- jejeyyy77 2y agoadd a test for it?
- shakna 2y agoKinda hard to right a test that a value that is null-checked, when that value may never actually be returned. For example, have a C function that reads in a file and returns you a string? The string can be checked that malloc actually succeeded, but how do you check that the file actually opened?
- jejeyyy77 2y agowhat?
- shakna 2y agoEvery time you call fopen, you need to do a null check. Every single time. You also need to check that fclose is there to match the call, every time. Writing a test for that, when it is generally just a call within the function you want to test, isn't really possible. It's not there in the arguments, or the return value, of the function. How do you check for the right checks in a function expected to do something like this: int foo() { FILE *fp = fopen("test.in", "r"); if(!fp) { return -1; } for(int i = 0; i < NUM; i ++){ if(matcher(fp, i)){ fclose(fp); return i; } } fclose(fp); return 0; }
- jejeyyy77 2y agoyou write 2 unit tests for your function... one with a test file that exists, and one with a fake file path. Assert not null and null respectively..
- joseferben 2y agoi found that the feedback loop between llm and test suite works really well, especially with sonnet 3.5 i wrote a similar tool the other day: https://github.com/joseferben/makeitpass https://github.com/joseferben/makeitpass it can make all kinds of commands pass by checking stdout/stderr and it’s language agnostic (you need npx to run makeitpass)
- jeylum22 2y ago[flagged]
- itchyjunk 2y agoCan you elaborate on why you think it won't? That might be more valuable use of everyone's time instead of the binary X will / won't happen stance comment.
- deleted 2y ago[deleted]
- bangaladore 2y agoAt least in my experience, possibly due to context limitations, or just architecture, SOTA LLMs aren't particularly good at iterating as they tend to loop back around to similar results with bad logic / errors
- ola_esponel 2y agoDoes this work with any llm?
- deleted 2y ago[deleted]
- Arubis 2y agoI'd much sooner accept and commit implementation code written by an AI against unit tests written by a human than the reverse.