3 ms·
Yesterday I got AI (a sota model) to write some tests for a backend I'm working on. One set of tests was for a function that does a somewhat complex SQL query t
by AstroBen 9mo ago
Yesterday I got AI (a sota model) to write some tests for a backend I'm working on. One set of tests was for a function that does a somewhat complex SQL query that should return multiple rows
In the test setup, the AI added a single database row, ran the query and then asserted the single added row was returned. Clearly this doesn't show that the query works as intended. Is this what people are referring to when they say AI writes their tests?
I don't know what to call this kind of thinking. Any intelligent, reasoning human would immediately see that it's not even close to enough. You barely even need a coding background to see the issues. AI just doesn't have it, and it hasn't improved in this area for years
This kind of thing happens over and over again. I look at the stuff it outputs and it's clear to me that no reasoning thing would act this way
- elcritch 9mo agoAs a counter I’ve had OpenAI Codex and Claude Code both catch logic cases I’d missed in both tests and codes. The tooling in the Code tools is key to useable LLM coding. Those tools prompt the models to “reason” whether they’ve caught edge cases or met the logic. Without that external support they’re just fancy autocompletes. In some ways it’s no different than working with some interns. You have to prompt them to “did you consider if your code matched all of the requirements?”. LLMs are different in that they’re sorta lobotomized. They won’t learn from tutoring “did you consider” which needs to essentially be encoded manually still.
- AstroBen 9mo ago> As a counter I’ve had OpenAI Codex and Claude Code both catch logic cases I’d missed in both tests and codes That has other explanations than that it reasoned its way to the correct answers. Maybe it had very similar code in its training data This specific example was with Codex. I didn't mention it because I didn't want it to sound like I think codex is worse than claude code I do realize my prompt wasn't optimal to get the best out of AI here, and I improved it on the second pass, mainly to give it more explicit instruction on what to do My point though is that I feel these situations are heavily indicative of it not having true reasoning and understanding of the goals presented to it Why can it sometimes catch the logic cases you miss, such as in your case, and then utterly fail at something that a simple understanding of the problem and thinking it through would solve? The only explanation I have is that it's not using actual reasoning to solve the problems
- grayhatter 9mo ago> In some ways it’s no different than working with some interns. You have to prompt them to “did you consider if your code matched all of the requirements?”. I really hate this description, but I can't quite filly articulate why yet. It's distinctly different because interns can form new observations independently. AIs can not. They can make another guess at the next token, but if it could have predicted it the 2nd time, it must have been able to predict it the first, so it's not a new observation. The way I think through a novel problem results in drastically different paths and outputs from an LLM. They guess and check repeatedly, they don't converge on an answer. Which you've already identified > LLMs are different in that they’re sorta lobotomized. They won’t learn from tutoring “did you consider” which needs to essentially be encoded manually still. This isn't how you work with an intern (unless the intern is unable to learn).
- expedition32 9mo agoThe whole point about an intern is that after a month they can act without coaching. Humans do actually learn- it is quite a revelation to see a child soak up data like an AI on steroids.
- grayhatter 9mo ago> Is this what people are referring to when they say AI writes their tests? yes > Any intelligent, reasoning human would immediately see that it's not even close to enough. You barely even need a coding background to see the issues. [nods] > This kind of thing happens over and over again. I look at the stuff it outputs and it's clear to me that no reasoning thing would act this way and yet there're so many people who are convinced it's fantastic. Oh, I made myself sad. The larger observation about it being statistical inference, rather than reason... but looks to so many to be reason is quite an interesting test case for the "fuzzing" of humans. In line with why do so many engineers store passwords in clear text? Why do so many people believe AI can reason?
- dorgo 9mo agoSounds like the AI was not dumb but lazy. I do it similarly when I don't feel like doing it.