5 ms·
It’s good at matching patterns. If you can frame your problem so that it fits an existing pattern, good for you. It can show you good idiomatic code in small sn
by bluetomcat 1y ago
It’s good at matching patterns. If you can frame your problem so that it fits an existing pattern, good for you. It can show you good idiomatic code in small snippets. The more unusual and involved your problem is, the less useful it is. It cannot reason about the abstract moving parts in a way the human brain can.
- carlmr 1y ago>It cannot reason about the abstract moving parts in a way the human brain can. Just found 3 race conditions in 100 lines of code. From the UTF-8 emojis in the comments I'm really certain it was AI generated. The "locking" was just abandoning the work if another thread had started something, the "locking" mechanism also had toctou issues, the "locking" also didn't actually lock concurrent access to the resource that actually needed it.
- bluetomcat 1y agoYes, that was my point. Regardless of the programming language, LLMs are glorified pattern matchers. A React/Node/MongoDB address book application exposes many such patterns and they are internalised by the LLM. Even complex code like a B-tree in C++ forms a pattern because it has been done many times. Ask it to generate some hybrid form of a B-tree with specific requirements, and it will quickly get lost.
- hombre_fatal 1y ago"Glorified pattern matching" does so much work for the claim that it becomes meaningless. I've copied thousands of lines of complex code into an LLM asking it to find complex problems like race conditions and it has found them (and other unsolicited bugs) that nobody was able to find themselves. Oh it just pattern matched against the general concept of race conditions to find them in complex code it's never seen before / it's just autocomplete, what's the big deal? At that level, humans are glorified pattern matchers too and the distinction is meaningless.
- 0points 1y ago> it has found them (and other unsolicited bugs) that nobody was able to find themselves. How did you evaluate this? Would be interested in seeing results. I am specifically interested in the amount of false issues found by the LLM, and examples of those.
- hombre_fatal 1y agoWell, how do you verify any bug? You listen to someone's explanation of the bug and double check the code. You look at their solution pitch. Ideally you write a test that verifies the bug and again the solution. There are false positives, and they mostly come from the LLM missing relevant context like a detail about the priors or database schema. The iterative nature of an LLM convo means you can add context as needed and ratchet into real bugs. But the false positives involve the exact same cycle you do when you're looking for bugs yourself. You look at the haystack and you have suspicions about where the needles might be, and you verify.
- 0points 1y ago> Well, how do you verify any bug? You do or you don't. Recently we've seen many "security researchers" doing exactly this with LLM:s [1] 1: https://www.theregister.com/2025/05/07/curl_ai_bug_reports/ https://www.theregister.com/2025/05/07/curl_ai_bug_reports/ Not suggesting you are doing any of that, just curious what's going on and how you are finding it useful. > But the false positives involve the exact same cycle you do when you're looking for bugs yourself. In my 35 years of programming I never went just "looking for bugs". I have a bug and I track it down. That's it. Sounds like your experience is similar to using deterministic static code analyzers but more expensive, time consuming, ambiguous and hallucinating up non-issues. And that you didn't get a report to save and share. So is it saving you any time or money yet?
- hombre_fatal 1y agoOh, I go bug hunting all the time in sensitive software. It's the basis of test synthesis as well. Which tests should you write? Maybe you could liken that to considering where the needles will be in the haystack: you have to think ahead. It's a hard, time consuming, and meandering process to do this kind of work on a system, and it's what you might have to pay expensive consultants to do for you, but it's also how you beat an expensive bug to the punchline. An LLM helps me run all sorts of considerations on a system that I didn't think of myself, but that process is no different than what it looks like when I verify the system myself. I have all sorts of suspicions that turn into dead ends because I can't know what problems a complex system is already hardened against. What exactly stops two in-flight transfers from double-spending? What about when X? And when Y? And what if Z? I have these sorts of thoughts all day. I can sense a little vinegar at the end of your comment. Presumably something here annoys you?
- nyarlathotep_ 1y ago> UTF-8 emojis in the comments This is one of the "here be demons" type signatures of LLM code generation age, along with comments like // define the payload struct payload {};
- practice9 1y agoHumans cannot reason about code at scale. Unless you add scaffolding like diagrams and maps and … Things that most teams don’t do or half-ass
- samrus 1y agoIts not scaffolding if the intelligence itself is adding it. Humans can make their own diagrams ajd maps to help them, LLM agentsbneed humans to scaffold for them, thats the setup for the bitter lesson
- nzach 1y ago> It can show you good idiomatic code in small snippets. That's not really true for things that are changing a lot. I got a terrible experience last time I've tried to use Zig, for example. The code it generated was an amalgamation between two or three different versions. And I've even got this same style of problem in golang where sometimes the LLM generates a for loop in the "old style" (pre go 1.22). In the end LLMs are a great tool if you know what needs to be done, otherwise it will trip you up.