7 ms·
I always see these reports about how much better AI is than humans now, but I can't even get it to help me with pretty mundane problem solving. Yesterday I gave
by mrdependable 1y ago
I always see these reports about how much better AI is than humans now, but I can't even get it to help me with pretty mundane problem solving. Yesterday I gave Claude a file with a few hundred lines of code, what the input should be, and told it where the problem was. I tried until I ran out of credits and it still could not work backwards to tell me where things were going wrong. In the end I just did it myself and it turned out to be a pretty obvious problem.
The strange part with these LLMs is that they get weirdly hung up on things. I try to direct them away from a certain type of output and somehow they keep going back to it. It's like the same problem I have with Google where if I try to modify my search to be more specific, it just ignores what it doesn't like about my query and gives me the same output.
- simonw 1y agoLLMs are difficult to use. Anyone who tells you otherwise is being misleading.
- __loam 1y ago"Hey these tools are kind of disappointing" "You just need to learn to use them right" Ad infinitum as we continue to get middling results from the most overhyped piece of technology of all time.
- simonw 1y agoThat's why I try not to hype it.
- mvdtnz 1y agoYou're the biggest hype merchant for this technology on this entire website. Please.
- simonw 1y agoI've been banging the drum about how unintuitive and difficult this stuff is for over a year now: https://simonwillison.net/2025/Mar/11/using-llms-for-code/ https://simonwillison.net/2025/Mar/11/using-llms-for-code/ I'm one of the loudest voices about the so-far unsolved security problems inherent in this space: https://simonwillison.net/tags/prompt-injection/ https://simonwillison.net/tags/prompt-injection/ (94 posts) I also have 149 posts about the ethics of it: https://simonwillison.net/tags/ai-ethics/ https://simonwillison.net/tags/ai-ethics/ - including one of the first high profile projects to explore the issue around copyrighted data used in training sets: https://simonwillison.net/2022/Sep/5/laion-aesthetics-weeknotes/ https://simonwillison.net/2022/Sep/5/laion-aesthetics-weekno... One of the reasons I do the "pelican riding a bicycle" thing is that it's a great way to deflate the hype around these tools - the supposedly best LLM in the world still draws a pelican that looks like it was done by a five year old! https://simonwillison.net/tags/pelican-riding-a-bicycle/ https://simonwillison.net/tags/pelican-riding-a-bicycle/ If you want AI hype there are a thousand places on the internet you can go to get it. I try not to be one of them.
- __loam 1y agoThe prompt injection articles you wrote really early in the tech cycle were really good and I appreciated them at the time.
- andai 1y agoCould a five year old do it in XML (SVG)? Could an artist? In one shot?
- annjose 1y agoI agree - the content you write about LLMs is informative and realistic, not hyped. I get a lot of value from it, especially because you write mostly as stream of consciousness and explains your approach and/or reasoning. Thank you for doing that.
- maleldil 1y agoIt's true that simonw writes a lot about LLMs, but I find his content to be mostly factual. Much of it is positive, but that doesn't mean it's hype.
- JohnKemeny 1y agoUh... You don't do anything but hype them. I literally don't know who anyone on HN are except you and dang, and you're the one that constantly writes these ads for your LLM database product.
- simonw 1y agoI think you and I must have different definitions of the word "hype". To me, it means LinkedIn influencers screaming "AGI is coming!", "It's so over", "Programming as a career is dead" etc. Or implying that LLMs are flawless technology that can and should be used to solve every problem. To hype something is to provide a dishonest impression of how great it is without ever admitting its weaknesses. That's what I try to avoid doing with LLMs.
- bluefirebrand 1y ago> without ever admitting its weaknesses I don't think this part is necessary "To hype something is to provide a dishonest impression of how great it is" is accurate. Marketing hype is all about "provide a dishonest impression of how great it is". Putting the weaknesses in fine print doesn't change the hype Anyways I don't mean to pile on but I agree with some of the other posters here. An awful lot of extremely pro-AI posts that I've noticed have your name on them I don't think you are as critical of the tech as you think you are. Take that for what you will
- tzumaoli 1y agoalso "They will get better in no time"
- simonw 1y agoThat one's provably correct. Try comparing 2023-era GPT-3.5 with 2025's best models.
- xboxnolifes 1y agoIt's not provably correct if the comment is made toward 2025 models.
- simonw 1y agoGemini 2.5 came out just over two weeks ago (25th March) and is a very significant improvement on Gemini 2.0 (5th February), according to a bunch of benchmarks but also the all-important vibes.
- torginus 1y agoLLMs are a casino. They're probabilistic models which might come up with incredible solutions at a drop of a hat, then turn around and fumble even the most trivial stuff - I've had this same experience from GPT3.5 to the latest and greatest models. They come up with something amazing once, and then never again, leading me to believe, it's operator error, not pure dumb luck or slight prompt wording that lead me to be humbled once, and then tear my hair out in frustration the next time. Granted, newer models tend to do more hitting than missing, but it's still far from a certainty that it'll spit out something good.
- pants2 1y agoIn my experience, most people who say "Hey these tools are kind of disappointing" either refuse to provide a reproducible example of how it falls short, or if they do, it's clear that they're not using the tool correctly.
- __loam 1y agoAd infinitum
- sksxihve 1y agoI'd love to see a reproducible example of these tools producing something that is exceptional. Or a clear reproducible example of using them the right way. I've used them some (sorry I didn't make detailed notes about my usage, probably used them wrong) but pretty much there are always subtle bugs that if I didn't know better I would have overlooked. I don't doubt people find them useful, personally I'd rather spend my time learning about things that interest me instead of spending money learning how to prompt a machine to do something I can do myself that I also enjoy doing. I think a lot of the disagreements on hn about this tech is that both sides are mostly on the extremes of either "it doesn't work and at and is pointless" or "it's amazing and makes me 100x more productive" and not much discussion about the mid-ground of it works for some stuff and knowing what stuff it works well on makes it useful but it won't solve all your problems.
- doug_durham 1y agoWhy are you setting the bar at "exceptional". If it means that you can write your git commit messages more quickly and with fewer errors then that's all the payoff most orgs need to make them worthwhile.
- bluefirebrand 1y ago> Why are you setting the bar at "exceptional" Because that is how they are being sold to us and hyped > If it means that you can write your git commit messages more quickly and with fewer errors then that's all the payoff most orgs need to make them worthwhile. This is so trivial that it wouldn't even be worth looking into, it's basically zero value
- TeMPOraL 1y agoNo, it's just you and yours. IDK, maybe there's a secret conspiracy of major LLM providers to split users into two groups, one that gets the good models, and the other that gets the bad models, and ensure each user is assigned to the same bucket at every provider. Surely it's more likely that you and me got put into different buckets by the Deep LLM Cartel I just described, than it is for you to be holding the tool wrong.
- KronisLV 1y ago> "Hey these tools are kind of disappointing" > "You just need to learn to use them right" Admittedly, the first line is also my reaction to the likes of ASM or system level programming languages (C, C++, Rust…) because they can be unpleasant and difficult to use when compared to something that’d let me iterate more quickly (Go, Python, Node, …) for certain use cases. For example, building a CLI tool in Go vs C++. Or maybe something to shuffle some data around and handle certain formatting in Python vs Rust. Or a GUI tool with Node/Electron vs anything else. People telling me to RTFM and spend a decade practicing to use them well wouldn’t be wrong though, because you can do a lot with those tools, if you know how to use them well. I reckon that it applies to any tool, even LLMs.
- zamadatix 1y agoI also think LLMs are more difficult to use for most tasks than is often flouted myself but I don't really jive with statements like "Anyone who tells you otherwise is being misleading". Most of the time I find they are just using them in a very different capacity.
- simonw 1y agoI intended those words to imply "being misleading even if they don't know they are being misleading" - I made a better version of that point here: https://simonwillison.net/2025/Mar/11/using-llms-for-code/ https://simonwillison.net/2025/Mar/11/using-llms-for-code/ > If someone tells you that coding with LLMs is easy they are (probably unintentionally) misleading you. They may well have stumbled on to patterns that work, but those patterns do not come naturally to everyone.
- slig 1y agoWas that on 3.7 Sonnet? I feel it's a lot worse than 3.5. If you can, try again but on Gemini 2.5.
- avandekleut 1y agoI'm glad I'm not the only one that has found 3.5 to be better than 3.7.
- johnisgood 1y agoWhen did 3.7 come out? I might have had the same experience. I think I have been using 3.5 with success, but I cannot remember exactly. I may have not used 3.7 for coding (as I had a couple of months break).
- simonw 1y ago3.7 came out on 24th February. My notes from that release: https://simonwillison.net/2025/Feb/24/claude-37-sonnet-and-claude-code/ https://simonwillison.net/2025/Feb/24/claude-37-sonnet-and-c... and https://simonwillison.net/2025/Feb/25/llm-anthropic-014/ https://simonwillison.net/2025/Feb/25/llm-anthropic-014/
- johnisgood 1y agoI will have to check, but apparently I have been using 3.5 with success, then. I will give 3.7 a try later, I hope it is really not that much worse, or is it? :(
- mrdependable 1y agoThis was 3.7. I did give Gemini a shot for a bit but it couldn’t do it either and the output didn’t look quite as nice. Also, I paid for a year of Claude so kind of feel stuck using it now. Maybe I will give 3.5 a shot next time though.
- namaria 1y agoIt's overfitting. Some people say they find LLMs very helpful for coding, some people say they are incredibly bad. I often see people wondering if the some coding task is performed well or not because of availability of code examples in the training data. It's way worse than that. It's overfitting to diffs it was trained on. "In other words, the model learns to predict plausible changes to code from examples of changes made to code by human programmers." https://arxiv.org/abs/2206.08896 https://arxiv.org/abs/2206.08896
- simonw 1y ago... which explains why some models are better at code than others. The best coding models (like Claude 3.7 Sonnet) are likely that good because Anthropic spent an extraordinary amount of effort cultivating a really good training set for them. I get the impression one of the most effective tricks is to load your training set up with as much code as possible that has comprehensive automated tests that pass already.
- torginus 1y agoI've often experienced that I had what I thought an obscure and very intellectually challenging coding problem, and after prompting the LLM, it basically one-shotted it. I've been profoundly humbled by the the experience, but then it occurred to me that what I thought to be an unique problem has been solved by quite a few people before and the model had plenty of references to pull from.
- zifpanachr23 1y agoDo you have any examples?
- suddenlybananas 1y ago[flagged]
- torginus 1y ago
- lispisok 1y agoThe PR articles and astroturfing will continue until investors get satisfactory returns on their many billions dumped into these things.