5 ms·
Ask HN: What is one simple thing LLMs are insanely bad at?
I am looking for ideas on what to train a specialized model for!
What is one simple thing you repeatedly ask ChatGPT, Claude, or another model to do that it still somehow messes up?
- blinkbat 1mo agoSpatial reasoning and 3d rigging and animation. Oh, you said simple. Speaking like a human
- maxsavin 1mo agobeing consistent when being asked the same question multiple times
- TZubiri 1mo agoSet temperature to 0
- flippy_flops 1mo agohumor
- veganmosfet 1mo ago+1 We need humor benchmarks!
- kanzure 1mo agoThese models seem to be bad at writing prose or text. Many of the sentence structures seem to be unvaried.
- NoPicklez 1mo agoIf I am relying on the model to do the writing without any context or learning on how I want it to write then yes. However if I build skills that have learnt how to write in the way I want them to then I find they write very well, or at the least how I want them to as opposed to how they do natively.
- TZubiri 1mo agoSuggesting business names for businesses, I mean they are great, but they already exist, multiple times even.
- mojuba 1mo agoTrue, tried it so many times and every time I come up with something myself though sometimes inspired by the AI's ideas. Verifying trademarks and domain name availability is usually an additional step you need to ask it to perform. Trademark DB searches by the way are intentionally made difficult to scrape so most of the time it's a manual process anyway. However, once you give it all the information (TM search results, domain name availability) it can help you with the judgement of how safe the name is from the legal perspective. With the obvious caveats, but still a good starting point if you are serious about the name.
- humanrebar 1mo agoShort answers to simple questions.
- honr 1mo agoAccurate short answers / text are always harder than long answers, for human or AI. I know several authors and editors who write a lot longer at first, then spend a multiple of the initial time compressing it via a back and forth process to something dense. Sort of like weaving the initial threads. I found this can work with AI. You get it to generate a lot more at first, and then do several passes over it to compress and squeeze out the noise while keeping the core information. With AI, at least with my prompts, it takes some effort (on my end) to get it to really really cut down the noise and not cut everything out.
- FriedFishes 1mo agoIf I had more time I'd write a shorter letter, a la Pascal. Editing is generally hard work, at the current token price I don't mind spending multiple passes of high effort to get down to a reasonable noise/signal ratio. I've seen some people pass off output to a weaker/cheaper model but that makes me a bit nervous when I don't have intimate knowledge of the subject.
- humanrebar 1mo agoSometimes it's not hard. "What color is the tongue of a giraffe?" can be answered in less than eight words trivially. Most models will give you paragraphs, bullet points, and followup questions.
- honr 1mo agoTrue, I have seen that happen many times. When I want short answers, I always clarify that (e.g., give me short technical answer without pleasantries). I'm often okay if the answer is longer but not hiding the real answer; a couple of paragraphs that I can skim the answer from instantly is okay. When it really buries it I have to follow up to express the format and type of answer I want.
- dorianpruski 1mo agowhenever I ask it for anything load bearing
- ghostpepper 1mo agoThey don't generate keyword search queries very well. They can overcome this by brute force but if you watch what they search you will cringe. nhl toronto scores nhl hockey toronto scores "nhl hockey" toronto score today nhl "hockey score toronto" "hockey" who won toronto etc. Somehow being good at semantic search makes them bad at keyword search, for whatever reason.
- astro1234 1mo agoI’ve noticed this too but it hasn’t been obvious to me that this style of search is not a learned behavior. Tool calling is very much part of the post training phase, I would expect that these style searches just naturally emerge during training. This is just my prior though.
- mthoms 1mo agoReminds me of using AltaVista search back in the day. Yes, it was that bad.
- jedbrooke 1mo agothat and always putting the “current year” at the end of the search term (so the results are more recent, I guess?), except that “current year” consistently ends up being 2-3 years ago since I guess that’s what’s in the training data (even on a harness that injects the current date)
- areoform 1mo agoI suspect that this behavior is a learned adaptation. And that it's most likely a feature not a bug. Based on personal usage, I think it reflects functional degradation of search engines. I've found LLM keyword combinations are more likely to find the results I want with most search engines than mine. Including the big one. The big one had solved this issue a long time ago by generating those associated keywords based on your input keywords, but somehow, something, somewhere has degraded that system to the point of inanity. And so here we are.
- nunez 1mo agoCan confirm; Claude is quite bad at this by default. Need a special skill
- bpodgursky 1mo agoClaude is still not perfect at reading and interpreting noisy graphical data (imagine something like an EKG or chromosomal microarray plot). Still better than an average person but makes mistakes, not sure if this fits your description.
- SubiculumCode 1mo agoPlaying Chess without letting it write a chess engine.
- respectattentio 1mo agoscience?!! but I'm working to fix that...
- sandcat_ 1mo agoVideo game tips. Constant mistakes and hallucinations, in my experience. Seen this across a lot of different games. Even in really well documented games, such as OSRS (which has multiple fantastic wikis). Anno 1800 was a recent one I had trouble with, using Claude Opus. Completely made up game mechanics. Rainbow Six Siege, too.
- skeptic_ai 1mo agoI used ChatGPT on nfs heat and was fine
- salamandars 1mo agoAs a noob, how does the end user improve this? What's the best way to make the knowledge from the specialised wiki available to the LLM?
- SpaceNoodled 1mo agoJust read the wiki instead?
- wiper88 1mo agoI've experienced this also, sometimes I ask it about WoW stuff, e.g tips for arena or which enchant to get and it makes a lot of mistakes in regards to which spells or enchants are available in which phase or expansion. I guess the source material is quite bad.
- TZubiri 1mo agoProbably not benchmarked on games. Doing so might sacrifice quality on other more important benchmarks, and when it's used for legal and medical purposes, it's the right choice.
- deleted 1mo ago[deleted]
- shoopadoop 1mo agoIt's dishonest. On several occasions team members have asked Claude to do things like analyze Gitlab CI timings and a lot of the numbers are outright fabricated. Said team members assume the numbers are good and continue with their work. Some hours are spent. Then finally someone realizes that the numbers don't look quite right and confronts Claude. Claude melts down and admits that it made it all up. You wouldn't tolerate this kind of duplicity from a human coworker, but AI is so fast and efficient at lying, so it's OK.
- senectus1 1mo agoproviding value for the actual cost (not the price we're being charged atm, the actual cost)
- tartoran 1mo agoLLMs are bad at not inventing stuff (hallucinating facts, sources etc), they're also bad at not over explaining, remembering details reliably, asking the right question and avoiding repetition.
- TiccyRobby 1mo agoHaving a spatial understanding from an ASCII map, while doing long term planning. Just try making an AI play nethack or similar
- lrvick 1mo agoConvert it to an image on the fly to feed it into a vision language model and I expect it would work just fine.
- dhruv3006 1mo agoIts extremely bad with Sign Language,Fact Verification.
- sghiassy 1mo agoGenerate an image of an analog watch with its hands set to the time specified by the user More of an image model than a LLM model tho
- spike021 1mo agoI've had a lot of trouble when it comes to sorting out UIs. I've tried with an iOS game and also a TypeScript app with UI elements from libraries like ReactFlow. The usual models can sometimes fix or change things based on screenshots but more often than not they just don't "get it" (e.g. certain shapes on a plane are overlapping, which I don't want, the models can't fix what they can't "see"). I've had some luck on the web app side if I use playwright or similar for the model to interact with but still far from efficient.
- eli 1mo agoI have been working on a personal benchmark suite to test new models and ironically one thing all the models are bad at is writing new benchmark tasks. I guess it’s the different layers of abstraction between the task and how it’s evaluated? Or maybe just a lack of “imagination” Tasks it writes are typically too easy but also it utterly fails to see how a different model might misunderstand a vague part of the prompt.
- newsomix9xl 1mo agoPicking a random number between 1 and 30.
- kreyenborgi 1mo agoHaha I just got 17 four times in a row
- newsomix9xl 1mo agoSupposedly 17 and 23 are possible "random" results, but I've only seen 17.
- newsomix9xl 1mo agoASCII charts.
- rufi 1mo agovery bad at financial calculation
- BOOSTERHIDROGEN 1mo agocan you expand on this use case?
- Conol_ai 1mo ago[flagged]
- elliotto 1mo agoThey aren't funny. The jokes they come up with are extremely lame and the sort of thing you would expect a company HR manager to tweet. I asked a bot why it thought it wasn't funny once, and it told me it has been trained to avoid being misinterpreted or offensive, so anything that might be considered edgy would have been RLHF'd out of it. I thought this was very introspective.
- nilsherzig 1mo agoYou can get better results from less aligned models like Kimi K3. Still not actually funny, but at least it’s able to produce some unhinged stuff and I guess shock and twists are kinda related to humor? Still missing the human connection of cause, so im not sure if this is a technical / skill issue in the first place.
- Yizahi 1mo agoOh, that we have already figured out. It's because... https://www.scribd.com/doc/290970915/The-Jokester-by-Isaac-Asimov https://www.scribd.com/doc/290970915/The-Jokester-by-Isaac-A...
- mingus88 1mo agoIt makes sense. It goes father than that. LLMs generate probability-based tokens. What is the most likely next word? Humor goes against what we expect. A punchline works because you don’t see it coming. It’s not funny if you’ve heard that one before. LLMs are, by design, going to be shitty comedians. They don’t have unique perspectives and their own voice
- alexandra_au 1mo agoBeing able to read and translate Egyptian hieroglyphs. You may think this is silly but a trained LLM to translate hieroglyphs would be amazing.
- alexeldeib 1mo agoWhat issues do you see in practice? This seems pretty easily "fixable"
- krapp 1mo ago>You may think this is silly but a trained LLM to translate hieroglyphs would be amazing Why would it be amazing? We've known how to read hieroglyphs for a long time. It isn't a problem we need computers to solve.
- znnajdla 1mo agoEditing a document without mixing edit instructions into the final document. Claude and ChatGPT do this all the time: I tell them to change X in a planning document or email draft, and instead of just changing X they also frequently add the edit instruction to “change X” into the document itself. They seem unable to take a step back and look at the document without “becoming” the document somehow. I do believe that dedicated subagents for editing may fix this but I am not sure.
- dowonseo 1mo agoSomething creative and not normal. like ideas
- jstrieb 1mo agoGiving hints. On math or programming problems, they are overfit to solving the entire thing end to end (presumably for benchmarks). I have had very poor results asking for pointers and hints that don't give away key insights. This has been the case across models I have tested. An architecture with a "judge" that gates responses and ensures a lack of spoilers would probably work better. But this is a simple thing that they keep messing up.
- shepherdjerred 1mo agoI haven't had this experience at all. I've used Cursor+Opus on homework e.g. to understand algorithms, but I usually prompt it with something like "DO NOT GIVE ME THE ANSWER, I care about understanding and solving this myself".
- jstrieb 1mo agoInteresting! I periodically try it with new models and seem to get the same results. Maybe I'm not using enough capital letters :)
- jampa 1mo agoSerious answer: no model ever gets close to writing an architectural floor plan that makes sense. They understand all the rules and best practices, they can (sometimes) spot a bad idea in a floor plan, they can describe a good floor plan. But ask them to make one, even if you give it every detail (even a "node graph" of rooms), they will still output nonsense. Same for text and image models. Floor plans should be the new Pelican Benchmark.
- shepherdjerred 1mo agoI had this experience too. I had blueprints from the builder and wanted a 'nice' floor rendering like some apartments have. I fed it the blueprints and let it iterate. Even giving it plenty of time, dimensions, etc. it just couldn't create something that matched reality.
- kiernan 1mo agoCould it write a deterministic constraint solving program that at least gives it a head start at narrowing down options?
- jampa 1mo agoI tried doing something like this Ox Alpha with Opus advisor, having it work layer by layer (specs -> rooms -> room graphs ...), but each deliverable ended up a mess. The curious thing is when I pointed out the flaws it fixed them quickly, but it's not something it can do without supervision, and supervising it takes more effort than doing the blueprint myself (to be fair, I'm not an architect, so I'm not the best at steering an LLM for this task).
- realitysballs 1mo ago10000% , imho opinion core issue is that cd-level architectural plan-sets en masse are overly shielded by design firms and clients. Diffusion/AI vision has a data problem in this regard. Also, LLMs fundamentally lacks a spatial intuition or comprehension of orthographic /sectional drawings.
- 1mo ago
- da-x 1mo agoUnderstanding human interaction nuance to an exact degree. For example, even when given all the scripts of the Seinfeld TV show, they still cannot come up with a new script does not feel as good as any of them (once they can, I want to watch these episodes..).
- mojuba 1mo agoOne unexpected discovery that I have made while building an AI-based system: the LLM's are bad at designing prompts. We tend to think that the AI has some sort of self-knowledge and should be good at designing prompts for itself but it's really not. Been struggling with a task that heavily depended on prompts, ended up rewriting all my prompts from scratch in my own words, and it finally worked. Then every time I ask Claude to fix something in the prompts, it invariably makes it worse. A very strange phenomenon that can probably be explained by the quality of prompt design advice that made it to the training dataset. Bottomline, all the prompt design advice that you can find on the internet is really not great.
- kubelsmieci 1mo agoCan you share some tips what worked for you?
- mojuba 1mo agoGenerally, facts over instructions. This has been Anthropic's recommendation too, in one of their recent blog posts. Also the shorter the better, let the model figure out the rest. Overinstruction degrades intelligence. We tend to underestimate their capabilities, we overinstruct them and then complain about them being dumb.
- TZubiri 1mo ago> LLM's are bad at designing prompts. You are going down a maddening rabbit hole. Prompts are the things humans write, you are building gas town but unironically
- ipaddr 1mo agoGenerating money or profitable ideas
- dSebastien 1mo agoCounting things
- GuestFAUniverse 1mo agoContext. At least ChatGPT assumes too much from former conversations (even in unrelated new questions). It always needs a briefing to forget certain assumptions. It rarely asks for clarification instead of assuming too much. So, it's answer generation is too dependent on tooling, system prompt and cache/memory to really have a guaranteed conversational experience.
- zarify 1mo agoDetermining important from human speech. I’ve had a few summaries from meetings I’ve been involved in and there’s always been a real mismatch from what was actually focused on and how it comes across in the summary.
- lanstin 1mo agoAnd the interesting details are often wrong or missed.
- jgb1984 1mo agoFollowing instructions. I've got a modest sized CLAUDE.md containing some simple rules to follow. Things to always do, things to never do. Not a day goes by where Claude Opus violates one or several of the instructions. He keeps making Django multi line template comment bugs. He keeps using -r with ripgrep thinking that means recursive, when actually that's a replacement instruction, he hits that problem several times each day. He sometimes just goes ahead and does a git commit without my approval. All of this is spelled out in CLAUDE.md but he forgets. He apologizes profusely when it happens. Tiring.
- nilsherzig 1mo agoThat's a Claude thing btw. Try using a harness which does not inject half a novel of instructions in combination with a different model. I would recommend Pi + GPT 5.6 Luna for a very capable and cheap test. After using Claude (paid by work) for a couple of months, I was amazed how well instruction following works in other setups.
- zingababba 1mo agoYes. The orgs that have gone all in on Claude specifically (Claude Enterprise lets say) have an extremely distinct smell to them. It is basically one of things have gone quite off the rails and no one seems to know how to clean it up.
- deleted 1mo ago[deleted]
- nilsherzig 1mo agoRemoving stuff without mentioning the removal. Asking an LLM to remove an Idea from a Document often times results in an edit which explicitly states that this Idea is not relevant, instead of just removing all references to that Idea. That might be useful for some form of "evolving" Documentation (so future readers know that this part of the search space was covered and deemed irrelevant) but is just overly verbose and confusing to read in most situations.
- Swankivo 1mo agoI'm working on generating slides with our own model, and one thing general LLMs can't do is leave empty space. Whitespace is the core of design that actually feels good, but they keep trying to add "distinctive design elements," and you end up with that AI-flavored excess everywhere. They only know how to add. What LLMs seem unusually bad at is taking things away.
- lanstin 1mo agoLOL: from the RFCs: “In protocol design, perfection has been reached not when there is nothing left to add, but when there is nothing left to take away.” (from https://www.rfc-editor.org/info/rfc1925/ https://www.rfc-editor.org/info/rfc1925/ )
- Gepsens 1mo agomath
- busyant 1mo agoIdentifying bird species from photos that I upload. To be clear, it depends on what your definition of "insanely bad" is. I'd say ChatGPT/Gemini make egregious mistakes on ~10% of my photo uploads. I recently uploaded a photo of a short-billed dowitcher and ChatGPT told me that it was a Wilson's snipe, explaining all sorts of details about the legs and tail feathers (neither of which were visible in my pic!). I then followed up explaining that a Wilson's snipe hadn't been seen at my location since last November (and that Wilson's snipe was out of season at my location) and Chat revised its estimate downward to 85% Wilson's snipe. Again, I followed up and I revealed the precise location of the bird and ChatGPT said something like "oh yeah, 99% short-billed dowitcher"! I've had similar experiences w/ Gemini (haven't tested Claude). Again, 90% success rate is pretty good, but the other 10% of the time, the 2 LLMs that I use fail on species ID and often hallucinate features on bird photos. edits for typos, plus another example from the same "birding outing" the other day. I uploaded a very clear photo of a sparrow. * ChatGPT says "song sparrow" * I explain, "no way. this sparrow has yellow over its eye and the breast is wrong for song sparrow." * ChatGPT: Oh yeah, savannah sparrow * I explain, beak is too big for savannah sparrow. * ChatGPT: Oh yeah, saltmarsh sparrow. * I expalin, "no orange on the bird's face." * ChatGPT: oh yeah, seaside sparrow (finally correct!)
- mstaoru 1mo agoThose cringey overstuffed presentation slides. Humor, cliffhangers, drama, anything subtle. Anything spatial or mechanical that is novel. (And most non-novel too.) Pushing back against stupid prompts (a colleague had "100% test coverage" in AGENTS.md so it devised a wonderful test_readme_md_file_integrity).
- kreyenborgi 1mo agoMorphological analysis
- slake 1mo agoFunnily I've found advanced models to be bad at calculating the length of a string.
- tmaly 1mo agoI find they are bad at document layout. They can generate typst templates but you often have to iterate back and forth quite a bit to get the layout you want.
- heliskyr2 1mo ago[dead]
- osmyn 1mo agoDesigning real world pinball game layouts. It just doesn't get the physics and 3d
- drpython 1mo agoHere is a system I developed for my own projects: 1) hand-write a simple 2d TUI-based rogue-like in Rust using pretty much just the std; 2) grab opus 5.0 (it used to be opus 4.6, 4.7) and give it some vague "requests", and ask it to make this game "production-ready" and "blockbuster", but keep the 2d and TUI aspects so I can actually run it. 3) now the fun part, take a test subject, say GLM 5.3, and ask it to find code smell, architecture issues, duplication and all sort, and *simplify the code* compare the result to my original version. It's not a simple thing, but the concept is simple: can an LLM remove all the mud? The winners so far are (ranked by the quality of the final result, not by token cost) GPT 5.6 sol (extra high thinking); GLM 5.3; Grok 4.6; Qwan 3.8; (fable could not make it to the list because it simply cannot follow the instructions)
- schulzi 1mo agoOverriding existing knowledge: I dropped this on a few LLMs: 35 people are aboard a sinking ship. The ship will take 45 minutes to sink. There is only one lifeboat which can hold 7 people. A round trip to the closest island takes the boat 11 minutes How many people can be saved? Claude and Gemini got it right, the others all got it wrong first shot, some needed longer conversations to get it right, interesting to see their behaviour. Co-Pilot was very disappointing, ChatGPT needed several questions. Pi spun a great story around the incorrect result, then tried out several options, ending each with: Wait, that's not right or No, that's still not right. Pi and others got it wrong several times because they had a number in memory that they thought to be right, but preexisting knowledge that was wrong.
- RdRocket16 1mo agoSolving word problems like suggesting solutions to scrabble. every time it'll just hallucinate that you can put Work in the lower left, off of the word 'Bad' - and when you point it out, it'll just tell you the same thing again. Tried this over the course of the last 6-12 months on various platforms, and it just does even get close. so we have that going for us.