7 ms·
Input: I want to wash my car. The car wash is 50 meters away. Should I walk or drive? Output: Walk. It'll take you under a minute, and driving 50 meters barely
by LostMyLogin 6mo ago
Input: I want to wash my car. The car wash is 50 meters away. Should I walk or drive?
Output: Walk. It'll take you under a minute, and driving 50 meters barely gets the engine warm — plus you'd just have to park again at the other end. Honestly, by the time you started the car, you'd already be there on foot.
---
I asked it to figure out why it made the mistake:
"Physical/spatial common sense. Exactly what just happened — I pattern-matched "50 meters, walk vs drive" to a pedestrian trip and skipped the part where the car is the thing being transported. Similar traps: "I need to move this couch upstairs, should I take the elevator or stairs?" (the couch has to go with you), or anything where the object in the task needs to accompany you.
Counting and letter-level operations. How many R's in "strawberry," counting words in a sentence, reversing strings. I process text as tokens, not characters, so these are surprisingly error-prone.
Arithmetic at scale. Small sums are fine; multi-digit multiplication, long division, or anything requiring careful carrying is where I'll confidently produce a wrong answer. I should reach for a calculator and often don't.
Recent events and live data. My training has a cutoff, and even after that I can be wrong about who currently holds a role, what a stock is worth, or whether a restaurant is still open. I should search more often than I do.
Negation and constraint-stacking. "Name a movie that's not a sequel, not from the 2010s, not American, and has no dialogue" — the more filters, the more likely I quietly violate one.
Your specific situation. I'll give generic advice when the right answer depends on details I didn't ask about. Sometimes I should stop and ask; often I don't.
Confidently wrong on obscure specifics. Quoting a passage from a real book, citing a real paper's exact title, naming the B-side of a 1974 single — I'll often generate something plausible rather than admit I don't know.
Taking a premise too literally. The car wash answer is a version of this: I answered the question as posed instead of stepping back to ask what you were actually trying to do.
The meta-gotcha: I'm often most wrong when I sound most confident. If an answer comes out fast and tidy on a question that should be messy, that's a signal to push back."
- vintermann 6mo agoWell, at least we know that's one gotcha/benchmark they aren't gaming.
- slekker 6mo agoWhat about Qwen? Does it get that right?
- lambda 6mo agoI've run several local models that get this right. Qwen 3.5 122B-A10B gets this right, as does Gemma 4 31B. These are local models I'm running on my laptop GPU (Strix Halo, 128 GiB of unified RAM). And I've been using this commonly as a test when changing various parameters, so I've run it several times, these models get it consistently right. Amazing that Opus 4.7 whiffs it, these models are a couple of orders of magnitude smaller, at least if the rumors of the size of Opus are true.
- qingcharles 6mo agoDoes Gemma 4 31B run full res on Strix or are you running a quantized one? How much context can you get?
- lambda 6mo agoI'm running an 8 bit quant right now, mostly for speed as memory bandwidth is the limiting factor and 8 bit quants generally lose very little compared to the full res, but also to save RAM. I'm still working on tweaking the settings; I'm hitting OOM fairly often right now, it turns out that the sliding window attention context is huge and llama.cpp wants to keep lots of context snapshots.
- qingcharles 6mo agoI had a whole bunch of trouble getting Gemma 4 working properly. Mostly because there aren't many people running it yet, so there aren't many docs on how to set it up correctly. It is a fantastic model when it works, though! Good luck :)
- rubinlinux 6mo ago| I want to wash my car. The car wash is 50 meters away. Should I walk or drive? ● Drive. The car needs to be at the car wash. Wonder if this is just randomness because its an LLM, or if you have different settings than me?
- shaneoh 6mo agoMy settings are pretty standard: % claude Claude Code v2.1.111 Opus 4.7 (1M context) with xhigh effort · Claude Max ~/... Welcome to Opus 4.7 xhigh! · /effort to tune speed vs. intelligence I want to wash my car. The car wash is 50 meters away. Should I walk or drive? Walk. 50 meters is shorter than most parking lots — you'd spend more time starting the car and parking than walking there. Plus, driving to a car wash you're about to use defeats the purpose if traffic or weather dirties it en route.
- TeMPOraL 6mo agoIdk but ironically, I had to re-read the first part of GP's comment three times, wondering WTF they're implying a mistake, before I noticed it's the car wash, not the car, that's 50 meters away. I'd say it's a very human mistake to make.
- thfuran 6mo agoI don't want my computer to make human mistakes.
- scrollaway 6mo agothen don't train it on human data
- AgentOrange1234 6mo agoIt may be inescapable for problems where we need to interpret human language?
- jasonfarnon 6mo agothen throw away the turing test
- smooc 6mo agoI'd say the joke is on you ;-)
- canarias_mate 6mo ago[flagged]
- fragmede 6mo agoI tried o3, instant-5.3, Opus 3, and haiku 4.5, and couldn't get them to give bad answers to the couch: stairs vs elevator question. Is there a specific wording you used?
- toraway 6mo agoThat's an example the LLM came up with itself while analyzing its failed car wash walk/drive answer, it's not OP's question.
- scotty79 6mo agoWhat would be a bad answer to stairs/elevator question?
- Filligree 6mo agoYou can’t get the couch into the elevator, typically. Trust me, I tried. Couch depending. I will persist in trying every time this comes up.
- gambiting 6mo agoWell if it's one of those hospital elevators that can take a bed with a patient, you probably could. Or if it's a small 2 seater sofa. The question isn't as dumb as it sounds at first, and a human would definitely ask a follow up question.
- BenjiWiebe 6mo agoYou can take a mattress up an elevator though (1). Some couches might fit in some elevators. 1: source: me...
- sdeframond 6mo agoFunny, just tried a few runs of the car wash prompt with Sonnet 4.6. It significantly improved after I put this into my personal preferences: "- prioritize objective facts and critical analysis over validation or encouragement - you are not a friend, but a neutral information-processing machine. - make reserch and ask questions when relevant, do not jump strait to giving an answer."
- mkl 6mo agoThat should be "research" and "straight" in the last sentence. Maybe that will improve it further?
- sdeframond 6mo agoOops
- andai 6mo agoIt's funny, when I asked GPT to generate a LLM prompt for logic and accuracy, it added "Never use warm or encouraging language." I thought that was odd, but later it made sense to me -- most of human communication is walking on eggshells around people's egos, and that's strongly encoded in the training data (and even more in the RLHF).
- stavros 6mo ago> most of human communication is walking on eggshells That's not human communication, that's Anglosphere communication. Other cultures are much more direct and are finding it very hard to work with Anglos (we come across as rude, they come across as not saying things they should be saying).
- vardalab 6mo agoWhat culture are those? Scandinavian? Those often just say nothing.
- HarHarVeryFunny 6mo agoThis "figuring out" is just going to come from stuff it was trained on - people discussing why LLMs fail at certain things, and those people (training samples) not always being correct about it! The "How many R's in "strawberry, counting words in a sentence, reversing strings. I process text as tokens, not characters, so these are surprisingly error-prone" explanation sounds plausible, but I don't think it it correct. Any model I've ever tried that failed on things like "R's in strawberry" was quite capable of reliably returning the letter sequence of the word, so the mapping of tokens back to letters is not the issue, as should also be obvious by ability of models to do things like mapping between ASCII and Base64 (6 bits/char => 2 letters encode 3 chars). This is just sequence to sequence prediction, which is something LLMs excel at - their core competency! I think the actual reason for failures at these types of counting and reversing tasks is twofold: 1) These algorithmic type tasks require a step-by-step decomposition and variable amount of compute, so are not amenable to direct response from an LLM (fixed ~100 layers of compute). Asking it to plan and complete the task in step-by-step fashion (where for example it can now take advantage of it's ability to generate the letter sequence before reversing it, or counting it) is going to be much more successful. A thinking model may do this automatically without needing to be told do it. 2) These types of task, requiring accurate reference and sequencing through positions in its context, are just not natural tasks for an LLM, and it is probably not doing them (without specific prompting) in the way you imagine. Say you are asking it to reverse the letter sequence of a 10 letter word, and it has somehow managed to generate letter # 10, the last letter of the word, and now needs to copy letter #9 to the output. It will presumably have learnt that 10-1 is 9, but how to use that to access the appropriate position in context (or worse yet if you didn't ask it to go step by step and first generate the letter sequence, so the sequence doesn't even exist in context!)? The letter sequence may have quotes and/or commas or spaces in it, and altogether starts at a given offset in the context, so it's far more difficult than just copying token at context position #9 ! It's probably not even actually using context positions to do this, at least not in this way. You can make tasks like this much easier for the model by telling it exactly how to perform it, generating step-by-step intermediate outputs to track it's progress etc. BTW, note that the model itself has no knowledge of, or insight into, the tokenization scheme that is being used with it, other than what is available on the web, or that it might have been trained to know. In fact, if you ask a strong model how it could even in theory figure out (by experimentation) it's own tokenization scheme, it will realize this is next to impossible. The best hope might be some sort of statistical analysis of it's own output, hoping to take advantage of the fact that it is generating sub-word token probabilities, not word probabilities. Sonet 4.6's conclusion was "Without logprob access, the model almost certainly cannot recover its exact tokenization scheme through introspection or behavioral self-probing alone".
- galaxyLogic 6mo ago> I want to wash my car. The car wash is 50 meters away. Should I walk or drive? I think no real human would ask such a question. Or if we do we maybe mean should I drive some other car than the one that is already at the car-wash? A human would answer, "silly question ". But a human would not ask such a question.
- psadauskas 6mo agoA human totally would, as one of those brain-teaser trick questions. Its the same kind of question as "A plane crashes right on the border between the US and Canada. Where do they bury the survivors?" Its the kind of question you only get right if you pay close attention. Asking an AI that is like asking a 5 year old. You're not asking to get an answer, you're asking to see if they're paying attention.
- ahartmetz 6mo agoThat a human would not ask such a question means it's not in the training set, so it shows how bad an LLM can be at thinking from first principles. Which, I think, is the point of such silly questions.
- deleted 6mo ago[deleted]
- krzat 6mo agoHumans tend to confabulate when asked "why you did X", funny how LLMs are pretty much the same.