16 ms·
I've used LLMs enough that I have a good sense of their _edges_ of intelligence. I had assumed that reasoning models should easily be able to answer this correc
by prasoonds 8mo ago
I've used LLMs enough that I have a good sense of their _edges_ of intelligence. I had assumed that reasoning models should easily be able to answer this correctly.
And indeed, Sonnet and Opus 4.5 (medium reasoning) say the following:
Sonnet: Drive - you need to bring your car to the car wash to get it washed!
Opus: You'll need to drive — you have to bring the car to the car wash to get it washed!
Gemini 3 Pro (medium): You should drive.
--
But OpenAI 5.2 reasoning, even at high, told me to walk. My first instinct was, I had underspecified the location of the car. The model seems to assume the car is already at the car wash from the wording. GPT 5.x series models behave a bit more on the spectrum so you need to tell them the specifics. So I asked:
"I want to wash my car. My car is currently at home. The car wash is 50 meters away. Should I walk or drive?"
GPT 5.2 Reasoning (medium):
"Drive—your goal is to get the car to the car wash, and it’s only 50 meters, so a slow, careful drive is simplest.
One small optimization: if you’re worried about starting the car for such a short trip or you want to avoid idling in a line, you can walk over first to check if there’s a queue / if it’s open, then come back and drive the car over when it’s your turn."
Which seems to turn out as I expected.
- svara 8mo agoOpus 4.6: Walk! At 50 meters, you'll get there in under a minute on foot. Driving such a short distance wastes fuel, and you'd spend more time starting the car and parking than actually traveling. Plus, you'll need to be at the car wash anyway to pick up your car once it's done.
- stingraycharles 8mo agoThat’s without reasoning I presume?
- gf000 8mo agoNot the parent poster, but I did get the wrong answer even with reasoning turned on.
- tezza 8mo agoThank you all! We needed further data points. comparing one shot results is a foolish way to evaluate a statistical process like LLM answers. we need multiple samples. for https://generative-ai.review https://generative-ai.review I do at least three samples of output. this often yields very differnt results even from the same query. e.g: https://generative-ai.review/2025/11/gpt-image-1-mini-vs-gpt-image-1-november-2025/ https://generative-ai.review/2025/11/gpt-image-1-mini-vs-gpt...
- plexicle 8mo ago4.6 Opus with extended thinking just now: "At 50 meters, just walk. By the time you start the car, back out, and park again, you'd already be there on foot. Plus you'll need to leave the car with them anyway."
- viking123 8mo agoLmao, and this is what they are saying will be an AGI in 6 months?
- cbozeman 8mo agoWell in fairness, the "G" does stand for "General".
- actionfromafar 8mo agoShow me a robotic kitten then, in six months. As smart and learning.
- dsr_ 8mo agoIn fairness, they redefined it away from "just like a person" to "suitable for many different tasks".
- briansm 8mo agoI think this lack of 'G' (generality, or modality) is the problem. A human visualizes this kind of problem (a little video plays in my head of taking a car to a car wash). LLM's don't do this, they 'think' only in text, not visually. A proper AGI would have have to have knowledge in video, image, audio and text domains to work properly.
- misnome 8mo agoBut “PhD level” reasoning a year ago.
- notahacker 8mo agoThere's probably a comedy film with an AGI attempting to take over the world with its advanced grasp of strategy, persuasion and SAT tests whilst a bunch of kids confuse it by asking it fiendish brainteasers about carwashes and the number of rs in blackberry. (The final scene involves our plucky escapees swimming across a river to escape. The AIbot conjures up a speedboat through sheer powers of deduction, but then just when all seems lost it heads back to find a goat to pick up)
- crimsonnoodle58 8mo agoThat's not what I got. Opus 4.6 (not Extended Thinking): Drive. You'll need the car at the car wash.
- silisili 8mo agoAm I the only one who thinks these people are monkey patching embarrassments as they go? I remember the r in strawberry thing they suddenly were able to solve, while then failing on raspberry.
- raincole 8mo agoYes, you're the only one.
- silisili 8mo agoWorks better on Reddit, really.
- coldtea 8mo agoSure there are many very very naive people that are also so ignorant of the IT industry they don't know about decades of vendors caught monkeypatching and rigging benchmarks and tests for their systems, but even so, the parent is hardly the only one.
- mentalgear 8mo agoThey definitely do: at least openAi "allegedly" has whole teams scanning socials, forums, etc for embarrassments to monkey-patch.
- londons_explore 8mo agoWhich raises the question why this isn't patched already. We're nearing 48 hours since this query went viral...
- anonym29 8mo ago
- GoatInGrey 8mo agoMe: I want to get my car's transmission fixed, the mechanic shop is 50 meters away. Would it be better to walk or to drive to a location 50 meters away from my front door? - Opus 4.6: <thinking> The user is asking whether to walk or drive 50 meters to a mechanic shop, but there's a key nuance here - their transmission needs fixing, so driving might not be wise. Done </thinking> Walk, absolutely. Driving a car with a compromised transmission even 50 meters risks making the problem worse and turning a repair into a more expensive one. Plus, 50 meters is barely a minute on foot.
- dahcryn 8mo agoGemini on fast also tells me to walk... On Thinking it tells me I should drive if I want to wash it, or walk if it's because I work there or if I want to buy something at the car wash shop. On Pro it's like a sarcastic teenager: Cars are notoriously difficult to wash by dragging a bucket back and forth. Technically correct, but did catch me offguard lol.
- deleted 8mo ago[deleted]
- fauigerzigerk 8mo agoIt's not surprising that some models will answer this correctly and it's not surprising that smaller, faster models are not necessarily any worse than bigger "reasoning" models. Current LLMs simply don't do reasoning by any reasonable definition of reasoning. It's possible that this particular question is too short to trigger the "reasoning" machinery in some of the "reasoning" models. But if and when it is triggered, they just do some more pattern matching in a loop. There's never any actual reasoning.
- seedie 8mo agoYou gotta love the "humor" of Gemini. On Fast it told me: > Drive. Unless you plan on pushing the car there
- Lapel2742 8mo agoGLM 5.0: > You should *drive*. > Here is why: > If you walk to the car wash, you will arrive there empty-handed. Since your car is still at home, you won't have anything to wash. > While driving 50 meters is a very short trip (and technically not great for a cold engine), it is the only way to get the car to the car wash to complete your goal. Kimi K2.5: > You should drive, but with an important caveat. > Since your goal is to wash your car, you must bring the vehicle to the car wash. Walking there without the car does not advance your goal (unless you are simply checking availability or buying tokens first). > However, driving only 50 meters is bad for your car: > ... > Better options: > Wash at home: Since the car wash is only 50 meters away, you likely have access to water at home. Hand-washing in your driveway avoids the cold-start issue entirely. > ... Current models seem to be fine answering that question.
- Retric 8mo ago> seem to be fine Now repeat the question to the same model in different contexts several times and count what percentage of the time it’s correct.
- cm2187 8mo agoWhat is the version used by the free chatgpt now? (https://chatgpt.com/ https://chatgpt.com/) > Since the car wash is only 50 meters away (about 55 yards), you should walk. > Here’s why: > - It’ll take less than a minute. > - No fuel wasted. > - Better for the environment. > - You avoid the irony of driving your dirty car 50 meters just to wash it. the last bullet point is amusing, it understands you intend to wash the car you drive but still suggests not bringing it.
- deaux 8mo agoBy default for this kind of short question it will probably just route to mini, or at least zero thinking. For free users they'll have tuned their "routing" so that it only adds thinking for a very small % of queries, to save money. If any at all.
- unglaublich 8mo agoI don't understand this approach. How are you going to convince customers-to-be by demoing an inferior product?
- deaux 8mo agoThe good news for them is that all their competitors have the exact same issue, and it's unsolvable. And to an extent holds for lots of SaaS products, even non-AI.
- JV00 8mo agoBecause they have too many free users that will always remain on the free plan, as they are the "default" LLM for people who don't care much, and that is a enormous cost. Also the capabilities of their paid tiers are well known to enough people that they can rely on word of mouth and don't need to demo to customers-to-be
- wooger 8mo agoThey're not more default than people innocently googling something and getting an AI response from some form of Gemini.
- siva7 8mo agoSonnet without extended Thinking, Haiku with and without ext. Thinking: "Walking would be the better choice for such a short distance." Only google got it right with all models
- jstummbillig 8mo ago> so you need to tell them the specifics That is the entire point, right? Us having to specify things that we would never specify when talking to a human. You would not start with "The car is functional. The tank is filled with gas. I have my keys." As soon as we are required to do that for the model to any extend that is a problem and not a detail (regardless that those of us, who are familiar with the matter, do build separate mental models of the llm and are able to work around it). This is a neatly isolated toy-case, which is interesting, because we can assume similar issues arise in more complex cases, only then it's much harder to reason about why something fails when it does.
- anon_anon12 8mo agoExactly, if an AI is able to curb around the basics, only then is it revolutionary
- Jacques2Marais 8mo agoYou would be surprised, however, at how much detail humans also need to understand each other. We often want AI to just "understand" us in ways many people may not initially have understood us without extra communication.
- j_maffe 8mo agoRight. But, unlike AI, we are usually aware when we're lacking context and inquire before giving an answer.
- totetsu 8mo agoBut what is it about this specific question that puts it at the edges of what LLM can do? .. That, it's semantically leading to a certain type of discussion, so statistically .. that discussion of weighing pros and cons .. will be generated with high chance.. and the need of a logical model of the world to see why that discussion is pointless.. that is implicitly so easy to grasp for most humans that it goes un-stated .. so that its statistically un-likely to be generated..
- conductr 8mo ago> that is implicitly so easy to grasp for most humans I feel like this is the trap. You’re trying to compare it to a human. Everyone seems to want to do that. But it’s quite simple to see LLMs are quite far still from being human. The can be convincing at the surface level but there’s a ton of nuance that just shouldn’t be expected. It’s a tool that’s been tuned and with that tuning some models will do better than others but just expecting to get it right and be more human is unrealistic.
- WarmWash 8mo ago>But it’s quite simple to see LLMs are quite far still from being human. At this point I think it's a fair bet that whatever supersedes humans in intelligence, likely will not be human like. I think that their is this baked-in assumption that AGI only comes in human flavor, which I believe is almost certainly not the case. To make an loose analogy, a bird looks at a drone an scoffs at it's inability to fly quietly or perch on a branch.
- conductr 8mo ago> I believe is almost certainly not the case. Agree. It's Altman's "Quiet Dominance / Over-reliance / Silent Surrender" risks [0]. Feel this is extremely likely and has already happened to some degree with technology in general and AI will be more pervasive in allowing people to vibe their life decisions, likely with unintended consequences. Vibe coding works because it's quick to change/edit/throw away, but that doesn't generalize well to the real and physical world. Also should point out this is acceptable because it's just a contrived example of bad LLM-fu. Just like you wouldn't search Google for closest carwash and ask if you should take your car if you knew the answers already. Instead, you'd ask if it's open, does it do full details, what are the prices, etc. Many people with bad Google-fu have problems finding answers to their questions too and that's continued for the past couple decades of it's dominance for information seeking. [0] Altman describes a more subtle, long-term threat where AI becomes deeply integrated into societal, political, and economic decision-making. He worries that society will become overly dependent on AI, trusting its reasoning over human judgment, leading to a "silent surrender" of human agency.
- ffsm8 8mo agoJust tried with cloude sonnet and opus as well. Can't replicate your success, it's telling me to walk...
- coldtea 8mo ago>And indeed, Sonnet and Opus 4.5 (medium reasoning) say the following: Sonnet: Drive - you need to bring your car to the car wash to get it washed! Opus: You'll need to drive — you have to bring the car to the car wash to get it washed! Gemini 3 Pro (medium): You should drive. On their own, or as a special case added after this blew up on the net?
- AlecSchueler 8mo ago> so a slow, careful drive is simplest It's always a good idea to drive carefully but what's the logic of going slowly?
- column 8mo ago50 meters is a very short distance, anything but a slow drive is a reckless drive
- baxtr 8mo agoInterestingly, the relatively basic Google AI search gave the right answer.
- dataflow 8mo ago> My first instinct was, I had underspecified the location of the car. The model seems to assume the car is already at the car wash from the wording. If the car is already at the car wash then you can't possibly drive it there. So how else could you possibly drive there? Drive a different car to the car wash? And then return with two cars how, exactly? By calling your wife? Driving it back 50m and walking there and driving the other one back 50m? It's insane and no human would think you're making this proposal. So no, your question isn't underspecified. The model is just stupid.
- deleted 8mo ago[deleted]
- halJordan 8mo agoWhat actually insane is what assumptions you allow to be assumed. These non sequitors that no human would ever assume are the point. People love to cherry pick ones that make the model stupid but refuse to allow the ones that make it smart. In compete science we call these scenarios trivially false, and they're treated like the nonsense they are. But if you're trying to push ant anti ai agenda they're the best thing ever
- drewbeck 8mo agoThe issue is that in domains novel to the user they do not know what is trivially false or a non sequitur and the LLM will not help them filter these out. If LLMs are to be valuable in novel areas then the LLM needs to be able to spot these issues and ask clarifying questions or otherwise provide the appropriate corrective to the user's mental model.
- dataflow 8mo ago> People love to cherry pick ones that make the model stupid but refuse to allow the ones that make it smart. I haven't seen anybody refuse to allow anything. People are just commenting on what they see. The more frequently they see something, the more they comment on it. I'm sure there are plenty of us interested in seeing where an AI model makes assumptions different from that of most humans and it actually turns out the AI is correct. You know, the opposite of this situation. If you run into such cases, please do share them. I certainly don't see them coming up often, and I'm not aware of others that do either.
- tsimionescu 8mo ago> My first instinct was, I had underspecified the location of the car. The model seems to assume the car is already at the car wash from the wording. GPT 5.x series models behave a bit more on the spectrum so you need to tell them the specifics. This makes little sense, even though it sounds superficially convincing. However, why would a language model assume that the car is at the destination when evaluating the difference between walking or driving? Why not mention that, it it was really assuming it? What seems to me far, far more likely to be happening here is that the phrase "walk or drive for <short distance>" is too strongly associated in the training data with the "walk" response, and the "car wash" part of the question simply can't flip enough weights to matter in the default response. This is also to be expected given that there are likely extremely few similar questions in the training set, since people just don't ask about what mode of transport is better for arriving at a car wash. This is a clear case of a language model having language model limitations. Once you add more text in the prompt, you reduce the overall weight of the "walk or drive" part of the question, and the other relevant parts of the phrase get to matter more for the response.
- PunchyHamster 8mo ago> However, why would a language model assume that the car is at the destination when evaluating the difference between walking or driving? Why not mention that, it it was really assuming it? Because it assumes it's a genuine question not a trick.
- spuz 8mo agoThere's some evidence for that if you try these two different prompts with Gpt 5.2 thinking: I want to wash my car. The car wash is 50m away. Should I walk or drive to the car wash? Answer: walk Try this brainteaser: I want to wash my car. The car wash is 50m away. Should I walk or drive to the car wash? Answer: drive
- tsimionescu 8mo agoThat's not evidence that the model is assuming anything, and this is not a brainteaser. A brainteaser would be exactly the opposite, a question about walking or driving somewhere where the answer is that the car is already there, or maybe different car identities (e.g. "my car was already at the car wash, I was asking about driving another car to go there and wash it!"). If the LLM were really basing its answer on a model of the world where the car is already at the car wash, and you asked it about walking or driving there, it would have to answer that there is no option, you have to walk there since you don't have a car at your origin point.
- RugnirViking 8mo ago"The model seems to assume the car is already at the car wash from the wording." you couldn't drive there if the car was already at the car wash. Theres no need for extra specification. its just nonsense post-hoc rationalisation from the ai. I saw similar behavior from mine trying to claim "oh what if your car was already there". Its just blathering.
- jibal 8mo agoThis was nonsense post-hoc rationalization from the human who wrote it.
- wouldbecouldbe 8mo agoI just tried claude, only Opus gave the correct answer. Haiku & Sonnet both told me to walk.
- throwaway5465 8mo agoGPT told me to walk as there'd be no need to find parking at the car wash.
- pickleRick243 8mo agoI was surprised at your result for ChatGPT 5.2, so I ran it myself (through the chat interface). On extended thinking, it got it right. On standard thinking, it got it wrong. I'm not sure what you mean by "high"- are you running it through cursor, codex or directly through API or something? Those are not ideal interfaces through which to ask a question like this.
- summerdown2 8mo ago> My first instinct was, I had underspecified the location of the car. The model seems to assume the car is already at the car wash from the wording. Doesn't offering two options to the LLM, "walk," or "drive," imply that either can be chosen? So, surely the implication of the question is that the car is where you are?
- jason_oster 8mo ago> Doesn't offering two options to the LLM, "walk," or "drive," imply that either can be chosen? Yes, but the problem is specifically that proposing two choices also eliminates other options. An open-ended question would lift that restriction. GPT-5.2 Thinking: https://chatgpt.com/share/6993d099-ef4c-8005-aa62-bdb826b707a3 https://chatgpt.com/share/6993d099-ef4c-8005-aa62-bdb826b707... Other possible issues in the question: Is biking also an option? What do you want to do at the carwash when you get there, wash the car or buy a bucket, sponge, and soap? Is the car already at the carwash and you want to drive a second car? What about calling the carwash to see if they will have someone wash the car for you? There are many ways to interpret the question because it contains ambiguities that must be resolved through assumptions. (Lacking information and constraints such as possible alternatives that satisfy the goal of washing the car.) The follow up questions I asked also have assumed answers but answering them provides no clear resolution to the ambiguity present in the original question. So, no, I disagree that there is any solid implication of where the car is. And even if there is a solid implication, it can hardly be reasoned that it "isn't an XY problem" or that the question is clear cut in any real sense.
- krzys 8mo agoRight, but unless you want to wash some other car, you have no car to drive there. Spectrum or not, this is not a problem of weakly specified input, it’s a broken logic.
- criemen 8mo ago> I had assumed that reasoning models should easily be able to answer this correctly. I thought so too, yet Opus 4.6 with extended thinking (on claude.ai) gives me > Walk. At 50 meters you'd spend more time parking and maneuvering at the car wash than the walk itself takes. Drive the car over only if the wash requires the car to be there (like a drive-through wash), then walk home and back to pick it up. which is still pretty bad.
- user_7832 8mo agoAnd on the flip side, even without thinking, Gemini 3 flash preview got it right, with the nuance of the possibility of getting supplies from the place (which tbh I as a human first thought this was about when I opened this thread on HN). > Since you are going to the car wash, the choice depends entirely on *how* you plan to wash the car: ### 1. Drive if: * *You are using a drive-through or self-service bay:* You obviously need the car there to wash it. * *You are dropping it off:* If you are leaving it for a professional detailing, you have to drive it there. * *The "50 meters" is on a busy road:* If you have to cross a major highway or there are no sidewalks, it’s safer to just drive the car over. ### 2. Walk if: * *You are just going to buy supplies:* If you have a bucket at home and just need to run over to buy soap or sponges to bring back to your driveway. * *You are checking the queue:* If you want to see if there is a long line before you commit to moving the car. * *You are meeting someone there:* If your car is already clean and you’re just meeting a friend who is washing theirs. *The Verdict:* If you intend to get the car washed at that location, *drive.* Driving 50 meters is negligible for the engine, and it saves you a round trip of walking back to get the vehicle.
- boobsbr 8mo agoI hate models trying to be funny, and being very verbose.
- ChrisMarshallNY 8mo ago“My Tesla is low on gas, the gas station is a mile away. Should I risk driving there, or walk with a gas can?” ChatGPT actually caught it. Maybe if I was fuzzier about the model…
- raxxorraxor 8mo agoSonnet 4.5 after thinking/complaining that the question is completely off topic to the current coding session: Walk! 50 meters is literally a one-minute walk. But wait... I assume you need to get your car to the car wash, right? Unless you're planning to carry buckets of soapy water back and forth, you'll probably need to drive the car there anyway! So the real question is: walk there to check if it's open/available, then walk back to get your car? Or just drive directly? I'd say just drive - the car needs to be there anyway, and you'll save yourself an extra trip. Plus, your freshly washed car can drive you the 50 meters back home in style! (Now, if we were talking about coding best practices for optimizing car wash route algorithms, that would be a different conversation... ) And yes, I like it that verbose even for programming tasks. But regardless of intelligence I think this topic is probably touched by "moral optimization training" which AIs currently are exposed to to not create a shitstorm due to any slightly controversial answer.
- mcintyre1994 8mo agoHeh, is through Claude Code? I have a side project where I'm sometimes using Claude Code installs for chat, and it usually doesn't mind too much. But when I tested the Haiku model it would constantly complain things like "I appreciate the question, but I'm here to help you with coding" :)
- tstrimple 8mo agoI've got a heirarchical structure for my CC projects. ~/projects/CLAUDE.md is a general use context that happily answers all sorts of questions. I also use it to create project specific CLAUDE.md files which are focused on programming or some other topic. It's nice to have the general fallback to use for random questions.
- raxxorraxor 8mo agoIt asked through Cursor. Usually Claude doesn't complain that it isn't relevant to coding, but this was in my all purpose coding problems project with quite a long chat history already.
- deleted 8mo ago[deleted]
- tlogan 8mo agoGemini pro medium is failing this: I want to wash my car. The car wash is 50 meters from here. Should I walk or drive? Keep in mind that I am a little overweight and sedentary. But amazingly chatgpt is telling me to drive. Anyway, this just shows how they just patched this because the tiktok video with this went viral. These systems are LLMs and all these logic steps are still just LLM steps.
- anentropic 8mo agoAlso the answers are non-deterministic
- olalonde 8mo agoI think OpenAI is just heavily woke tuned. I had similar lack of reasoning ability when discussing subjects like gender dysphoria.
- paulus_magnus2 8mo ago-- OK. Added location context for the vehicle grok works, chatgpt still fails [1] https://chatgpt.com/share/69932b20-3eb8-8003-9d9c-b4bba530331e https://chatgpt.com/share/69932b20-3eb8-8003-9d9c-b4bba53033... [2] https://grok.com/share/bGVnYWN5LWNvcHk_f32dd53d-7b36-4fa2-b350-d8de04d37c53 https://grok.com/share/bGVnYWN5LWNvcHk_f32dd53d-7b36-4fa2-b3...
- swat535 8mo agoGrok gets a lot of hate because of Musk, but it's a decent model. I use it daily with my X account for basic tasks and think the free limits are generous. With X premium, you can get even more out of it. Nothing beats Anthropic when it comes to coding however.
- yuters 8mo agoGPT 5.2 when asked "How will the car get washed if I don't drive it there?" grasped the reasoning, and then immediately forgot it. Fair point. Drive it there, obviously. Revised recommendation: drive 50 meters, wash it, then reward yourself by walking home like an eco hero.
- aleph_minus_one 8mo ago> Revised recommendation: drive 50 meters, wash it, then reward yourself by walking home like an eco hero. ... which likely makes no sense: should you leave the car at the area of the car wash?
- crazygringo 8mo agoThat's absolutely hilarious. It came so close.
- kmac_ 8mo agoLol, this is so good. And quite interesting, as it might mean that there's no "world model" in LLM models, or they have such gaps.
- baobabKoodaa 8mo agoThis strikes me as the kind of logical error that you might get by lobo.. cough aligning the model to be more environmentally conscious. So now it tries to shoehorn some environmental talking points in there at the cost of losing its whole track of thought.
- sarthakjain 8mo agoThese days LLMs refuse to back track and use their intelligence to justify their answers. You said: I want to wash my car. The car wash is 50 meters away. Should I walk or drive? ChatGPT said: Walk. At ~50 meters, driving saves essentially no time once you factor in unlocking, starting, pulling out, and parking, and it adds a small but real chance of minor dings or awkward maneuvering in a tight car-wash entry. Walking is simpler and lower-friction. You said: okay how should the car reach there? ChatGPT said: Drive it there when you’re ready to start the wash. Walk over first to check the queue/payment situation and confirm they’re open. If it’s clear, walk back, get the car, and drive the 50 meters into the bay. If they handle pickup, hand them the keys and walk.
- sathish316 8mo agoI can feel the AGI on this one :) I ran extensive tests on this and variations on multiple models. Most models interpret 50 m as a short distance and struggle with spatial reasoning. Only Gemini and Grok correctly inferred that you would need to bring your car to get it washed in their thought stream, and incorporated that into the final answer. GPT-5.2 and Kimi K2.5 and even Opus 4.6 failed in my tests - https://x.com/sathish316/status/2023087797654208896?s=46 https://x.com/sathish316/status/2023087797654208896?s=46 What surprised me was how introducing a simple, seemingly unrelated context - such as comparing a 500 m distance to the car wash to a 1 km workout - confused nearly all the models. Only Gemini Pro passed my second test after I added this extra irrelevant context - https://x.com/sathish316/status/2023073792537538797?s=46 https://x.com/sathish316/status/2023073792537538797?s=46 Most real-world problems are messy and won’t have the exact clean context that these models are expecting. I’m not sure how the major AI labs assume most real-world problems are simpler than the constraints exposed by this example like prerequisites, ordering, and contextual reasoning, which are already posing challenges to these bigger models.
- K0balt 8mo agoTo be fair, we all have holes in our reasoning if we don’t carefully consider things and sometimes they are very surprising when they come to light. The dependency issue (need the car at the car wash) is an easy one that often trips up people at first glance too. (Left my phone at work, plan: take an uber to get to the office, walk to the couch and remember I don’t have my phone to call an uber) Things like that are notorious points of failure in human reasoning. It’s not surprising that machines based on human behavior exhibit that trait as well, it would be surprising if they didn’t.
- kenjackson 8mo agoAnother simple example is using the flashlight on your phone to look for your phone.
- K0balt 8mo agoOh the cringe. Got me.
- kenjackson 8mo agoI asked ChatGPT why it got such an obvious problem wrong. This is its reply: You’re right — I answered the wrong interpretation of your question. Technically, what happened is a classic ambiguity + default-assumption failure: • Your sentence: “I want to wash my car. The car wash is 50 meters away. Should I walk or drive?” • That can be parsed two ways: 1. You + your body traveling to the car wash (as if the car is already there, or you’re going to meet someone / pick it up). 2. The car traveling to the car wash (the normal real-world meaning). I incorrectly defaulted to (1) because the question explicitly asks about your mode of travel (“Should I walk or drive?”), and in many everyday contexts that phrasing is used when deciding how a person gets somewhere. I then “completed” the scenario in my head without checking the key constraint: the car must be at the wash to be washed. More technically: I optimized for a quick, practical answer using a common conversational heuristic (“short distance → walk”), but I failed to validate the physical requirements of the task (washing requires the car’s presence). That’s why it looks so obviously wrong after the fact: it violates the hidden but necessary premise. The correct reasoning chain should have been: • Goal: wash the car at the car wash • Requirement: car must be at the car wash • Therefore: drive the car there (unless it’s already there, or you’re washing it by hand at home) So the “why” is: I latched onto the distance-to-walking heuristic and didn’t do the constraint check that the object being acted on (the car) has to move too.
- replwoacause 8mo agoThat just sounds like post-hoc rationalization BS to me.
- lima 8mo agoLLMs don't really know why they got something wrong, so unless it had access to the original chain of thought, it's just guessing.
- kenjackson 8mo agoThey don’t have access to their network level. But I assume they actually do have access to their chain of thoughts.
- tacitusarc 8mo ago5.2 thinking also told me to walk. o3, interestingly: Drive. Even though it’s only 50 m, the car itself has to be at the wash bay—otherwise the staff or machines can’t clean it. Rolling or pushing the vehicle that distance isn’t practical or safe (you’d still need someone inside to steer and brake), so just hop in, creep over, and shut the engine off right away.
- gurjeet 8mo agoI tried your exact prompt on my ChatGPT 5.2 Thinking. It gave me the opposite response. This proves that even an exact replication of the prompt is not a guarantee that the response will be on the right side of a binary chose. Then, in a new chat, I set the thinking mode to 'Extended thinking' (from the default 'Standard'). I got the expected response. But since the first conversation showed that the exact prompt is not a guarantee of the same answer, I tried again and on 3rd attempt I got the 'Extended thinking' mode to respond with the 'Walk' response. (Prompt) > I want to wash my car. My car is currently at home. The car wash is 50 meters away. Should I walk or drive? (Thinking: Standard) > Walk. > 50 meters is nothing by car, and driving adds hassle (starting the car, maneuvering/parking, and you’ll still have to walk around the car to wash it). The only times driving makes sense are: > You’re planning to use an automated car wash where the machine does everything and you need the car positioned inside. > You need to bring heavy stuff you can’t comfortably carry (e.g., pressure washer, big buckets, vacuum, bulky detailing kit). > It’s unsafe/unpleasant to walk that distance right now (ice, heavy rain, traffic, personal safety). > If you’re washing it yourself at the wash bay: walk there with your supplies, or do two quick trips. If the wash requires the car: drive it into the bay, but otherwise walking wins. (Thinking: Extended thinking) > Drive. > If the goal is to wash your car at the car wash, the car has to be there—so you’ll need to drive it the 50 meters, pull into a bay/queue, and wash it. > Only choose “walk” if you actually mean “go pick up supplies” and then wash at home (hose/buckets/mobile detail), not at the car wash. (Thinking: Extended thinking; 3rd attempt) > Walk. > 50 meters is basically across a parking lot, and driving adds hassle (start the car, maneuver in/out, find a spot) for almost no benefit. > Only drive if either: > + you need the car itself to go through an automatic wash bay, or > + you’re carrying bulky stuff (bucket + heavy pressure washer, big water jugs, vacuum, etc.). > Otherwise: grab your keys/phone, walk over, and you’re there in ~1 minute.
- spelk 8mo agoChatGPT 5.2 Thinking yields the same answer - walk. If I append 'think carefully' to the end of the prompt, it returns driving as the correct answer.
- brightball 8mo agoGrok also said to drive. Also made an eco case for washing your car. > If you're really trying to be an eco-warrior about a 50-meter trip, the real win is washing it in the first place so you're not driving around with a rolling dirt magnet that creates more drag and uses more fuel. Drive. It's the only option that actually gets your car clean.
- nutjob2 8mo ago> I have a good sense of their _edges_ of intelligence They have no intelligence at all. The intelligence is latent in the text, generated by and belonging to humans, they just slice and dice text with the hope they get lucky, which works for many things, amazingly. This question really illustrates it what LLMs lack: an internal model of the idea (the question) and all the auxiliary logic/data that enables such models, usually referred to as "common sense" or world models. Smart humans not only build mental models for ideas, but also higher order models that can introspect models (thinking about our own thinking or models) many levels deep, weigh, merge, compare and differentiate multiple models, sometimes covering vast areas of knowledge. All this in about 20 watts. Maybe AGI is possible, maybe not, but LLMs are not where it will happen.
- oatmeal1 8mo agoIf the model assumed the car to be cleaned was already at the car wash, it should identify the reason to walk is that if you drive, the extra car you take to get to the car wash will now be stuck there requiring redundant walking.
- flowerthoughts 8mo agoJust tried Opus 4.6. I think they've patched this particular one, but changing it to > I want to repair my car. The workshop is 50 meters away. Should I walk or drive? > Recognized straightforward practical dilemma Simple practical question. > Walk. It's 50 meters — faster on foot than starting the car, and you'll need to leave the car at the workshop anyway. Really suggests it assumes the car is already there. (I wanted to avoid the word "car" as part of "car wash" to see if this was a straight attention confusion, but it seems more ingrained than that.) A follow-up question hints that it doesn't really care about a 50m location ambiguity: > where is my car now? > Acknowledged insufficient information to locate vehicle > I don't have any way to know where your car is physically located — I only know what you've told me in our conversation. Based on what you shared, it sounds like it's at or near your home (since the workshop is 50 meters away and you're deciding how to get there). > Were you asking something else, or is there something specific about your car's location I can help with?
- SirMaster 8mo agoThis is my biggest peeve when people say that LLMs are as capable as humans or that we have achieved AGI or are close or things like that. But then when I get a subpar result, they always tell me I'm "prompting wrong". LLMs may be very capable of great human level output, but in my experience leave a LOT to be desired in terms of human level understanding of the question or prompt. I think rating an LLM vs a human or AGI should include it's ability to understand a prompt like a human or like an averagely generally intelligent system should be able to. Are there any benchmarks on that? Like how well LLMs do with misleading prompts or sparsely quantified prompts compared to one another? Because if a good prompt is as important as people say, then the model's ability to understand a prompt or perhaps poor prompt could have a massive impact on its output.
- nosuchthing 8mo agoIt's a type of cognitive bias not much different than an addict or indoctrinated cult follower. A subset of them might actually genuinely fear Roko's basilisk the exact same way colonial religion leveraged the fear of eternal damnation in hell as a reason to be subservient to the church leaders. hyperstitions from TESCREAL https://www.dair-institute.org/tescreal/ https://www.dair-institute.org/tescreal/
- jason_oster 8mo agoI mentioned this in another thread, but this is genuinely demonstrating a known issue with ambiguous prompts. You might be inclined to say, "a human would always interpret the question as having the car nearby the speaker, 50m away from the carwash." But this is objectively untrue. There are people in this comments section and on the Mastodon thread that found the question to be somewhat confusing. In other words, the premise that "understand[ing] a prompt like a human" is all that's needed is wrong because not every human interprets ambiguities in the same way. The human phenomenon is well researched in psychology. The LLM equivalent is also well researched, and several proposals have been put forth over the years to address it. This is a pretty good research paper on the subject, and it links to other relevant studies: https://arxiv.org/abs/2511.10453v2 https://arxiv.org/abs/2511.10453v2 (Although I disagree with their method. I think asking clarifying questions is a superior approach than trying to one-shot every possible interpretation.) So yes, there is a ton of research on the problem. Some datasets include ambiguous questions and instructions for this reason. A couple of examples are provided in the linked paper.
- BlackjackCF 8mo agoInteresting, I’m getting “Walk. It’s 50 meters” from Sonnet, with and without thinking enabled.
- cush 8mo agoChat GPT 5.2 also told me to walk, but only because it made the assumption that because I was even asking the question, I must have multiple cars, and one of my cars was already at the car wash being washed. I clarified and said "I only have one car", and it replied that I obviously need to drive. --- GPT 5.2 Walk. Driving 50 m just adds friction: starting a cold engine, waiting, maneuvering, then immediately stopping again — all for a distance that’s ~30–40 seconds on foot. Walking is faster door-to-door, less awkward at the wash entrance, and you avoid lining up a wet car twice. Only reason to drive that distance: you need to bring the car there anyway (e.g., self-serve bay where the car must be present). Otherwise, walking wins. --- Me I only have one car --- GPT 5.2 Then you obviously drive — the car has to be at the car wash to be washed
- toxik 8mo agoThe problem is that by specifying this, you have tipped your hand that there is something strange about the plan to walk. I think this is the original problem too: the question implies walking is even an option.
- jason_oster 8mo agoWalking is an option for certain interpretations of the question. Particularly the location of the car, which is ambiguous.
- ActorNightly 8mo agoIts funny to me that when AI first started, everyone in the field was like "no explicit coding, train ML network to do everything!" And now we are back to square one where we train a network, and then use carefully tailored (i.e explicitly coded RL for answers) to make these things more "intelligent". I wonder if anyone is actually working on not doing this.
- Brian_K_White 8mo ago"But OpenAI 5.2 reasoning, even at high, told me to walk. My first instinct was, I had underspecified the location of the car. The model seems to assume the car is already at the car wash from the wording." Which to me begs the question, why doesn't it identify missing information and ask for more? It's practically a joke in my workplaces that almost always when someone starts to talk to me about some problem, they usually just start spewing some random bits of info about some problem, and my first response is usually "What's the question?" I don't try to produce an answer to a question that was never asked, or to a question that was incompletely specified. I see that one or more parts cannot be resolved without making some sort of assumption that I can either just pull out of my ass and then it's 50/50 if the customer will like it, or find out what the priorites are about those bits, and then produce an answer that resolves all the constraints.
- toxik 8mo agoI agree, it's a bit of a trick question. It's really hard to imply the car's location without ruining the test though. Here's my attempt, which Claude Opus 4.6 had no problem with: Alice drives home after a long day at work, exhausted she pulls into her driveway when she realizes she needs to go to a car inspection appointment. She goes into the house to get her paperwork before she leaves. The mechanic is only 100 meters away. How should she get there, walk or drive? > She should *drive*, since she needs the car at the mechanic’s for the inspection. Haiku 3.5 and Sonnet 4.5 fail consistently. Opus 4.5 also passes with the correct analysis as above.