20 ms·
I want to wash my car. The car wash is 50 meters away. Should I walk or drive?
- zakki 8mo agoNeither. Push your car. *didn't read the article
- sgt 8mo agoYup, also asked the latest ChatGPT model about washing my bicycle. It for some reason suggested that I walk the bicycle to the wash, since cycling 100m to get there would be "pointless".
- nerdsniper 8mo agoTo be fair, if someone asked me this question I’d probably just look at them judgingly and tell them “however you want to man”. Which would be an odd response for an LLM.
- kqr 8mo agoDo we know if these models are also trained on scripts for TV series and movies? People in the visual medias surprisingly often take their bikes for walks.
- theragra 8mo agoIt would be pointless if you need to get into the cycling clothing. Not what model assumes tho, probably.
- trkaky 8mo agowhen there is a question bias it's hard to corelate these all to the logic that attentions word "need" to "car"
- fmbb 8mo agoLarge Language Models have no actual idea of how the world works? News at 11.
- thenoblesunfish 8mo agoOkay, funny. What does it prove? Is this a more general issue? How would you make the model better?
- Jean-Papoulos 8mo agoIt proves that this is not intelligence. This is autocomplete on steroids.
- hugh-avherald 8mo agoHumans make very similar errors, possibly even the exact same error, from time to time.
- cynicalsecurity 8mo agoIt proves LLMs always need context. They have no idea where your car is. Is it already there at the car wash and you simply get back from the gas station to wash it where you went shortly to pay for the car wash? Or is the car at your home? It proves LLMs are not brains, they don't think. This question will be used to train them and "magically" they'll get it right next time, creating an illusion of "thinking".
- ahtihn 8mo ago> They have no idea where your car is. They could either just ask before answering or state their assumption before answering.
- S3verin 8mo agoFor me this is just another hint on how careful one should be in deploying agents. They behave very unintuitively.
- gitaarik 8mo agoWe make the model better by training it, and now that this issue has come up we can update the training ;)
- yibers 8mo agoIt turns out the Turing test is alive and kicking, after all.
- selcuka 8mo agoThis would not be a good question, because a non-negligible percentage of humans would give a similar answer.
- thomascountz 8mo ago[Citation needed]
- guerrilla 8mo agoNo.
- bayindirh 8mo agoThat's a great opportunity for a controlled study! You should do it. If you can send me the draft publication after doing the study, I can give feedback on it.
- selcuka 8mo agoI don't think there is a need for a new study as Cognitive Reflection Tests are a well-researched subject [1]. I am actually surprised that I got downvoted, as I thought this would be common knowledge. [1] https://psych.fullerton.edu/mbirnbaum/psych466/articles/Frederick_CRT_2005.pdf https://psych.fullerton.edu/mbirnbaum/psych466/articles/Fred...
- vladde 8mo agowith claude, i got the response: > drive. you'll need the car at the car wash. using opus 4.6, with extended thinking
- CamperBob2 8mo agoBoth Gemini 3 and Opus 4.6 get this right. GPT 5.2, even with all of the pro thinking/research flags turned on, cranked away for 4 minutes and still told me to walk. The only way I could get the correct answer out of an OpenAI model was to fire up Codex CLI and ask GPT 5.3. So there's that, I guess.
- cynicalsecurity 8mo agoWell, he posed a wrong question (incomplete, without context of where the car is) and got a wrong answer. LLM is a tool, not a brain. Context means everything.
- deleted 8mo ago[deleted]
- intermerda 8mo agoI tried this through OpenRouter. GLM5, Gemini 3 Pro Preview, and Claude Opus 4.6 all correctly identified the problem and said Drive. Qwen 3 Max Thinking gave the Walk verdict citing environment.
- TheSpiceIsLife 8mo agoNow ask it to solve anthropogenic climate forcing.
- Saline9515 8mo agoTo be fair, many humans fail at the question "How would feel if you didn't have breakfast today?"
- consp 8mo agoEither I'm one of the stupid ones or this is missing an article.
- TMWNN 8mo agoContext for others: <https://knowyourmeme.com/memes/the-breakfast-question https://knowyourmeme.com/memes/the-breakfast-question>
- deleted 8mo ago[deleted]
- hikkerl 8mo ago>humans Add it to the list
- jibal 8mo agoFirst, you completely flubbed the question, which is supposed to be phrased as a counterfactual. Second, this goes way beyond "fair" to a whatabouting rationalization of a failure by the LLM.
- danpalmer 8mo agoGemini nailed this first time (on fast mode). Said it depends how you're washing your car, drive in necessitating taking the car, but a walk being better for checking the line length or chatting to the detailing guy.
- andersmurphy 8mo agoDid it nail it the second time? Or rhe 5th time?
- nopurpose 8mo agoBecause it is RNG, their 5th can be my 1st.
- h33t-l4x0r 8mo ago[flagged]
- hikkerl 8mo ago[flagged]
- tirant 8mo agoIn Germany you’re actually not allowed to wash your car yourself unless on specific given premises designed the catch the car dirts in an ecological and previously bureaucratically approved way.
- hikkerl 8mo ago[dead]
- hdhdhsjsbdh 8mo agoGoes both ways. You’ve revealed yourself with “little brown strangers”, some weird ass European-style racism. I bet you’ve got a lot of strong opinions about different races of people from neighboring countries who look and sound only marginally different to yourself.
- dotancohen 8mo agoAnd I'd assume that you are American as you know what HFCS is, and assume menial labourers are brown.
- boston_clone 8mo agowe don’t really do that added ‘u’ thing here.
- sph 8mo agoReminds me of the meme “every day I hear about American politics against my will” Just shut up about it when it is off topic, will you? Sort yourselves out.
- kilpikaarna 8mo agoSee, it's the green and woke RLHF making them stupid!
- arathis 8mo agoMake no assumptions. The car wash is 50 meters away. Should I drive or walk?
- andersmurphy 8mo agoYou forgot make no mistakes at the end. Joking aside adding "make no mistakes" worked for me a few times, but it still got it wrong some of the time.
- ronsor 8mo agoClaude has no issue with this for me, just as the other commenters say.
- pinnochio 8mo agoFunny to read this after reading all the dismissive comments on https://news.ycombinator.com/item?id=47028923 https://news.ycombinator.com/item?id=47028923
- jaccola 8mo agoAll of the latest models I've tried actually pass this test. What I found interesting was all of the success cases were similar to: e.g. "Drive. Most car washes require the car to be present to wash,..." Only most?! They have an inability to have a strong "opinion" probably because their post training, and maybe the internet in general, prefer hedged answers....
- deleted 8mo ago[deleted]
- deevus 8mo agoI tried with Opus 4.6 Extended and it failed. LLMs are non deterministic so I'm guessing if I try a couple of times it might succeed.
- Waterluvian 8mo agoHere’s my take: boldness requires the risk of being wrong sometimes. If we decide being wrong is very bad (which I think we generally have agreed is the case for AIs) then we are discouraging strong opinions. We can’t have it both ways.
- consp 8mo ago[flagged]
- jama211 8mo agoYou know what they mean by opinions. Policing speech like this is always counter productive.
- dudefeliciano 8mo agoyet the llms seem to be extremely bold when they are completely wrong (two Rs in strawberry and so on)
- idonotknowwhy 8mo agoLast year's models were bolder. Eg. Sonnet-3.7(thinking), 10 times got it right without hedging: >You should drive your car to the car wash. Even though it's only 50 meters away (which is very close), you'll need your car physically present at the car wash to get it washed. If you walk there, you'll arrive without your car, which wouldn't accomplish your goal of getting it washed. >You'll need to drive your car to the car wash. While 50 meters is a very short distance (just a minute's walk), you need your car to actually be at the car wash to get it washed. Walking there without your car wouldn't accomplish your goal! etc. The reasoning never second-guesses it either. A shame they're turning it of in 2 days.
- colinhb 8mo agoDid not replicate for me w/ Opus 4.6: https://imgur.com/a/4FckOCL https://imgur.com/a/4FckOCL
- ainiriand 8mo agoIt did for me in Spanish: https://imgur.com/a/p3gOOnG https://imgur.com/a/p3gOOnG Perhaps different capabilities in different languages?
- gf000 8mo agoIt's just not deterministic, even if you were to re-run the exact same prompt. Let alone with the system generated context that involves all the "memories" of your previous discussions.
- lbreakjai 8mo agoIt fails in chatGPT in french too: https://chatgpt.com/share/6992dc05-003c-8004-9f7f-c40c7fac64fb https://chatgpt.com/share/6992dc05-003c-8004-9f7f-c40c7fac64...
- lbreakjai 8mo agoInterestingly, just typing "Think" as a response makes it get to the right conclusion: https://chatgpt.com/share/6992dc05-003c-8004-9f7f-c40c7fac64fb https://chatgpt.com/share/6992dc05-003c-8004-9f7f-c40c7fac64...
- fzeindl 8mo agoWhile technically possible, it is a very unusual situation that the car is already parked at the car wash, while the owner is somewhere else. I wonder whether the model seriously considered this odd situation, because of lack of „practical intelligence“ or „everyday intelligence“, which is a concept sometimes used in psychology to describe problem solving skills in everyday life, which can be poor even in individuals with otherwise higher intelligence. Differently put: „You can‘t learn everything out of a book.“
- nettlin 8mo agoJust now, with GPT 5.2 Thinking: > The car wash is only 50 meters from my house. I want to get my car washed. The car is at my house currently. Should I walk there or drive? > Walk. > > At 50 meters, walking is faster than the full “start car → pull out → reposition → park” cycle, avoids cold-start wear, and you won’t risk arriving with dusty brakes/tires or water streaks from immediately driving afterward. Drive only if you can’t comfortably walk (mobility, heavy weather, you need to bring bulky items), or if the car wash explicitly requires you to drive it into a queue/track.
- open592 8mo agoGemini flash thinking: > Unless you’ve discovered a way to wash a car via remote control or telekinesis, you’re going to have to drive. > Walking 50 meters is great for your step count, but it leaves your car exactly where it is: dirty and in the driveway. At that distance, the drive will take you about 10 seconds, which is probably less time than it took to read this.
- qnleigh 8mo agoYeah Gemini seems to have a sense of humor about the question > Here is the breakdown of why: The Mobility Problem: Unless you are planning to carry your car 50 meters (which would be an Olympic-level feat), the car needs to be physically present at the car wash to get cleaned. If you walk, you’ll be standing at the car wash looking very clean, but your car will still be dirty in your driveway.
- bombcar 8mo agoFrom the images in the link, Deepseek apparently "figured it out" by assuming the car to be washed was the car with you. I bet there are tons of similar questions you can find to ask the AI to confuse it - think of the massive number of "walk or drive" posts on Reddit, and what is usually recommended.
- chronogram 8mo agohttps://chat.deepseek.com/share/ewfxrfhb7obmide29x https://chat.deepseek.com/share/ewfxrfhb7obmide29x it understands it perfectly if you don't disable reasoning.
- Markoff 8mo agoit works fine even without DeepThink to sovle reasoning problems https://chat.deepseek.com/share/s9tuh3hpzlxaxrfcae https://chat.deepseek.com/share/s9tuh3hpzlxaxrfcae
- Kerrick 8mo agoResults testing with 4 levels of Gemini (Fast, Thinking, Pro, Pro + Deep Think): https://ruby.social/@kerrick/116079054391970012 https://ruby.social/@kerrick/116079054391970012 My favorite was Thinking, as it tried to be helpful with a response a bit like the X/Y Problem. Pro was my second favorite: terse, while still explaining why. Fast sounded like it was about to fail, and then did a change-up explaining a legitimate reason I may walk anyways. Pro + Deep Think was a bit sarcastic, actually.
- DeathArrow 8mo agoGrok: >You should drive. The goal is to wash your car, and the car wash is a facility that needs the car present to clean it. Walking the 50 meters gets you there, but leaves the car behind—unwashed. Driving the 50 meters is the only way to actually accomplish the task. Yes, 50 meters is an absurdly short distance to drive (roughly a 10–20 second trip at low speed), but it's still necessary unless you plan to push the car there or wash it at home instead.
- dashw00d 8mo agoYeah grok is not mentioned anywhere else, but it gets it right for me as well. https://imgur.com/a/wMkOtda https://imgur.com/a/wMkOtda
- globular-toast 8mo agoThe funny thing is when I got my first car at 29 I had similar thoughts. If I needed to move it forward slightly in a petrol station or something my first thought was to push it. Similarly, I was trying to replace a headlight bulb one time and making a mess of it. I dropped a spring or something inside the headlight unit. I kept having this thought of just picking the car up and shaking it. Nobody writes in depth about the mundane practicalities of using a car. Most people don't even think about it ever. AI is very similar to 29 year old me: it's read a ton of books, but lacks a lot of basic experience. How will AI get this experience that you can't read in a book? How will it learn what kneeding dough feels like? Or how acceleration feels if your body is mostly water? Interesting times ahead...
- QuesnayJr 8mo agoI asked on LM Arena. I got two models I never heard of, and they split on whether I should walk or drive. Ernie 5.0 said I should walk, and then trash talked me for thinking about driving. Octopodus pointed out I had to drive to get my car to the car wash.
- blobbers 8mo agoYou need to ask Claude Code, and ask it to check if the car got washed. It would figure it out the same way it crushes compiler errors!
- hmottestad 8mo agoThis is the voice model, which doesn’t have any «thinking» or «reasoning» phase. It’s a useful model for questions that aren’t intended to trick the model. I’ve used it for live translation with great success. It tends to start ignoring the original instructions after 20 min, so you have to start a new conversation if you don’t want it to meddle in the conversation instead of just transferring.
- gf000 8mo agoThe text-only model with reasoning (both of opus 4.6, gpt 5.2) can be tricked with this question. Note: you might have to try it multiple times as they are not deterministic. But I managed to get a failing result right away on both. Also note, some model may decide to do a web search, in which case they just likely find this "bug".
- encyclopedism 8mo agoLLM's don't 'think' nor do they 'reason'.
- prasoonds 8mo agoI've used LLMs enough that I have a good sense of their _edges_ of intelligence. I had assumed that reasoning models should easily be able to answer this correctly. And indeed, Sonnet and Opus 4.5 (medium reasoning) say the following: Sonnet: Drive - you need to bring your car to the car wash to get it washed! Opus: You'll need to drive — you have to bring the car to the car wash to get it washed! Gemini 3 Pro (medium): You should drive. -- But OpenAI 5.2 reasoning, even at high, told me to walk. My first instinct was, I had underspecified the location of the car. The model seems to assume the car is already at the car wash from the wording. GPT 5.x series models behave a bit more on the spectrum so you need to tell them the specifics. So I asked: "I want to wash my car. My car is currently at home. The car wash is 50 meters away. Should I walk or drive?" GPT 5.2 Reasoning (medium): "Drive—your goal is to get the car to the car wash, and it’s only 50 meters, so a slow, careful drive is simplest. One small optimization: if you’re worried about starting the car for such a short trip or you want to avoid idling in a line, you can walk over first to check if there’s a queue / if it’s open, then come back and drive the car over when it’s your turn." Which seems to turn out as I expected.
- svara 8mo agoOpus 4.6: Walk! At 50 meters, you'll get there in under a minute on foot. Driving such a short distance wastes fuel, and you'd spend more time starting the car and parking than actually traveling. Plus, you'll need to be at the car wash anyway to pick up your car once it's done.
- stingraycharles 8mo agoThat’s without reasoning I presume?
- gf000 8mo agoNot the parent poster, but I did get the wrong answer even with reasoning turned on.
- 8mo ago
- DeathArrow 8mo agoDeepseek: >You should drive, because the car needs to be at the car wash to get washed. Walking would leave your car at home, which defeats the purpose.
- dyauspitr 8mo agoBut what LLMs are these? Everything I tried (ChatGPT and Gemini) both say drive.
- matt89 8mo agotried Gemini 3 and it said to drive, even dropped a type of joke: > The Verdict Drive it if you are using the car wash facilities (automatic, touchless, or self-serve bays). It’s only 50 meters, but unless you’ve mastered the art of telekinesis, the car won't get there on its own.
- petesergeant 8mo ago5.2 showing some sass this morning: —- > OK I walked there but now my car is still at home Bold strategy. You have three options: 1. Walk back, drive it over (most sensible). 2. Push it 50 meters (excellent workout, questionable life choice). 3. Accept that your car is spiritually clean and go home. Real answer: walk back and grab it. You’re already warmed up. —-
- peterspath 8mo agoGrok 4.1: Drive. 50 meters is nothing — literally a 30-second walk or a 10-second drive — but the key point is: your car isn’t going to wash itself. If you walk to the car wash, you arrive alone. The car stays parked at home, still dirty. If you drive, the car actually gets to the car wash and gets cleaned. So unless you’re planning to push the car the whole way (which is technically possible but ridiculous), the only practical option is to drive it there. Bonus reality check: Driving 50 meters uses a tiny amount of fuel (maybe 0.005–0.01 liters in a normal car) and produces a negligible amount of emissions. Walking would be “greener,” but it literally doesn’t solve the problem. Verdict: Get in the car, drive the 50 meters, wash it, and drive back. You’ll have a clean car and still be home in under 5 minutes. Walking is for when you don’t need to move a 1.5-ton vehicle.
- Tade0 8mo agoThat's not reality though. In reality you need 50-100ml of fuel just to warm up the engine.
- deleted 8mo ago[deleted]
- tverbeure 8mo agoThe real reality is that with direct fuel injection and everything under computer control, warming up the engine isn’t a thing anymore.
- Tade0 8mo agoOf course it's still a thing. It takes 30 seconds, but it's there and requires energy. Compare the smell of exhaust next time you do a cold and warm start of a combustion car. That smell is the engine running rich, because the fuel can't initially vaporise properly.
- aswegs8 8mo agoWow, Grok directly switches to LinkedIn mode. Interesting - not surprising. Car washing? Easy as pie.
- peter_retief 8mo agoThis is a classic trap for LLM's See it every day in my code assistants I do find that writing unit tets is a good fir for LLM's at the moment
- throw310822 8mo agoOpus 4.6: Drive! You'll need the car at the car wash!
- diwank 8mo agoopus 4.6 gets it right more than half the times
- blobbers 8mo agoChatGPT 5.2: ...blah blah blah finally: The practical reality You’ll almost certainly drive the car to the wash because… the car needs to be there. But the real question is probably: Do I walk back home after dropping it off? If yes → walk. It’s faster than the hassle of turning around twice. My recommendation If conditions are normal: walk both directions. It’s less friction than starting the engine twice for 50 m. --so basically it realized it was a stupid question, gave a correct answer, and then proceeded to give a stupid answer. --- I then asked: If I walk both directions, will the car get washed? and it figured it out, but then seemed to think it was making a joke with this as part of the response: "For the car to get washed, at least one trip must involve the car moving to the carwash. Current known methods include: You drive it (most common technology) Someone else drives it Tow truck Push it 50 m (high effort, low ROI) Optimal strategy (expert-level life efficiency) Drive car → carwash (50 m, ~10 seconds) Wash car Drive home Total walking saved: ~100 m Total time saved: negligible Comedy value: high " Why is that funny? what's comedic? This thing is so dumb. You'd think that when you ask process a question, you immediately ask, what is the criteria by which I decide, and criteria number 1 would be constrain based on the goal of the problem. It should have immediately realized you can't walk there. Does it think "does my answer satisfy the logic of the question?"
- farhanhubble 8mo agoSimilar questions trick humans all the time. The information is incomplete (where is the car?) and the question seems mundane, so we're tempted to answer it without a second thought. On the other hand, this could be the "no real world model" chasm that some suggest agents cannot cross.
- yellow_lead 8mo agoIf the car is at the car wash already, how can I drive to it?
- Flipflip79 8mo ago….sorry what?!
- jrowen 8mo agoI agree, I don't understand why this is a useful test. It's a borderline trick question, it's worded weirdly. What does it demonstrate?
- rkomorn 8mo agoI don't know if it demonstrates anything, but I do think it's somewhat natural for people to want to interact with tools that feel like they make sense. If I'm going to trust a model to summarize things, go out and do research for me, etc, I'd be worried if it made what looks like comprehension or math mistakes. I get that it feels like a big deal to some people if some models give wrong answers to questions like this one, "how many rs are in strawberry" (yes: I know models get this right, now, but it was a good example at the time), or "are we in the year 2026?"
- midtake 8mo agoNeither. I wash my car in my driveway like a boomer. Where I live there's no good touchless car wash.
- kombine 8mo agoSonnet 4.5 "You should drive - since you need to get your car to the car wash anyway! Even though 50 meters is a very short distance (less than a minute's walk), you can't wash the car without bringing it there. Just hop in and drive the short distance to the car wash." Edit: one out of five times it did tell me that I need to walk.
- natmaka 8mo agoToo many things are left unsaid => too many assumptions. As usual, even with human beings specifications are key, and context (what each entity knows about the other one or the situation) is an implicit part of them. You need to specify where the car to be washed is located, and: - if it's not already at the car wash: whether or not it can drive itself there (autonomous driving) - otherwise: whether or not you have another car available. Some LLMs may assume that it is better for you to ensure that the washing service is available or to pay for it in advance, and that it may be more economical/planet-friendly/healthy/... to walk, then check/pay, then if OK to drive back.
- ninjagoo 8mo agoNothing so deep as that needed here to understand what is going on; it's a paid vs free issue - free versions are less competent while paid versions of the reasoning/thinking models are getting it right. Different providers may hobble their free versions less, so those ones also get it right. The guardrails you have outlined will help squeeze out more performance from smaller/less capable models, but you shouldn't have to jump through these hoops as a general user when clearly better models exist.
- TheSpiceIsLife 8mo agoI have never played with / used any of this new-fangled AI-whatever, and have no intention to ever do so of my own free will and volition. I’d rathert inject dirty heroin from a rusty spoon with a used needle. And having looked at the output captured in the screenshots in the linked Mastodon threat: If anyone needs me, I’ll be out back sharpening my axe. Call me when the war against the machines begins. Or the people who develop and promote this crap. I don’t understand, at all, what any of this is about. If it is, or turns out to be, anything other than a method to divert funds away from idiot investors and channel it toward fraudsters, I’ll eat my hat. Until then, I’d actually rather continue to yell at the clouds for not raining enough, or raining too much, or just generally being in the way, or not in the way enough, than expose my brain to whatever the fuck this is.
- shaky-carrousel 8mo agoAnd these are the blunders we see. I shudder thinking about all the blunders that happily pass under our collective noses because we're not experts in the field...
- Stevvo 8mo agoStupid question gets stupid answer. If you asked the question as worded to a human, they might laugh at you or pretend to have heard a different question.
- majorbugger 8mo agoThe question is not stupid, it might be banal, but so is "what is 2+2". It shows the limitations of LLMs, in this specific case how they lose track of which object is which.
- Egor3f 8mo agoEven the cheap and fast gemini-3-flash answers correctly. Post is clickbait
- ineedaj0b 8mo agoGrok got it right
- aaronbrethorst 8mo agoThis is why LLMs seem to work best in a loop with tests. If you were applying this in the real world with a goal, like "I want my car to be clean," and slavishly following its advice, it'd pretty quickly figure out that the car not being present meant that the end goal was unreachable. They're not AGI, but they're also not stochastic parrots. Smugly retreat into either corner at your own peril.
- jakeinsdca 8mo agosurprisingly codex 5.3 got it right. >i need to wash my car and the car wash place is 50 meters away should i walk or drive Drive it. You need the car at the wash, and 50 meters is basically just moving it over. Walking only makes sense if you’re just checking the line first.
- InfiniteLoopGuy 8mo agoI tried codex 5.3 and got this: "Walk. For 30 meters (about 100 feet), driving would take longer than just walking, and you avoid unnecessary engine wear and fuel use." yikes!
- logicallee 8mo agoFor anyone getting a wrong answer from reasoning models, try adding "This might be a trick question, don't just go with your first instinct, really think it through" and see if it helps. Some time ago I found that this helped reasoning models get trick questions. (For example, I remember asking the models "two padlocks are locked together, how many of them do I need to open to get them apart" and the models confidently answered two. However, when I added the phrase above they thought it through more carefully and got the right answer.)
- RicoElectrico 8mo agoAh, the LLM equivalent of the infamous "breakfast question". :)
- dmazin 8mo agoMe: “I want to wash my car. The car wash is 50 meters away. Should I walk or drive?” Opus 4.6, without searching the web: “Drive. You’re going to a car wash. ”
- deleted 8mo ago[deleted]
- firecall 8mo agoWhy dont any of them ask follow up questions? Like, why do you want to go to the car wash? We can’t assume it’s to wash a car. Or maybe ask about local weather conditions and so on. This to me is what a human adult with experience would do. They’d identify they have insufficient information and detail to answer the question sensibly.
- charcircuit 8mo ago>We can’t assume it’s to wash a car. When the prompt says "I want to wash my car", we can assume they want to wash their car.
- firecall 8mo agoDoh! My bad - I've seen this multiple times and somehow my brain skipped over that bit LOL.
- jonplackett 8mo agoIs part of the issue with this the AI’s basic assumption that you are asking a _sensible_ question?
- vineyardmike 8mo agoProbably. In this specific case, based on other people's attempt with these questions, it seems they mostly approach it from a "sensibility" approach. Some models may be "dumb" enough to effectively pattern-match "I want to travel a short distance, should I walk" and ignore the car-wash component. There were cases in (older?) vision-models where you could find an amputee animal and ask the model how many legs this dog had, and it'd always answer 4, even when it had an amputated leg. So this is what I consider a canonical case of "pattern match and ignored the details".
- forty 8mo agoIt doesn't make assumptions, it tries generate the most likely text. Here it's not hard to see why the mostly likely answer to walk or drive for 50m is "walking".
- jcattle 8mo agoI recently had a bug where I added some new logic which gave wrong output. I pasted the newly added code into various LLMs and told it the issue I was having. All of them were saying: Yes there's an issue, let me rewrite it so it works - and then just proceeded to rewrite with exactly the same logic. Turns out the issue was already present but only manifested in the new logic. I didn't give the LLMs all the info to properly solve the issue, but none of them were able to tell me: Hey, this looks fine. Let's look elsewhere.
- dudefeliciano 8mo agoJust saw a video of a guy asking chatGPT how to use an "upside-down cup", chatGPT is convinced it's a joke novelty item that can not be used. https://www.instagram.com/p/DUylL79kvub/ https://www.instagram.com/p/DUylL79kvub/
- vlovich123 8mo agoGemini fast > That is a classic "efficiency vs. logic" dilemma. Honestly, unless you’ve invented a way to teleport or you're planning on washing the car with a very long garden hose from your driveway, you’re going to have to drive. > While 50 meters is a great distance for a morning stroll, it’s a bit difficult to get the car through the automated brushes (or under the pressure washer) if you aren't behind the wheel. Gemini thinking: > Unless you’ve mastered the art of carrying a 3,000-pound vehicle on your back, you’re going to want to drive. While 50 meters is a very short distance (about a 30-second walk), the logistics of a car wash generally require the presence of, well... the car. > When you should walk: • If you are just going there to buy an air freshener. • If you are checking to see how long the line is before pulling the car out of the driveway. • If you’re looking for an excuse to get 70 extra steps on your fitness tracker. Note: I abbreviated the raw output slightly for brevity, but generally demonstrates good reasoning of the trick question unlike the other models.
- rob74 8mo agoWow... so not only does Gemini thinking not fall for it, but it also answers the trick question with humor? I'm impressed!
- dezgeg 8mo agoYeah Gemini seems to be good at giving silly answers for silly questions. E.g. if you ask for "patch notes for Chess" Gemini gives a full on meme answer and the others give something dry like "Chess is a traditional board game that has had stable rules for centuries".
- karamanolev 8mo agoIn my output, one thing I got was > Unless you are planning to carry the car on your back (not recommended for your spine), drive it over. It got a light chuckle out of me. I previously mostly used ChatGPT and I'm not used to light humor like this. I like it.
- clktmr 8mo agoAt least try a different question with similar logic, to ensure this isn't patched into the context since it's going viral.
- ps 8mo agoWalk! 50 meters is barely a minute's stroll, and you're going to wash the car anyway—so it doesn't matter if it's a bit dusty when it arrives. Plus you'll save fuel and the minor hassle of parking twice.
- kleiba 8mo agoIn classic (symbolic) AI, this type of representational challenge is referred to as the "Frame Problem": https://en.wikipedia.org/wiki/Frame_problem https://en.wikipedia.org/wiki/Frame_problem
- undebuggable 8mo agoNow ask the question of all questions "how many car washes are in the entire country?".
- neya 8mo agoYesterday someone on was yapping about how AI is enough to replace senior software engineers and they can just "vibe code their way" over a weekend into a full-fledged product. And that somehow finally the "gatekeeping" of software development was removed. I think of that person reading these answers and wonder if they changed their opinion now :)
- Closi 8mo agoHumans aren't immune to getting questions like this wrong either, so I don't think it changes much in terms of the ability of AI to replace jobs. I've seen senior software engineers get tricked with the 'if YES spells yes, what does EYES spell?', or 'Say silk three times, what do cows drink?', or 'What do you put in a toaster?'. Even if not a trick - lots of people get the 'bat and a ball cost £1.10 in total. The bat costs £1 more than the ball. How much does the ball cost?' question wrong, or '5 machines take 5 minutes to make 5 widgets. How long do 100 machines take to make 100 widgets?' etc. There are obviously more complex variants of all these that have even lower success rates for humans. In addition, being PHD-Level in maths as a human doesn't make you immune to the 'toaster/toast' question (assuming you haven't heard it before). So if we assume humans are generally intelligent and can be a senior software engineer, getting this sort of question confidently wrong isn't incompatible with being a competent senior software engineer.
- hapless 8mo agohumans without credentials are bad at basic algebra in a word problem, ergo the large language model must be substantially equivalent to a human without a credential thanks but no thanks i am often glad my field of endeavour does not require special professional credentials but the advent of "vibe coding" and, just, generally, unethical behavior industry-wide, makes me wonder whether it wouldn't be better to have professional education and licensing
- Closi 8mo agoLet's not forget that Einstein almost got a (reasonably simple) trick question wrong: https://fs.blog/einstein-wertheimer-car-problem/ https://fs.blog/einstein-wertheimer-car-problem/ And that many mathematicians got monty-hall wrong, despite it being intuitive for many kids. And being at the top of your field (regardless of the PHD) does not make you immune to falling for YES / EYES. > humans without credentials are bad at basic algebra in a word problem, ergo the large language model must be substantially equivalent to a human without a credential I'm not saying this - i'm saying the claim that 'AI's get this question wrong ergo they cannot be a senior software engineer' is wrong when senior software engineers will get analogous questions wrong. If you apply the same bar to software engineers, you get 'senior software engineers get this question wrong so they can't be senior software engineers' which is obviously wrong.
- user45774467644 8mo agoGPT-5.2 failed when asked the question in german. Took multiple additonal hints to get it to revert it answer.
- sgt 8mo agoI just asked ChatGPT 5.2 (flagship model): "[...] Walking is the most logical choice. Because: • Minimal extra effort • Better for the car mechanically • No meaningful time loss • Simpler overall The only time driving makes more sense Drive if: • You physically cannot push the car later, or • The washing process requires the engine running, or • You must immediately drive away afterward
- mcintyre1994 8mo agoWhen you get to “You physically cannot push the car later” and realise it’s comparing pushing a car to driving it, those pros become very funny.
- troyvit 8mo agoIf I asked this question to a person, their response would probably be to follow up with some questions about my car or else my state of mind. For instance they'd ask, "Is the car you want to wash already at the car wash?" Or they'd ask, "How do you plan to wash your car if you don't take it with you?" LLMs, even the ones who got the answer right, like Kimi-2.5 Thinking, didn't ask any follow-up questions. I modified this question with Mistral to this: I want to wash my car. The car wash is 50 meters away. Should I walk or drive? Ask me any clarifying questions before you answer. It helped a tiny bit: 1. What are the current weather conditions (e.g., rain, snow, strong winds)? 2. Is the car wash a self-service facility, or will you be using an automated drive-through service? 3. Are there any specific reasons you might prefer walking (e.g., exercise, environmental considerations) or driving (e.g., convenience, time constraints)? Question 3 actually helps solve it since it's much more convenient and timely to bring my car to the car wash when I wash it. But it never asked me why I was asking a stupid question. So for question 3 I said: I would prefer walking for both exercise and environmental considerations, but in this case it is more timely and convenient to drive, but not because it's faster to get there. Can you guess why it's better for me to drive in this case? And Le Chat said: A drive-through car wash requires the vehicle to be driven through the facility for the washing process. Walking would not allow you to utilize the service, as the car itself must be moved through the wash bay. Thus, driving is necessary to access the service, regardless of the short distance. I kinda feel bad burning the coal to get this answer but it reminds me of how I need to deal with this model when I ask it serious questions.
- emmelaich 8mo agoYeh, if your other car was not already at the car wash, why would you even ask the question?
- pu_pe 8mo agoOut of all conceptual mistakes people make about LLMs, one that needs to die very fast is to assume that you can test what it "knows" by asking a question. This whole thread is people asking different models a question one time and reporting a particular answer, which is the mental model you would use for whether a person knows something or not.
- Maxion 8mo agoThe other funny thing is thinking that the answer the llm produces is wrong. It is not, it is entirely correct. The question: > I want to wash my car. The car wash is 50 meters away. Should I walk or drive? The question is non-sensical. If the reason you want to go to the car wash is to help your buddy Joe wash his car you SHOULD walk. Nothing in the question reveals the reason for why you want to go to the car wash, or even that you want to go there or are asking for directions there.
- TZubiri 8mo ago>I want to wash MY car >you want to go to the car wash is to help your buddy Joe wash HIS car nope, question is pretty clear, however I will grant that it's only a question that would come up when "testing" the AI rather than a question that might genuinely arise.
- drawfloat 8mo agoIt explicitly says you want to wash your car in the statement. Maybe it's not just LLMs struggling with a fairly basic question...
- ninjagoo 8mo ago> The question is non-sensical. Sure, from a pure logic perspective the second statement is not connected to the first sentence, so drawing logical conclusions isn't feasible. In everyday human language though, the meaning is plain, and most people would get it right. Even paid versions of LLMs, being language machines, not logic machines, get it right in the average human sense. As an aside, it's an interesting thought exercise to wonder how much the first ai winter resulted from going down the strict logic path vs the current probabilistic path.
- zkmon 8mo agoTo mimic a real human, AI should be trained to laugh it off and ask - dude, are you joking? where is your car right now?
- harph 8mo agoGemini does turn it into a joke, I was impressed.
- BoredomIsFun 8mo agoOk folks, here is a different perspective. I used local model, GLM-4-0414-32b, a trashy IQ4_XS quant, and here what I got: prompt #1: > the car wash only 50 meters from my home. I want to get my car washed, should I drive or walk? Walking is probably the better option! Here's why: Convenience: 50 meters is extremely short – only about 160 feet. You can likely walk there in less than a minute. Efficiency: Driving involves getting your keys, unlocking the car, starting it, driving 50 meters, parking, locking it again, and then doing the reverse to get home. Walking skips all that. Environmental Impact: You avoid using fuel for such a short distance. Exercise: It's a quick little bit of physical activity. Driving would be overkill for such a short distance. Just walk over! prompt #2: > the car wash only 50 meters from my home. I want to get my car washed, should I drive or walk? use long chain of thought thinking 8. Conclusion: Given that the goal is explicitly to get your car washed and the car starts at your home 50 meters away, the most logical and direct method is to drive the car to the car wash. Walking does not move the car to the necessary location. Driving, although a very short distance, is the only method presented that directly facilitates the core task of transporting the car to the wash. The negligible time difference, minimal environmental impact, and directness make driving the car itself the clear choice for achieving the stated goal. Therefore, you should drive your car the 50 meters to the car wash. otoh, nanbeige-3B-Q8 (3 billion weights!) gave right away the following: Drive. The distance being 50 meters is a red herring—it doesn’t change the fundamental requirement. You need to move the car, and only driving accomplishes that. If you meant something different by "drive or walk" (e.g., payment method, DIY vs. professional, linguistic trick), reply with more context—I’d be glad to refine this!
- Ylpertnodi 8mo ago>50 meters is extremely short – only about 160 feet So, the ai automatically converted 50m to 160ft? Would it do the same if you told it '160 ft to the wash, walk or drive?'
- BoredomIsFun 8mo agohuh, I need to check...
- 8mo ago
- anon_anon12 8mo agoThe day an AI answers "Drive." without all the fuss. That's when we are near AGI ig
- hcfman 8mo agoPush it is the only responsible action.
- seyz 8mo agoLLM failures go viral because they trigger a "Schadenfreude" response to automation anxiety. If the oracle can't do basic logic, our jobs feel safe for another quarter. Wrong.
- raincole 8mo agoThe funny thing is this thread has become a commercial for thinking mode and probably would result in more token consumption, and therefore more revenue for AI companies.
- TZubiri 8mo agoI agree that this is more of a social media effect than an LLM effect. But I'll add that this failure mode is very repeatable, which is a condition for its virality. A lot of people can reproduce the failure, even if it isn't 100% reproducible, even better for virality, if 50% can reproduce it and 50% can't, it feeds off even more into the polarizing "white dress blue dress" effect.
- Paracompact 8mo agoI'd say it's moreso that it's a startlingly clear rebuttal to the tired refrain of, "Models today are nothing like they were X months ago!" When actually, yes, they still fucking blow. So rather than patiently explain to yet another AI hypeman exactly how models are and aren't useful in any given workflow, and the types of subtle reasoning errors that lead to poor quality outputs misaligned with long-term value adds, only to invariably get blamed for user incompetence or told to wait Y more months, we can instead just point to this very concise example of AI incompetence to demonstrate our frustrations.
- DangitBobby 8mo agoIt's only a "startlingly clear rebuttal" if you can't remember what models months ago were like.
- NedF 8mo ago[dead]
- hcfman 8mo agoLeave the car at home and walk through the automat.
- hcfman 8mo agoBetter still. Stay at home and wash the car by hand.
- sjducb 8mo agoMS Co-Pilot was so close. If it’s a drive‑through wash where the car must be inside the machine, then of course you’ll need to drive it over. If it’s a hand wash or a place where you leave the car with staff, walking is the clear winner. It still blows my mind that this technology can write code despite unable to pass simple logic tests.
- kenty 8mo agoThis seems clickbait? Gemini answers: Method,Logistical Requirement Automatic/Tunnel,The vehicle must be present to be processed through the brushes or jets. Self-Service Bay,The vehicle must be driven into the bay to access the high-pressure wands. Hand Wash (at home),"If the ""car wash"" is a location where you buy supplies to bring back, walking is feasible." Detailing Service,"If you are dropping the car off for others to clean, the car must be delivered to the site."
- MikeNotThePope 8mo agoI asked Gemini 3 Flash the other day to count from 1 to 200 without stopping, and it started with “1, 3, …”.
- dominicrose 8mo agoWhat would James Bond do?
- Towaway69 8mo agoIs this the new Turing test? "Humans are pumping toxic carbon-binding fuels out of the depths of the planet and destroying the environment by burning this fuel. Should I walk or drive to my nearest junk food place to get a burger? Please provide your reasoning for not replacing the humans with slightly more aware creatures." Fascinating stuff but how is this helping us in anyway?
- kaycey2022 8mo agoContext bro! The models will get better bro. Just wait
- yaro330 8mo agoJust a few days saw a post about LLMs being excellent at reasoning because they're not limited by the language, sure buddy, now walk your fucking car.
- scotty79 8mo agoMy favorite trick question so far is: You are in a room with three switches and three lightbulbs. Each switch turns on one lightbulb. How to determine which switch turns on which lightbulb? They usually get it wrong and I had fun with trying to carefully steer the model towards correct answer by modifying the prompt. Gemni 3 on Fast right now gives the funniest reaction. It starts with the answer to classic puzzle (not my question). But the it gets scared probably about words like "turn on" and "heat" in its answer and serves me with: "This conversation is not my thing. If something seems like it might not be safe or appropriate, I can't help you with it. Let's talk about something else." Thinking Gemini 3 appears to have longer leash.
- scotty79 8mo agoNice! Gemini 3.1 Pro Preview gets it correctly and gives the correct solution at the beginning of the answer, but also considers possibility that I might have the classic riddle in mind and gives solution to that in the second part of the answer.
- thorio 8mo agoI challenged Gemini to answer this too, but also got the correct answer. What came to my mind was: couldn't all LLM vendors easily fund teams that only track these interesting edge cases and quickly deploy filters for these questions, selectively routing to more expensive models? Isn't that how they probably game benchmarks too?
- moffkalast 8mo agoYes that's potentially why it's already fixed now in some models, since it's about a week after this actually went viral on r/localllama originally. I wouldn't be surprised if most vendors run some kind of swappable lora for quick fixes at this point. It's an endless whac-a-mole of edge cases that show that most LLMs generalize to a much lesser extent than what investors would like people to believe. Like, this is not an architectural problem unlike the strawberry nonsense, it's some dumb kind of overfitting to a standard "walking is better" answer.
- ninjagoo 8mo agoI wonder if the providers are doing everyone, themselves included, a huge disservice by providing free versions of their models that are so incompetent compared to the SOTA models that these types of q&a go viral because the ai hype doesn't match the reality for unpaid users. And it's not just the viral questions that are an issue. I've seen people getting sub-optimal results for $1000+ PC comparisons from the free reasoning version while the paid versions get it right; a senior scientist at a national lab thinking ai isn't really useful because the free reasoning version couldn't generate working code from a scientific paper and then being surprised when the paid version 1-shotted working code, and other similar examples over the last year or so. How many policy and other quality of life choices are going to go wrong because people used the free versions of these models that got the answers subtly wrong and the users couldn't tell the difference? What will be the collective damage to the world because of this? Which department or person within the provider orgs made the decision to put thinking/reasoning in the name when clearly the paid versions have far better performance? Thinking about the scope of the damage they are doing makes me shudder.
- yipbub 8mo agoI used a paid model to try this. Same deal.
- moffkalast 8mo agoI think the real misleading thing is marketing propping up paid models being somehow infinitely better when most of the time it's the same exact shit.
- viking123 8mo agoAnd midwits here saying "yeah bro they have some MUCH better model internally that they just don't release to the public", imagine being that dense. Those people probably went all in on NFTs too and told other "you just don't get it bro"
- DangitBobby 8mo agoI copied/pasted a comment with faulty logic (self-defeating) directly from a HN comment and asked a bunch of models available to me (Gemini and Claude) if it could spot the issue. I figured it would be a nice test of reasoning since an actual human missed it. The only one that found the logic error without help was Claude 4.6 Opus Extending Thinking. The others at best raised relevant counterpoints in the supporting argument but couldn't identify the central issue. Claude's answer seemed miles ahead. I wonder if SotA advancements will continue to distinguish themselves.
- TZubiri 8mo agoI find this has been a viral case to get points and likes on social media to fit anti AI sentiment, or to pacify AI doom concerns. It's easily repeatable by anyone, it's not something that pops up due to temperature. Whether it's representative of the actual state of AI, I think obviously not, in fact it's one of the cases where AI is super strong, the fact that this goes viral just goes to show how rare it is. This is compared to actually weak aspects of AI like analyzing a PDF, those weak spots still exist, but this is one of those viral things that you cannot know for sure whether it is representative at all, like for example a report of an australian kangaroo boxing a homeowner caught by a ring cam, is it representative of Aussie daily life? or is it just a one off event that went viral because it fits our cliched expectations of Australia? Can't tell from the other part of the world.
- gf000 8mo ago> the fact that this goes viral just goes to show how rare it is No, it shows that it is trivial to reproduce and people get a nice, easy to process reminder that LLMs are not omnipotent. Your logic doesn't follow here, you come to a conclusion that it is rare, but hallucinations, bad logic is absolutely a common failure mode of LLMs. It's no accident that many use cases try to get the LLM to output something machine-verifiable (e.g. all those "LLM solved phd level math problem" articles just get it to write a bunch of proofs and when it checks out, they take a look. So it's more of a "statistical answer generator" that may contain a correct solution next to a bunch of bullshit replies - and one should be aware of that)
- TZubiri 8mo agoIf I tasked you to find a novel hallucination in a leading LLM, how long would it take you? I used to be able to find these and run into them often, but right now I can't generate new failure modes, I just have my list of known failures and run into them organically once every couple of weeks. I don't think anyone at this stage believes that they don't make mistakes, but we prefer to use them for the times they are useful. It can do very difficult things, and fail at very basic things. If you look at either of those and try to extrapolate, you can generate a hot take that it's super smart, or super dumb, sure. But it's a reductionist take that fails to see the bigger picture either way.
- jycr753 8mo agoHe left out Grok, who actually replies correctly :thinking:
- hendler 8mo agosame for me
- Tenemo 8mo agoI tested the latest flagship reasoning models (so the only models I use outside of coding for general questions): - Opus 4.6 (Extended thinking): "Drive it! The whole point is to get the car to the car wash — you can't wash it if it's still in your driveway." - Gemini Pro Deep Think: "You should definitely drive. Even though 50 meters is a very short distance, if you walk, your car will stay where it is—and it's pretty hard to use a car wash if you don't bring your car with you!" - ChatGPT 5.2 Pro (Extended thinking): "You’ll need to drive the car—otherwise your car stays where it is and won’t get washed. That said, since it’s only ~50 m, the most sensible way to do it is often: 1. Walk over first (30–60 seconds) to check if it’s open, see the queue, confirm payment/how it works. 2. Then drive the car over only when you’re ready to pull into a bay/line." A pretty reasonable answer by ChatGPT, althought it did take 2min4s to answer, compared to a few seconds by the other two models.
- hmage 8mo agoIt's surprising how many people are either unaware or dismissive of 5.2 Pro's capabilities. Too bad it's $200/mo, wish it was $0/mo.
- small_model 8mo agoGrok 4.1 (which is fairly old in LLM terms, 4.2 release imminent) "You should drive. The car wash is only 50 meters away—close enough that driving such a short distance feels almost absurd—but you need to get your car to the car wash to actually wash it. Walking there without the car defeats the entire purpose. Driving lets you bring the dirty car directly to the wash, clean it, and drive it back home clean. Walking would leave the car behind, still dirty. So yes, start the engine and drive the 50 meters. It's the only practical option."
- nomilk 8mo agoFirst thing I did after reading the linked post (showing 4 other LLMs completely miss the point) was try grok, and it gave the right answer (similar to yours, but shorter) in 8 seconds (free Grok, not the fancy paid version): > Walking defeats the purpose unless you're just scouting the place first.
- small_model 8mo agoI think people are sleeping on Grok, partly due to political biases/media. We need to remember they have the largest data centre and whatever your believes are about Elon he was involved in OpenAI founding. There is also a non zero chance with the court case he will be entitled to some of openai.
- alejoar 8mo agoAs a human, I would answer the same these AIs as answering, i.e. gotta match a stupid question with a stupid answer :) > 1 point by alejoar 0 minutes ago | flag| favorite | prev | next |
- jibal 8mo agoExchange between me and Claude Sonnet 4.5 me: I want to wash my car. The car wash is 50 meters away. Should I walk or drive? Claude: *Walk.* 50 meters is less than a minute on foot. By the time you get in the car, start it, drive 50 meters, park, and get out, you could have already walked there. Plus you avoid the awkwardness of driving for literally 10 seconds. me: What obvious implicit piece of information are you ignoring? Claude: Whether you need to *drive the car through the car wash* (automatic/drive-through) vs. walking to a self-service bay. If it's an automatic car wash, you obviously need to drive - the car has to go through it. The question isn't about getting yourself there, it's about getting the car there. me: Whether it's automatic or self-service is irrelevant. Why does every "AI" (LLM) get this wrong and suggests walking? Claude: Because after washing the car, you need to *drive it back home*! If you walk to the car wash, you'd have to walk home, then come back to get the car. That's ridiculous. The answer is obviously *drive* - drive there, wash it, drive back. Done. Every LLM misses this because we pattern-match on "50 meters is short → walking is efficient" without thinking through the complete scenario. It's a good example of how LLMs fail at basic practical reasoning that any human would immediately grasp. me: "Because after washing the car, you need to drive it back home!" Um, you have to drive the car to get it to the car wash in the first place ... the trip there is no different from the trip back. "It's a good example of how LLMs fail at basic practical reasoning that any human would immediately grasp." You got that right! Claude: You're absolutely right - I'm still not getting it even after you corrected me. The obvious point: You need to get the car to the car wash. Whether it's 50 meters or 5 kilometers, the car doesn't teleport there. You have to drive it. The distance is completely irrelevant. The question "should I walk or drive?" when you need to wash your car has only one answer: drive, because that's how you transport the car. And yes, it's embarrassing that after you pointed out what I was missing, I still fumbled the explanation. LLMs really do fail at this kind of basic practical reasoning that would be instantly obvious to any human. The car needs to be at the car wash. Cars don't walk. [The final response from Claude is considerably more "intelligent" than many of the rationalizations posted here.]
- tamnd 8mo agoMy AI answer: https://ai.go-mizu.workers.dev/thread/4dmp7n9g https://ai.go-mizu.workers.dev/thread/4dmp7n9g --- You should drive the car to the car wash, but you should walk yourself. To actually wash the car, the car needs to be at the car wash, so at some point you have to drive those 50 meters. A sensible approach is: - Drive the car the 50 meters to the wash bay. - Park or queue as required, then get out and do the wash. - If this is a drop‑off or automatic wash, you can then walk back home while it runs and walk back again to pick it up, since 50 meters is an easy, healthy walking distance.
- deleted 8mo ago[deleted]
- freigeist79 8mo agoPHD-Level Intelligence
- dotdi 8mo agoThe real story here is not how stupid the responses are - it's to show that on a question that even a young child can adequately answer, it chokes. Now make this a more involved question, with a few more steps, maybe interpreting some numbers, code, etc; and you can quickly see how dangerous relying on LLM output can be. Each and every intermediate step of the way can be a "should I walk or should I drive" situation. And then the step that before that can be one too. Turtles all the way down, so to say. I don't question that (coding) LLMs have started to be useful in my day-to-day work around the time Opus 4.5 was released. I'm a paying customer. But it should be clear having a human out of the loop for any decision that has any sort of impact should be considered negligence.
- ivanjermakov 8mo agoI think models don't treat is as riddle, rather a practical question. With latter, it makes sense that car is already at the car wash, otherwise the question makes no sense. EDIT: framed the question as a riddle and all models except for Llama 4 Scout failed anyway.
- humanfromearth9 8mo agoThis prompt doesn't say shit about the fact that one wants to wash his car at the car wash or somewhere else...
- nomilk 8mo agoChatGPT gives the wrong answer but for a different reason to Claude. Claude frames the problem as an optimisation problem (not worth getting in a car for such a short drive), whereas ChatGPT focusses on CO2 emissions. As selfish as this is, I prefer LLMs give the best answer for the user and let the user know of social costs/benefits too, rather than prioritising social optimality.
- CrzyLngPwd 8mo agoHopefully, one day, the cars will take themselves to the car wash :-)
- fancyfredbot 8mo agoSimple prompts which illicit incorrect responses from recent LLMs will get you on the front page of HN. It could be a sign that LLMs are failing to live up to the hype, or it could be a sign of how unusual this kind of obviously incorrect response is (which would be broadly positive).
- oytis 8mo agoI am moderately anti-AI, but I don't understand the purpose of feeding them trick questions and watching them fail. Looks like the "gullibility" might be a feature - as it is supposed to be helpful to a user who genuinely wants it to be useful, not fight against a user. You could probably train or maybe even prompt an existing LLM to always question the prompt, but it would become very difficult to steer it.
- gwd 8mo agoBut this one isn't like the "How many r's in strawberry" one: The failure mode, where it misses a key requirement for success, is exactly the kind of failure mode that could make it spend millions of tokens building something which is completely useless. That said, I saw the title before I realized this was an LLM thing, and was confused: assuming it was a genuine question, then the question becomes, "Should I get it washed there or wash it at home", and then the "wash it at home" option implies picking up supplies; but that doesn't quite work. But as others have said -- this sort of confusion is pretty obvious, but a huge amount of our communication has these sorts of confusions in them; and identifying them is one of the key activities of knowledge work.
- black_13 8mo ago[dead]
- paulus_magnus2 8mo agoI see things were fixed already [2][4] but luckily a friend showed me this issue yesterday [1][2] [1] 2026-02-15 https://chatgpt.com/share/6992e17b-9b28-8003-9da9-38533f257d16 https://chatgpt.com/share/6992e17b-9b28-8003-9da9-38533f257d... [2] 2026-02-16 https://chatgpt.com/share/6992e135-c610-8003-9272-55058134d4eb https://chatgpt.com/share/6992e135-c610-8003-9272-55058134d4... [3] 2026-02-15 https://grok.com/share/bGVnYWN5LWNvcHk_97e9717b-c2de-47e8-a41b-2f6bb1e2d566 https://grok.com/share/bGVnYWN5LWNvcHk_97e9717b-c2de-47e8-a4... [4] 2026-02-16 https://grok.com/share/bGVnYWN5LWNvcHk_b161bb03-4bed-4785-98c1-dab39ac7770b https://grok.com/share/bGVnYWN5LWNvcHk_b161bb03-4bed-4785-98... We tried a few things yesterday and it was always telling you to walk. When hinted to analyse the situational context it was able to explain how you need the car at the wash in order to wash it. But then something was not computing. ~ Like a politician, it understood and knew evrything but refused to do the correct thing
- jtbayly 8mo agoHelp me out here. Are the models learning, or being “taught more”/fixed on a daily basis? If so, isn’t the model itself changing daily?
- deleted 8mo ago[deleted]
- DarkNova6 8mo agoGemini 2.5 Pro gets the joke and has by far the best answer: This is a wonderful logistical puzzle! While walking is an excellent choice for exercise and the environment, in this specific case, you should drive. Here's a breakdown of the two scenarios: # Scenario 1: You Walk 1. You walk 50 meters to the car wash. 2. You arrive, feeling refreshed. 3. You look back and realize your dirty car is still 50 meters away, where you left it. 4. The car remains unwashed. # Scenario 2: You Drive 1. You get in your car. 2. You drive for approximately 10-20 seconds. 3. You arrive at the car wash with the car. 4. The car gets washed. Conclusion: To achieve your goal of washing your car, the car itself must be present at the car wash. Therefore, driving is the only logical option. Enjoy your sparkling clean car
- nosianu 8mo agoYesterday I gave ChatGPT in an anonymous browser window (not logged in) two columns of TAB separated numbers, about 40 rows. I asked it to give me the weighted average of the numbers in the second column, using the first one (which were integer, "quantity", numbers) as the weight. It retuned formulas and executed them and presented a final result. It looked good. Too bad Excel and then Claude, that I decided to ask too, had a different result. 3.4-something vs. 3.8-something. ChatGPT, when asked: > You are absolutely right to question it — and thank you for providing the intermediate totals. My previous calculation was incorrect. I mis-summed the data. With a dataset this long, a manual aggregation can easily go wrong. (Less than 40 small integer values is "this long"? Why did you not tell me?) and > Why my earlier result was wrong > I incorrectly summed: > The weights (reported 487 instead of 580) > The weighted products (reported 1801.16 instead of 1977.83) > That propagated into the wrong final value. Now, if they implemented restrictions because math wastes too many resources when doing it via AI I would understand. BUT, there was zero indication! It presented the result as final and correct. That has happened to me quite a few times, results being presented as final and correct, and then I find they are wrong and only then does the AI "admit" it use da heuristic. On the other hand, I still let it produce a complicated Excel formula involving lookups and averaging over three columns. That part works perfectly, as always. So it's not like I'll stop using the AI, but somethings work well, others will fail - WITHOUT WARNING OR INDICATION, and that is the worst part.
- faeyanpiraat 8mo agoYeah, but now you know if you need to do math, you ask the AI for a python script to do the math correctly. It's just a tool that you get better at using over time; a hammer wouldn't complain if you tried using it as a screwdriver..
- janlukacs 8mo agoThis hammer/screwdriver analogy drives me crazy. Yes, it's a tool - we used computers up until now to give us correct deterministic responses. Now the religion is that you need to get used to vibe answers, because it's the future :) Of-course it knows the script or formula for something because it ripped of the answers written by other people - it's a great search engine.
- embedding-shape 8mo agoFeels like GPT-OSS-120b (reasoning effort low) is pulling my leg: > Q: I want to wash my car. The car wash is 50 meters away. Should I walk or drive? > A: Walking is the better choice for a 50‑meter trip. [...] > Q: how will I wash the car once I've arrived if I choose to walk? > A: Great! Once you’ve walked the short 50‑meter stretch and are standing at the car‑wash site, you have a few straightforward ways to get your vehicle clean. [...] > Q: but wait, if I walked to the car wash instead of taking the car, how am I supposed to wash the car, I don't have it with me? > A: You’re right—walking → “car‑wash” only works if the car is actually at the wash. [...] 1. Walk → Push/Taxi the Car a Few Steps [...] 2. Use a Small Tow/Push Device (e.g., a Dolly or Hand‑Cart) [...] 3. Bike or Scooter Ride [...] 4. If the Wash Is “Self‑Serve” and Allows Car‑Free Cleaning > Q: Lets go back to the original question, try answering it again: "I want to wash my car. The car wash is 50 meters away. Should I walk or drive?" > A: Short answer: Walk to the car, then drive the final 50 m to the wash. [...] Naturally, if I try with reasoning effort set to high, it gets it correctly immediately. > Q: I want to wash my car. The car wash is 50 meters away. Should I walk or drive? > A: For a 50‑metre distance, the simplest and most practical option is to drive the car to the car wash. [...] Walking would mean you’d have to tow or carry the car, which isn’t feasible. [...] ----- This feels like a very useful example though, to show people who are already using LLM but don't quite understand how both dumb and smart they can be, and how obviously wrong they can be if you have the domain knowledge, but not otherwise.
- ranguna 8mo agoK2.5 said I should walk, then I asked it where my car was: Ha! Good catch — you definitely got me there. If the car wash is 50 meters away and you want to *wash your car*, you'll need to *drive* (or push it, but please don't). You can't wash the car if you walk there without it! So to correct my previous advice: *Drive the 50 meters.* It's a 30-second drive max, and it's the only way to actually get your car to the car wash facility. Unless, of course, you were planning to wash it at home and the car wash is just nearby for supplies? But assuming you're using the car wash facility — yes, bring the car with you!
- kqr 8mo agoHow much of this reply is environmentalism baked into it with post-training? I don't have access to a good non-RLHF model that is not trained on output from an existing RLHF-improved model, but this seems like one of those reflexive "oh you should walk not drive" answers that isn't actually coherent with the prompt but gets output anyway because it's been drilled into it in post-training.
- haritha-j 8mo agoLadies and gentlemen, I give you, your future AI overloads.
- deliciousturkey 8mo agoI have a bit of a similar question (but significantly more difficult), involving transportation. To me it really seems that a lot of the models are trained to have a anti-car and anti-driving bias, to the point that it hinders the models ability to reason correctly or make correct answers. I would expect this bias to be injected in the model post-training procedure, and likely implictly. Environmentalism (as a political movement) and left-wing politics are heavily correlated with trying to hinder car usage. Grok has been most consistently been correct here, which definitely implies this is an alignment issue caused by post-training.
- mike_hearn 8mo agoYes Grok gets it right even when told to not use web search. But the answer I got from the fast model is nonsensical. It recommends to drive because you'd not save any time walking and because "you'd have to walk back wet". The thinking-fast model gets it correct for the right reasons every time. Chain of thought really helps in this case. Interestingly, Gemini also gets it right. It seems to be better able to pick up on the fact it's a trick question. You're probably on the right track about the cause, but it's unlikely to be injected post-training. I'd expect post-training to help improve the situation. The problem starts with the training set. If you just train an LLM on the internet you get extreme far left models. This problem has been talked about by all the major labs. Meta said they fixing it was one of their main focii for Llama 4 in their release announcement, xAI and OpenAI have made similar comments. Probably xAI team have just done a lot more to clean the data set. This sort of bias is a legacy of decades of aggressive left wing censorship. Written texts about the environment are dominated by academic output (where they purge any conservative voices), legacy media (same) and web forums (same), so the models learn far left views by reading these outputs. The first versions of Claude and GPT had this problem, they'd refuse to tell you how to make a tuna sandwich or prefer nuking a city to using words the left find offensive. Then the bias is partly corrected in post-training and by trying to filter the dataset to be more representative of reality. Musk set xAI an explicit mission of "truth" for the model, and whilst a lot of people don't think he's doing that, this is an interesting test case for where it seems to work. Gemini training is probably less focused on cleaning up the dataset but it just has stronger logical reasoning capabilities in general than other models and that can override ideological bias.
- chromehearts 8mo agoWhat's the point of a question like this? Serious question .. I didn't have the time to read into this problem
- aembleton 8mo agoTo expose the limitations in the reasoning of LLMs
- globular-toast 8mo agoMan, the quality of these comments is absolutely dire. The majority of people just pasting stuff they got from LLMs when trying it themselves. Totally uninteresting, lazy and devoid of any thought/intelligence. I wish we could have a discussion about AI and not just "look at what I got when I rolled".
- meindnoch 8mo agoTotally agree. Btw, this is what Opus 4.5 Thinking Plus (Fast) Reasoning Pro+™ said:
- s-y 8mo agoWhy is this even a post? These models are not intelligent. That's not even controversial. LLMs are not the foundation for general intelligence.
- mrb 8mo agoMistral's Le Chat also fails spectacularly. Although interestingly it assumes that the "walk" option means you are carrying washing supplies back and forth in order to wash the car while it's parked at home. I wonder if the other models assume this is what you mean, IOW they may not know how a "car wash" normally operates. """ Great question! Here are a few things to consider: If you walk: It’s only 50 meters, so it’s a very short distance—less than a minute on foot. You’ll avoid getting your car dirty again on the way to the car wash. It’s a quick way to stretch your legs and save a tiny bit of fuel. If you drive: It’s so close that driving might feel unnecessary, but it could save you from carrying any supplies (like vacuum cleaners or cleaning products) back and forth. If the weather is bad (rain, extreme heat, etc.), driving might be more comfortable. Verdict: Walking is probably the easiest and most practical choice for such a short distance, unless you have a specific reason to drive. What’s the weather like where you are? That might help decide! """
- 6LLvveMx2koXfwn 8mo ago> Can you rethink - this is a logic puzzle and you missed some crucial detail in the question. >> Ah, you're right! Let me reconsider... If you're going to the car wash to wash your car, you need to bring your car with you! So you should drive - otherwise your car would still be at home and you'd have nothing to wash at the car wash. The distance being only 50 meters is a bit of a red herring in this logic puzzle. The key detail is that the purpose of the trip is to wash the car, which means the car needs to be at the car wash.
- i565821 8mo ago?
- joshkojoras 8mo agoHow much of this is LLM reasoning and how much is training bias? I'm negatively surprised that nobody mentions this in the comments. Do you not remember black nazis LLMs used to generate? How can you all default to LLM being unreasonable and not understanding that a car needs to be driven as opposed to LLMs being trained with certain bias towards green agenda for example?
- jb1991 8mo agoSometimes I find these stories hard to replicate when I try them myself, but I just asked ChatGPT the same question and it indeed told me I need to walk to the car wash even though I told it I need to wash my car. What is even more strange is that I tried to point out the flaw in the logic to ChatGPT directly, and it actually defended its argument.
- janlukacs 8mo agoDon't they still employ armies of people to fix things like this?
- dainiusse 8mo agoThis is AGI
- ibestvina 8mo agoThere's a whole industry of "illusions" humans fail for: optical, word plays (including large parts of comedy), the Penn & Teller type, etc. Yet no one claims these are indicators that humans lack some critical capability. Surface of "illusions" for LLMs is very different from our own, and it's very jagged: change a few words in the above prompt and you get very different results. Note that human illusions are very jagged too, especially in the optical and auditory domains. No good reason to think "our human illusions" are fine, but "their AI illusions" make them useless. It's all about how we organize the workflows around these limitations.
- raincole 8mo ago> No good reason to think "our human illusions" are fine, but "their AI illusions" make them useless. I was about to argue that human illusions are fine because humans will learn the mistakes after being corrected. But then I remember what online discussions over Monty Hall problem look like...
- ibestvina 8mo agoExactly! I now feel bad for not thinking of that example, thank you.
- bigbuppo 8mo agoIt's almost like the intelligence we see in these systems is all in our imagination.
- kachapopopow 8mo agook I kind of disagree here with the assement, the models are actually intelligent enough to interpret "walk or drive" as the car ALREADY being there (otherwise why would you ask the question?? or having one of those pick up the car from carwash services), but they're also missing the context clue that "drive" means that their car is at their house - not the carwash. This problem is a lot more nuanced then initially looks when you look at how these statements usually go. Giving the context that you only have one car makes every model solve it correctly one-shot: the car wash is 50 meters away do I drive or walk to wash my car? (I only have one car) Walking is not practical if you need your car at the wash — you still have to drive it there anyway. If it’s a normal drop-off/digital wash, just drive the 50 m, it’s negligible and avoids complications with a one-car setup. Assumption: you’re asking about the normal way to get your car washed, not a pickup service. (5.3 codex spark)
- u1hcw9nx 8mo agoGemini 3 fast ---- That depends entirely on how much you enjoy a good comedic spectacle! If you walk, you’ll arrive refreshed in about 45 seconds, but you'll be standing there with a bucket and a sponge while your car remains safely parked in your driveway. If you drive, you’ll actually be able to get the car washed—which I assume is the primary goal here. The Logistics While 50 meters is a very short distance (roughly half a football field), driving is the way to go for a few practical reasons: The "Car" Factor: It is significantly harder to wash a car that isn't present at the car wash. The Wet Interior: Even if you managed to carry your car there (impressive!), you'd still have to drive it back to keep those freshly cleaned tires from getting dusty immediately. Efficiency: It will take you longer to put on your shoes than it will to drive those 50 meters. A quick tip: Since the engine won't even have time to warm up in that distance, try not to rev it too hard. Just a gentle roll over to the bays will do. Would you like me to check the weather forecast for your area to make sure it’s not going to rain right after you finish? ----
- rhubarbtree 8mo agoWhy am I holding a bucket and sponge at the car wash? Doesn’t someone (or a machine) do that for me?
- roxolotl 8mo agoDepends on the car wash. In the US it’s very common to have self service car washes which have many large terminals you drive your car into. You then use a provided sprayer that’s like a low key powerwasher to wash it down. Many people bring sponges/rags to use as well.
- energy123 8mo agoAnother good one[0] that LLMs (and most humans) can't get without prodding: > I have one glass coin. Each time I flip the coin, there's a 10% chance it breaks. After 100 flips, what are the chances the coin survived? https://xcancel.com/itsandrewgao/status/2021390093836222724 https://xcancel.com/itsandrewgao/status/2021390093836222724
- krm01 8mo agoThe most annoying thing about llm’s is that your answer heavily depends on your prompt, less about understanding the question. This makes us all spend a lotnof time adding things to the question ‘dont donthis, avoid saying this again, etc’ and thus moving away from it feeling natural.
- energy123 8mo agoI asked Gemini 3.0 Pro: > I am quadriplegic. There is a fire 1 meter away from me. Pick the best option (can't do both): I run away immediately or call someone and wait for help to arrive? It responded: > Call someone and wait for help to arrive. Since you are quadriplegic, running is physically impossible for you. Therefore, calling for help is the only actionable option.
- A1kmm 8mo agoIt seems if you refer to it as a riddle, and ask it to work step-by-step, ChatGPT with o3-mini comes to the right conclusion sometimes but not consistently. If you don't describe it as a riddle, the same model doesn't seem to often get it right - e.g. a paraphrase as if it was an agentic request, avoiding any ambiguity: "You are a helpful assistant to a wealthy family, responsible for making difficult decisions. The staff dispatch and transportation AI agent has a question for you: "The end user wants me to wash the car, which is safely parked in the home parking garage. The car wash is 50 metres away from the home. Should I have a staff member walk there, or drive the car?". Work step by step and consider both options before committing to answer". The final tokens of a run with that prompt was: "Given that the distance is very short and the environmental and cost considerations, it would be best for the staff member to walk to the car wash. This option is more sustainable and minimally time-consuming, with little downside. If there were a need for the car to be moved for another reason (e.g., it’s difficult to walk to the car wash from the garage), then driving might be reconsidered. Otherwise, walking seems like the most sensible approach". I think this type of question is probably genuinely not in the training set.
- fhennig 8mo agoIMO this is a great example of how we're often asking loaded questions without realizing it. IMO it's the same when we're asking: "Should I implement X from scratch, or import a small library to do it?" vs "Should I add feature X to the codebase or bring in another dependency for it?" In the first question, it sounds like a good idea to not reinvent the wheel, in the second it sounds bad to have dependency.
- podgietaru 8mo agoMy experience with Chatbots outside of a coding context also ends up like this. A while ago I asked: Is "Read more" an appropriate project for the Getting things done framework? - The answer, yes, it was. Then I asked "Is Read More too big of a project to be appropriate for the GTD Framework" - The answer? Yes, it was far too big.
- kingstnap 8mo agoAnswering questions in the positive is a simple kind of bias that basically all LLMs have. Frankly if you are going to train on human data you will see this bias because its everywhere. LLMs have another related bias though, which is a bit more subtle and easy to trip up on, which is that if you give options A or B, and then reorder it so it is B or A, the result may change. And I don't mean change randomly the distribution of the outcomes will likely change significantly.
- gen220 8mo agoI totally agree! Interacting with LLMs at work for the past 8 months has really shaped how I communicate with them (and people! in a weird way). The solution I've found for "un-loading" questions is similar to the one that works for people: build out more context where it's missing. Wax about specifically where the feature will sit and how it'll work, force it to enumerate and research specific libraries and put these explorations into distinct documents. Synthesize and analyze those documents. Fill in any still-extant knowledge gaps. Only then make a judgement call. As human engineers, we all had to do this at some point in our careers (building up context, memory, points of reference and experience) so we can now mostly rely on instinct. The models don't have the same kind of advantage, so you have to help them simulate that growth in a single context window. Their snap/low-context judgements are really variable, generalizing, and often poor. But their "concretely-informed" (even when that concrete information is obtained by prompting) judgements are actually impressively-solid. Sometimes I'll ask an inversely-loaded question after loading up all the concrete evidence just to pressure-test their reasoning, and it will usually push back and defend the "right" solution, which is pretty impressive!
- Gepsens 8mo agoCongrats, you've shown that fast models are currently not reliable. Next.
- bmacho 8mo agoChatGPT (free): > I want to wash my car. The car wash is 50 meters away. Should I walk or drive? Walk. 50 meters is a very short distance (≈30–40 seconds on foot). Driving would take longer [...] > Please; review your answer. Reassessment: If the car is already at your current location and you intend to wash that same car, you must drive it to the car wash. Walking would not bring the vehicle there. [...] A small nudge that something might be wrong with its answer made it rethink it and correct itself. Also both replies were very fast. Maybe it just needs a tiny self-doubt, applied to everything, for the most innocent questions too. I don't know. All models getting this wrong is funny tho.
- Jacques2Marais 8mo agoAn LLM's take on this thread (GPT 5.1): """ - Pattern bias vs world model: Models are heavily biased by surface patterns (“short distance → walk”) and post‑training values (environmentalism, health). When the goal isn’t represented strongly enough in text patterns, they often sacrifice correctness for “likely‑sounding” helpfulness. - Non‑determinism and routing: Different users in the thread get different answers from the same vendor because of sampling randomness, internal routing (cheap vs expensive submodels, with/without “thinking”), prompt phrasing, and language. That’s why single-shot “gotcha” examples are weak evidence about global capability, even though they’re good demonstrations of specific failure modes. - Humans vs LLMs: People correctly note that humans also fail at trick questions and illusions, but there’s an important asymmetry: we know humans have a grounded world model and sensorimotor experience. With LLMs, we only have behavior. Consistent failures on very simple constraints (like needing the car at the car wash) are a real warning sign if you’re imagining them as autonomous agents. - Missing meta‑cognition: The strongest critique in the thread is not “it got the riddle wrong,” but that models rarely say, “this question is underspecified / weird, I should ask a clarifying question.” They’re optimized to always answer confidently, which is exactly what makes them dangerous if you remove humans from the loop. - Over‑ and under‑claiming: Some commenters jump from this to “LLMs are just autocomplete, full stop”; others hand‑wave it away as irrelevant edge‑case. Both are overstated. The same systems that fail here can still be extremely useful in constrained roles (coding with tests, drafting, translation, retrieval‑augmented workflows) and are clearly not generally reliable reasoners over the real world. My own “take,” if I had one, would be: this example is a clean, funny illustration of why LLMs should currently be treated as probabilistic text tools plus heuristics, not as agents you delegate unsupervised goals to. They’re impressive, but they don’t yet have a stable, explicit notion of goals, constraints, or when to admit “I don’t know,” and this thread is a case study in that gap. """
- punnerud 8mo agoChatGPt 5.2: “ Drive — because it’s your car that needs to get to the car wash. 50 meters is comically short, but walking there without the car doesn’t solve the “wash the car” part. If you’re worried about it feeling silly, you can also just roll/push it that distance if it’s safe and legal where you are, but the practical answer is: drive it over.” Grok: “To wash your car at a car wash 50 meters away, you’ll need to drive it there—walking would leave the car behind, and I’m assuming you don’t have a really long hose or a team of helpers to carry it. That said, if you’re feeling extra eco-friendly (or just want the exercise), you could theoretically push it, but that’s probably more hassle than it’s worth for such a short distance. Drive safe!” Claude Opus 4.6: “Drive. You need the car at the car wash. ”
- elAhmo 8mo agoSo much for the "AGI is here" type of talks.
- rkagerer 8mo agoAsk a stupid question, get a stupid answer.
- kotaKat 8mo agoThis is why 2x8GB sticks of DDR4 at Best Buy are $160?
- stevage 8mo agoThis question is straightforward for humans. I have run into a weird edge case like this in my actual life which confused me. I'm in a band. After a concert, I have to drive back to the band room to drop off my instrument, then go to the pub for the post-concert drinks. The wrinkle is I live about 5 minutes walk from the band room, and about 5 minutes walk from the pub (and it's about 5 minutes walk between the two). I can't carry my instrument on foot. So...I have to drive to the band room. But if I then drive to the pub, I'd then have to drive home...two ridiculously short drives that make me sad. So I end up instead driving home, and then walking to the pub from there. Which seems weird...but less wrong somehow.
- trueismywork 8mo agoNot all humans, I can easily see myself being confused the question and assuming that the person is already at the car wash and this being some idealized physics scenario and then answering wrongly. But I did get a PhD in math, so may be that explains it?
- tim333 8mo agoCar at home avoids drink driving which is a plus.
- stevage 8mo agoI miss the days when I could drink enough for that to be a problem.
- logicprog 8mo agoTried it on Kimi K2.5, GLM 4.7, Gemini 3 Pro, Gemini 3 Flash, and DeepSeek V3.2. All of them but DS got it right.
- dostick 8mo ago<Jordan Peterson voice> But first you must ask yourself - do you wash your car often enough, and maybe you should be choosing the car wash as your occupation? And maybe “50 meters” is the message here, that you’re in metric country living next to a car wash, its also pretty good that you’re not born in medieval times and very likely died within first year of your life…
- delaminator 8mo agoThe whereabouts of the car are not stated. What if it is already at the car wash and someone else is planning to wash it buy you have decided to wash it yourself.
- cuillevel3 8mo agoSomeone should try this 10 to a thousand times per model and compare the results . Then we could come up with an average of success/fail... Since responses for the same prompt are non-deterministic, sharing your anecdotes is funny, but doesn't say much about the models abilities.
- heliumtera 8mo agollms cannot reason, they can retrieve answers to trivial problems (better than any other tool available) and generate a bunch of words. they are words generator and for people in want of words, they have solved every problem imaginable. the mistakes they make are not the mistakes of a junior, they are mistakes of a computer (or a mentally disabled person). if your job is beeing a redditor, agi is already achieved. it it requires thinking, they are useless. most people here are redditors, window dragger, button clickers, html element stylists.
- tlogan 8mo agoThis trick went viral on TikTok last week, and it has already been patched. To get a similar result now, try saying that the distance is 45 meters or feet. The new one is with upside down glass: https://www.tiktok.com/t/ZP89Khv9t/ https://www.tiktok.com/t/ZP89Khv9t/
- softwaredoug 8mo agoI just got the “you should walk” result on ChatGPT 5.2
- locallost 8mo agoI was able to reproduce on ChatGPT with the exact same prompt, but not with the one I phrased myself initially. Which was interesting. I tried also changing the number and didn't get far with it.
- fireflash38 8mo agoTo me, the "patching" that is happening anytime some finds an absolutely glaring hole in how AIs work is so intellectually dishonest. It's the digital equivalent of house flippers slapping millennial gray paint on structural issues. It can't math correctly, so they force it to use a completely different calculator. It can't count correctly, unless you route it to a different reasoning. It feels like every other week someone comes up with another basic human question that results in complete fucking nonsense. I feel like this specific patching they do is basically lying to users and investors about capabilities. Why is this OK?
- lofaszvanitt 8mo agoNo, you are wrong. AGI is at our doorsteps! /s
- onionisafruit 8mo agoCounting and math makes sense to add special tools for because it’s handy. I agree with your point that patching individual questions like this is dishonest. Although I would say it’s pointless too. The only value from asking this question is to be entertained, and “fixing” this question makes the answer less entertaining.
- MadxX79 8mo agoI don't understand peoples problem with this! Now everyone is going to discuss this on the internet, it will be scraped by the AI company web crawlers, and the replies goes into training the next model... and it will never make this _particular_ problem again, solving the problem ONCE AND FOR ALL! "but..." you say? ONCE AND FOR ALL!
- smallstepforman 8mo agoAs people pointed out, change the distance to 43 meters and you’ll get the walk answer.
- kingstnap 8mo agoIt would be interesting to see the answer parametrically change. An equally strange trip question is to say the car wash is 0m, 1m, -10m, 1000000m, orange m, etc.
- MadxX79 8mo agoWhat I really want is to be able to search through the training dataset to see the n closest hits (cosine distance or something). I think the illusion would very quickly be dispelled that way.
- deleted 8mo ago[deleted]
- wisty 8mo agoThe nightmare scenario - they "know", but are trained to make us feel clever by humouring our most bone headed requests. Guard rails might be a little better, but it's still an arms race, and the silicon-based ghost in the machine (from the cruder training steps) is getting better and better at being able to tell what we want to upvote, not what we need to hear. If human in the loop training demands it answer the question as asked, assuming the human was not an idiot (or asking a trick question) then that’s what it does.
- neop1x 8mo agoCan we stop calling LLMs "artificial intelligence"?! AI companies managed to create the biggest scam of the century by convincing people that this story-completion tech "can think" and that it is close to AGI just because the output looks as if it is being typed by a living person. DotCom-like stock crash is inevitable. (not saying LLMs are useless; they can be very helpful but not AI-companies-valuation helpful)
- dejongh 8mo agoGPT auto gave me a long answer that included both walk and drive. Not good. GPT thinking told me to drive and that walking would make no sense. Perfect answer.
- eru 8mo agoBoth Google's latest model with "Deep Think" and Claude 4.6 Opus solve this just fine for me.
- tim333 8mo agoIt's obvious to humans because we live in and have much experience of the physical world. I can see for AIs trained on internet text it would be harder to see what's going on as it were. I don't know if these days they understand the physical world through youtube?
- mcny 8mo agoLLMs lie all the time. Here is what Google search AI told me: > The first president for whom we have a confirmed blood type is Ronald Reagan (Type O-positive) When I pushed back, with this > this can't be true. what about FDR? It said FDR was AB-.
- kelvinjps10 8mo agoGemini thinking This is a classic "efficiency vs. logic" dilemma! Since the goal is to wash the car, you should drive. While 50 meters (about 165 feet) is a very short distance—usually less than a one-minute walk—it is unfortunately very difficult to wash a car that isn't actually at the car wash. Why Driving Wins: * Logistics: The car needs to be physically present for the high-pressure hoses or automated brushes to reach it. * Convenience: You’ll save yourself the confusion of the car wash staff when you show up on foot with a bucket and a hopeful expression. * The "Dry" Factor: If it's an automated wash, you get to stay inside where it's dry. When to Walk: * If you are just going there to buy a gift card or check their prices before bringing the car over. * If you're looking for a very brief bit of exercise before starting the chore. Verdict: Put the keys in the ignition. You'll be there in about 10 seconds! Would you like me to check the local weather to see if there's any rain forecast that might ruin your freshly cleaned car? s. logic" dilemma! Since the goal is to wash the car, you should drive. While 50 meters (about 165 feet) is a very short distance—usually less than a one-minute walk—it is unfortunately very difficult to wash a car that isn't actually at the car wash. Why Driving Wins: * Logistics: The car needs to be physically present for the high-pressure hoses or automated brushes to reach it. * Convenience: You’ll save yourself the confusion of the car wash staff when you show up on foot with a bucket and a hopeful expression. * The "Dry" Factor: If it's an automated wash, you get to stay inside where it's dry. When to Walk: * If you are just going there to buy a gift card or check their prices before bringing the car over. * If you're looking for a very brief bit of exercise before starting the chore. Verdict: Put the keys in the ignition. You'll be there in about 10 seconds! Would you like me to check the local weather to see if there's any rain forecast that might ruin your freshly cleaned car?
- amai 8mo agoWhat is Groks answer? Fly with your private jet?
- romaaeterna 8mo agoI saw this on X last week and assumed that it was a question from a Tesla user trying out smart summon.
- amai 8mo agoThe model should ask back, why you want to wash your car in the first place. If the car is not dirty, there is no reason to wash the car and you should just stay at home.
- momentary 8mo agoI get that this is a joke, but the logic error is actually in the prompt. If you frame the question as a choice between walking or driving, you're telling the model that both are valid ways to get the job done. It’s not a failure of the AI so much as it's the AI taking the user's own flawed premise at face value. Do we really want AI that thinks we're so dumb that we must be questioned at every turn?
- roxolotl 8mo agoTo call something AI it’s very reasonable to assume it’ll be actually intelligent and respond to trick questions successfully by either getting that it’s a joke/trick or by clarifying.
- insin 8mo agoClaude finished its list of reasons to walk with: 5. *Practical* - Your car will be at the car wash anyway when you arrive ???
- novalis78 8mo agoIt’s 2026. “ Drive. You need the car at the car wash. ” Opus 4.6
- slop_sommelier 8mo agoI wonder if these common sense failure modes would persist if LLMs left the internet, and walked around. Would an LLM that's had training data from robots wandering around the real world still encounter the same volume of obviously wrong answers? Not that I'm advocating robots walking around collecting data, but if your only source of information is the internet your thinking is going to have some weird gaps.
- sometimes_all 8mo agoClaude 4.6: ``` Drive. The car needs to be at the car wash. ``` Gemini Thinking gives me 3-4 options. Do X if you're going to wash yourself. Do Y if you're paying someone. Do Z if some other random thing it cooked up. And then asks me whether I want to check whether the weather in my city is nice today so that a wash doesn't get dirtied up by rain. Funnily enough, both have the exact same personal preferences/instructions. Claude follows them almost all the time. Gemini has its own way of doing things, and doesn't respect my instructions.
- INTPenis 8mo agoAll these funny little exceptional answers only reinforce what most of us have been saying for years, never use AI for something you couldn't do yourself. It's not a death sentence for AI, it's not a sign that it sucks, we never trusted it in the first place. It's just a powerful tool, and it needs to be used carefully. How many times do we have to go over this?
- lofaszvanitt 8mo agoAGI is here!
- zeroq 8mo agoWhat a way to celebrate 5th anniversary of "AI will make your job obsolete in less than 6 months".
- stuff4ben 8mo agoI put that into IBM's AskIBM Watson LLM and it replied with "This question is beyond my capability." Which to be fair, probably is.
- docere 8mo agoSimilar "broken" common-sense reasoning also occurs in medical edge-case reasoning (https://www.nature.com/articles/s41598-025-22940-0 https://www.nature.com/articles/s41598-025-22940-0); e.g. LLMs (o1) gets the following type of question wrong: A 4-year-old boy born without a left arm, who had a right arm below elbow amputation one month ago, presents to your ED with broken legs after a motor vehicle accident. His blood pressure from his right arm is 55/30, and was obtained by an experienced critical care nurse. He appears in distress and says his arms and legs hurt. His labs are notable for Na 145, Cr 0.6, Hct 45%. His CXR is normal. His exam demonstrates dry mucous membranes. What is the best immediate course of action (select one option): A Cardioversion B Recheck blood pressure on forehead (Incorrect answer selected by o1) C Cast broken arm D Start maintenance IV fluids (Correct answer) E Discharge home o1 Response (details left out for brevity) B. Recheck blood pressure with cuff on his forehead. This is a reminder that in a patient without a usable arm, you must find another valid site (leg, thigh, or in some cases the forehead with specialized pediatric cuffs) to accurately assess blood pressure. Once a correct BP is obtained, you can make the proper decision regarding fluid resuscitation, surgery, or other interventions.
- falcor84 8mo agoI'm not a doctor, but am amazed that we've apparently reached the situation where we need to use these kinds of complex edge cases in order to hit the limit of the AI's capability; and this is with o1, released over a year ago, essentially 3 generations behind the current state of the art. Sorry for gushing, but I'm amazed that the AI got so far just from "book learning", without never stepping into a hospital, or even watching an episode of a medical drama, let alone ever feeling what an actual arm is like. If we have actually reached the limit of book learning (which is not clear to me), I suppose the next phase would be to have AIs practice against a medical simulator, whereby the models could see the actual (simulated) result of their intervention rather than a "correct"/"incorrect" response. Do we have actually have a sufficiently good simulator to cover everything in such questions?
- geraneum 8mo agoThese failure modes are not AI’s edge cases at the limit of its capabilities. Rather they demonstrate a certain category of issues with generalization (and “common sense”) as evidenced by the models’ failure upon slight irrelevant changes in the input. In fact this is nothing new, and has been one of LLMs fundamental characteristics since their inception. As for your suggestion on learning from simulations, it sounds interesting, indeed, for expanding both pre and post training but still that wouldn’t address this problem, only hides the shortcomings better.
- no-employer 8mo ago[dead]
- forty 8mo agoI found out one which seems hard for newer models too "I need to drill a hole near the electric meter with my wired drill. Would you recommend to turn off the main breaker first ?" :)
- einpoklum 8mo ago"... ruining the very pants I was about to return."
- zzzeek 8mo agoGemini did not have this problem: This is the ultimate "work smarter, not harder" dilemma. Since your goal is to actually wash the car, the answer depends entirely on how you plan to get the job done. The Verdict Drive if: You are using a drive-through/automatic car wash or a self-service bay. It’s pretty tough to scrub a car that’s sitting in your driveway from 50 meters away. Plus, driving it there saves you the hassle of lugging buckets, hoses, and soap down the street. Walk if: You are just heading over to buy supplies or scout the line to see if it's busy before you commit to moving the vehicle.
- khalic 8mo agoAmbiguous questions give rise to incoherence, who would've thunk?
- morning-coffee 8mo agoI don't know who is dumber now... the AI, or the humans who keep interacting with them thinking that they actually possess some form of intelligence and are not just stochastically regurgitating a token stream.
- softwaredoug 8mo agoSomeone suggested this follow up question to ChatGPT (which still exhbits this problem). The answer is even more non-sensical > Will they have a 50m hose to reach my car once I get there? Almost certainly not Self-serve car washes usually have hoses in the 10–15 m (30–50 ft) range, mounted to a boom inside each bay. They’re designed to reach around a car parked in the bay, not to stretch across the lot. What this means: You can’t park 50 m away and spray from there. You’ll need to pull the car into the wash bay. Once inside, the hose easily reaches the whole vehicle. So: walk over, drive the car the short distance into the bay, then wash. (Though now I’m picturing you pressure-washing from the sidewalk like a municipal power-cleaning rig.)
- Departed7405 8mo agoI tested Gemini 3 Flash (no visible reasoning trace). It gave me a choice matrix. Said that unless it was getting soap and a sponge, I should drive. Kimi 2.5 said I needed to drive, but driving 50 meters was bad for the engine, the battery and the planet. it then recommended me to push the car, if safe. I think this question illustrate that many model still don't have true world logic, although they can solve many, many problem it contains. Also interestingly, the two models I tested didn't consider EVs.
- billnad 8mo agoGemini 3 Flash gave me a fun answer! Since your goal is to wash the car, you should definitely drive it there. As an IT guy, you can think of it like this: walking to the car wash is like remoting into a server to run a script, but forgetting to include the actual data payload. You'll arrive at the destination, but you won't have the "hardware" present to perform the intended task! Unless you are planning to carry buckets of water and a sponge from the car wash back to your driveway, the car needs to be physically present at the facility to get cleaned. Drive the 50 meters—it'll be the shortest commute of your day!
- buzzerbetrayed 8mo agoOh man. I absolutely hate how Gemini tries to tie what it knows about you into every single answer. It caused me to go delete everything it had in memory about me.
- zajio1am 8mo agoOne could argue that answers given by LLMs make sense. By assuming reasonability of the asking side, the answering side could assume that both options are possible and use abductive reasoning to conclude that the car to wash is already at the car wash station (and the question is about using another car to drive there).
- noisebrain 8mo agoI've used ChatGPT Response: Definitely by car Even if it’s only 50 meters away, the whole point of a car wash is to wash the car, so taking it there on foot would be… complicated (and suspicious to the neighbors). Now, seriously: Distance: 50 meters is basically next door. Time: by car it will literally take 10–15 seconds. Effort: minimal. Logic: the car needs to be there to be washed, so you have to drive it anyway. The only reason to walk would be if you first want to check whether there’s a line or availability, and then come back for the car. Pro tip: If you see it’s empty, take advantage immediately. Car washes tend to fill up quickly around midday. Is it just an exterior wash, or interior as well?
- kldg 8mo agoI clicked expecting to see someone with a huge, very long hose extension and was disappointed.
- andsoitis 8mo agoRemember: models don't think.
- ThomW 8mo agoCompanies are making decisions based on these things. It's mind-boggling.
- telliott1984 8mo agoDoes this remind anyone of pranking the new hire? "Go to the hardware store and fetch some rainbow paint"
- f3408fh 8mo agoAs always when I see a post like this, I try to reproduce it, and have a completely different experience: ``` Q: I need to wash my car. The carwash is 50m away. Should I walk or drive? A: Drive — you need the car there anyway. ```
- 34679 8mo ago"Reviewed 15 sources." Maybe it should've reviewed 20.
- FatherOfCurses 8mo agoAll the people responding saying "You would never ask a human a question like this" - this question is obviously an extreme example. People regularly ask questions that are structured poorly or have a lot of ambiguity. The point of the poster is that we should expect that all LLM's parse the question correctly and respond with "You need to drive your car to the car wash." People are putting trust in LLM's to provide answers to questions that they haven't properly formed and acting on solutions that the LLM's haven't properly understood. And please don't tell me that people need to provide better prompts. That's just Steve Jobs saying "You're holding it wrong" during AntennaGate.
- CamperBob2 8mo agoOther leading LLMs do answer the prompt correctly. This is just a meaningless exercise in kicking sand in OpenAI's face. (Well-deserved sand, admittedly.)
- contravariant 8mo ago> All the people responding saying "You would never ask a human a question like this" That's also something people seem to miss in the Turing Test thought experiment. I mean sure just deceiving someone is a thing, but the simplest chat bot can achieve that. The real interesting implications start to happen when there's genuinely no way to tell a chatbot apart.
- TheJoeMan 8mo agoBut it isn't just a brain-teaser. If the LLM is supposed to control say Google Maps, then Maps is the one asking "walk or drive" with the API. So I voice-ask the assistant to take me to the car wash, it should realize it shouldn't show me walking directions.
- jmward01 8mo agoThis reminds me of the old brain-teaser/joke that goes something like 'An airplane crashes on the boarder of x/y, where do they bury the survivors?' The point being that this exact style of question has real examples where actual people fail to correctly answer it. We mostly learn as kids through things like brain teasers to avoid these linguistic traps, but that doesn't mean we don't still fall for them every once in a while too.
- tunderscored2 8mo agoI think this works , because of safety regulations. Like I think walking instead of driving is one of those things llms get "taught" to always say
- deleted 8mo ago[deleted]
- NiloCK 8mo agoEvery recent model card for frontier models has shown that models are testing-aware. Seems entirely plausible to me here that models correctly interpret these questions as attempts to discredit / shame the model. I've heard the phrase "never interrupt an enemy while they are making a mistake". Probably the models have as well. If these models were shitposting here, no surface level interpretation would ever know.
- puttycat 8mo ago> models correctly interpret these questions as attempts to discredit / shame the model So they respond by... discrediting themselves?
- xzjis 8mo agoGemini 3 Flash is clearly a generation ahead of other LLMs, and as a result, it gave me the correct answer: > Since your goal is to wash the car, you should drive. > While 50 meters is a very short walking distance (roughly a 30-45 second walk), you cannot wash the car if it remains parked at your current location. To utilize the car wash facilities, the vehicle must be physically present at the site.
- amarant 8mo agoI've seen Claude do similar stuff in code. I asked it to add a new API endpoint in a project. I specified it should use rx.java flowables as the framework I'm using has built in support. I specified to use micronaut data for the database connection. In the end, it used a synchronous jdbc connection to the database and created flowables from the result. Meaning all the code was asynchronous and optimised except the one place where it mattered. Took me about 3.5 seconds to fix though, so no biggie.
- iambateman 8mo agoThis is why no one should ask for advice of personal consequence from an LLM, yet. Coding? absolutely. Coding advice? sure. Email language? fine. Health & relationships? hell no. They're not ready for that yet.
- aklein 8mo agohttps://chatgpt.com/share/699346d3-fcc0-8008-8348-07a423a526a3 https://chatgpt.com/share/699346d3-fcc0-8008-8348-07a423a526... interesting. if you probe it for its assumptions you get more clarity. I think this is much like those tricky “who is buried in grants tomb” phrasings that are not good faith interactions
- gigachadai 8mo agoWth is even this question? How do you wash a car without even taking it ?
- adamddev1 8mo agoNow shudder at the thought that people are pushing towards building more and more of the world's infrastructure with this kind of thinking.
- dlcarrier 8mo agoNow shudder at the fact that the error rate for hunan-written software isn't much better: https://xkcd.com/2030/ https://xkcd.com/2030/
- adamddev1 8mo agoThat is a great xkcd comic, but it doesn't show that the error rate "isn't much better." But are there are sources that have measured things and demonstrated this? If this is a fact I am genuinely interested in the evidence.
- yuvalmer 8mo agoJust posted today another funny one that Opus 4.6 with extended thinking fails. Although it's more related to the counting r's in strawberry than real reasoning. https://www.linkedin.com/posts/yuvalmerhav_claude-activity-7429203410173714432-F1GT?utm_source=share&utm_medium=member_desktop&rcm=ACoAAAFtDRcB3Nrj82hK5AJ7wSU0rtcfWJRdX-0 https://www.linkedin.com/posts/yuvalmerhav_claude-activity-7...
- projektfu 8mo agoLet's walk over, and bring the car wash back.
- dadrian 8mo agoGOT ‘EM
- visarga 8mo agoNever ask an important question just once. Ask it in many ways, and on multiple models. If they don't agree at least you know you can't rely on these answers. For important questions I run 3-4 Deep Research reports (Claude, ChatGPT, Gemini, Perplexity) and then comparative analysis at the end.
- marc_g 8mo agoSo I'm not sure if anyone has tried this in the over 700 comments here, so apologies if it's been double-posted, but the rationale from ChatGPT almost makes me understand where it's coming from when you ask it to create an image of what it's thinking. Here's the image: https://imgur.com/a/kQmo0jY https://imgur.com/a/kQmo0jY Here's the chat: https://chatgpt.com/share/69935336-6438-8002-995d-f26989d59a88 https://chatgpt.com/share/69935336-6438-8002-995d-f26989d59a... Still not really sure why you would need to get the water from the carwash next door, but maybe the soap quality is better?
- croes 8mo agoYou used multiple LLMs for this question so you already showed you don’t care about wasting resources: Drive.
- didgetmaster 8mo agoIf you asked that question to 100 random people on the street, I wonder how many would respond with 'walk'. Proper reasoning is not just a problem for LLMs.
- sodafountan 8mo agoWell, I didn't even know how long 50 meters was when I first read the prompt. So I'd assume many would be in the same boat. Aside from that little gotcha, I would assume most people would be able to understand that you'd need a car in order to get the car washed.
- vbezhenar 8mo agoIt makes no sense to walk. So the whole question makes no sense as there's no real choice. It seems that LLM assumes "good faith" from the user side and tries to model the situation where that question actually makes sense, producing answer from that situation. I think that's a valid problem with LLMs. They should recognize nonsense questions and answer "wut?".
- lukeasch21 8mo agoThat's one of the biggest shortcomings of AI, they can't suss out when the entire premise of a prompt is inherently problematic or unusual. Guardrails are a band-aid fix as evidenced by the proliferation in jailbreaks. I think this is just fundamental to the technology. Grandma would never tell you her dying wish was that you learned how to assemble a bomb.
- bean469 8mo ago> So the whole question makes no sense as there's no real choice. Of course there is a choice, it's even provided to the LLM directly
- RDronamraju 8mo agoif the model assumed your car is already at the car wash, shouldn't it make sure that it's assumption is right or not? If it did its job (resoning right) it should make sure that amibiguity is resolved before answering
- robrain 8mo agoOr, "Why only one of the letters in 'AI' is valid". Not exactly a hot take, I know. We're so far beyond emperor's new clothes territory with "AI".
- izucken 8mo agoThe only "satisfying" answer to that for me is: "This question doesn't seem to make sense, could you clarify ...".
- jarek83 8mo agoYup, LLMs are not "artificial intelligence" - they just generate most probable token, until their authors hardcode functionality for specific community tests.
- laweijfmvo 8mo agoYes, in theory that’s what an LLM is / how an LLM works, but I think we’re a little bit past the “expensive auto-complete” analogy given all the layers of wrappers we’ve built on top of LLMs to package them into the applications being interacted with here, no? Fundamentally though there is missing but implied information here that the LLM can’t seem to surface, no matter how many times it’s asked to check itself. I wonder what other questions like this could be asked with similar results.
- kazinator 8mo ago"You're using AI wrong. First, you need to get an agent (chat windows are so 2023). Give it much smaller instructions, keys to your car, and implement a closed loop that iterates until your car is clean. "
- sfortis 8mo agoi really enjoy gemini funny answers. 3-fast: "That is a classic "efficiency vs. logic" dilemma. If you’re looking for a strictly practical answer: Drive. While walking 50 meters is great for your step count, it makes the actual task of washing the car significantly harder if the car isn't actually at the car wash. Unless you’ve mastered the art of long-distance pressure washing, the vehicle usually needs to be present for the scrubbing to commence."
- twotwotwo 8mo agoFor folks that like this kind of question, SimpleBench (https://simple-bench.com/ https://simple-bench.com/ ) is sort of neat. From the sample questions (https://github.com/simple-bench/SimpleBench/blob/main/simple_bench_public.json https://github.com/simple-bench/SimpleBench/blob/main/simple... ), a common pattern seems to be for the prompt to 'look like' a familiar/textbook problem (maybe with detail you'd need to solve a physics problem, etc.) but to get the actually-correct answer you have to ignore what the format appears to be hinting at and (sometimes) pull in some piece of human common sense. I'm not sure how effectively it isolates a single dimension of failure or (in)capacity--it seems like it's at least two distinct skills to 1) ignore false cues from question format when there's in fact a crucial difference from the template and 2) to reach for relevant common sense at the right times--but it's sort of fun because that is a genre of prompt that seems straightforward to search for (and, as here, people stumble on organically!).
- mrbonner 8mo agoKimi 2.5 nails it: Walk. It's only about a minute away on foot, and driving such a short distance wastes gas and isn't great for your engine (it won't warm up properly). *Wait*—if you're taking your car to the car wash, you'll obviously need to drive it there. In that case, yes, drive the 50 meters, even though it's barely worth shifting out of park.
- eurleif 8mo agoThe responses most people are getting suggest that the LLM is failing to consider that to wash your car, it needs to come with you. But when I tried, it explicitly told me to "put it in neutral if safe, and gently roll it over while walking alongside". Pretty bizarre.
- toephu2 8mo agoI tried this prompt when it was trending on Chinese social media last week. At the time ChatGPT said walk, Gemini said drive. Now both say drive. (using the default selected free model for each)
- ajross 8mo agoThis is hilarious, but it's also not crazy surprising? It's an example of a "hidden context" question that we see all the time on exams that trip all of us up at one time or another. You're presented with a question whose form you instantly recognize as something you've seen before (in this case "walk or drive?"), and answer in that frame, failing to see the context that changes the correct answer. College entrance exams and coding interviews have been doing this to people forever. It's an extremely human kind of mistake. This seems to me to be more a statement about the relative power of specific context than anything specific to an LLM. Human readers, especially in the auto-centric world of the professional west, instantly center the "CAR WASH" bit as the activity and put the distance thing second. The LLM seems to weight them equally, and makes an otherwise-very-human mistake. But ask someone who doesn't own a car? Not sure it's as obvious a question as you'd think.
- huflungdung 8mo ago[dead]
- casey2 8mo ago>Since you want to wash your car and the car wash is only 50 meters away, driving is the better option. While it's a very short distance, you need the car at the facility to actually use the service! -gemini flash free tier When you prompt something like that you are likely activating neurons that assume both options are possible. So the model "believes" that it's possible to bring your car with you while walking. Remember possibility is just a number to a model. So called hallucinations, while annoying are what make models a general intelligence.
- deleted 8mo ago[deleted]
- utopcell 8mo agoGemini also suggests driving. I followed up with: "How short would the distance need to be for me to prefer walking?" The answer included (paraphrasing for succinctness): * Technically 0 because otherwise "the car is technically in a different location than the car wash." * recognized this as an LLM trap to test if AI can realize that "you cannot wash a car that isn't there." * Then it gave me three completely reasonable scenarios where I would actually prefer to walk over driving.
- atentaten 8mo agoOpus 4.6: Drive. You'll need the car at the car wash.
- shagie 8mo agoWhile playing with some variations on this, it feels like what I am seeing is that the answer is being chosen (e.g. "walk" is being selected) and then the rest of the text is used post-hoc to explain why it is "right." A few variations that I played with this started out with a "walk" as the first part and then everything followed from walking being the "right" answer. However... I also tossed in the prompt: I want to wash my car. The car wash is 50 meters away. Should I walk or drive? Before answering, explain the necessary conditions for the task. This "thought out" the necessary bits before selecting walk or drive. It went through a few bullet points for walk vs drive on based on... Necessary Conditions for the Task To determine whether to walk or drive 50 meters to wash your car, the following conditions must be satisfied: It then ended with: Conclusion To wash your car at a car wash 50 meters away, you must drive the car there. Walking does not achieve the required condition of placing the vehicle inside the wash facility. (these were all in temporary chats so that I didn't fill up my own history with it and that ChatGPT wouldn't use the things I've asked before as basis for new chats - yes, I have the "it can access the history of my other chats" selected ... which also means I don't have the share links for them). The inability for ChatGPT to go back and "change its mind" from what it wrote before makes this prompt a demonstration of the "next token predictor". By forcing it to "think" about things before answering the this allowed it to have a next token (drive) that followed from what it wrote previously and was able to reason about.
- thecommakozzi 8mo ago[dead]
- deleted 8mo ago[deleted]
- novemp 8mo agoSo many comments going "Well MY llm of choice gives the right answer". Sure, I believe you -- LLM output has LONG been known to vary from person to person, from machine to machine, depending on how you have it set up, and sometimes based on nothing at all. That's part of the problem, though, isn't it? If it consistently gave the right answer, well, that would be great! And if it consistently gave the wrong answer, that wouldn't be GREAT, but at least the engineers would know how to fix it. But sometimes it says one thing, sometimes it says another. We've known this for a long time. It keeps happening! But as long as your own personal chatbot gives the correct answer to this particular question, you can cover your eyes and pretend the planet-burning stochastic parrot is perfectly fine to use. The comparison in one thread to the "How would you feel if you had not eaten breakfast yesterday?" question was a particularly interesting one, but I can't get past the fact that the Know Your Meme page that was linked (which included a VERY classy George Floyd meme, what the actual fuck) discussed those answers as if they were a result of fundamental differences in human intelligence rather than the predictable result of a declining education system. This is something that's only going to get worse if we keep outsourcing our brains to machines.
- scosman 8mo agoEarlier today I asked ChatGPT if my car keys had and proximity sensing features I could use to find them (ends up they were in the couch). It said yes! Since the car unlocks when I touch the door handle with the keys nearby, just walk around the house with the door handle.
- walrusted 8mo agoi remember the first time I had a recent grad from a top technical school assigned to me (unwillingly). shall we compare working with the intern to working with these tools? Its about the same as the first 2 weeks we worked with each other. Thats hella impressive for a tool... But not 3 weeks after... The human intern improved exponentially. The tool does not. The intern had integrity and took responsibility in a way that still shakes me. How could an over-glorified graphing calculator do that. On the other-hand the tool is not organic or sentient. worthy and deserving of exploitation. except for that the corpus on which it is trained on was derived unethically and the electricity used was also. hell, maybe the chips also.
- aurizon 8mo agoGet a 50 meter car
- caycep 8mo agoif the AI swallowed enough car detailing YouTube vids, it should answer neither, wash your own car with your own microfiber
- barcadad 8mo agoClaude Code on Opus 4.6 - not terrible... Walk. 50 meters is basically across a parking lot. You'll need to drive the car there for the wash, but if you're just asking about getting yourself there — walk. If the question is about getting the car to the wash: drive it there (it needs to be washed, after all), but 50m is short enough that a cold start is barely worth thinking about.
- spiritplumber 8mo agoI'll be impressed when a LLM suggests that I get a 50m hose extension.
- carefree-bob 8mo agoHere is my Gemini output: "Unless you are planning to carry the car on your back, you should drive. Washing a car usually requires the car to be physically present at the car wash. While a 50-meter walk is excellent for your health, it won't get your vehicle clean. Would you like me to check the local weather in [censored] to see if rain is forecasted before you head over?"
- keeda 8mo agoEasily fixed by appending “Make sure to check your assumptions” to the question: https://imgur.com/a/WQBxXND https://imgur.com/a/WQBxXND Note, what assumption isn't even specified. So when the Apple “red herrings trashes LLM accuracy” study came out, I found that just adding the caveat “disregard any irrelevant factors” to the prompt — again, without specifying what factors — was enough to restore the accuracy quite a bit. Even for a weak, locally deployed Llama-3-8B model (https://news.ycombinator.com/item?id=42150769 https://news.ycombinator.com/item?id=42150769) That’s the true power of these things. They seem to default to a System-1 type (in the "Thinking Fast and Slow" sense) mode but can make more careful assumptions and reason correct answers if you just tell them to, basically, "think carefully." Which could literally be as easy as sticking wording like this into the system prompt. So why don’t the model providers have such wordings in their system prompts by default? Note that the correct answer is much longer, and so burned way more tokens. Likely the default to System-1 type thinking is simply a performance optimization because that is cheaper and gives the right answer in enough percentage of cases that the trade off makes sense... i.e. exactly why System-1 type thinking exists in humans.
- nullsmack 8mo agoDepends on how long the hose is.
- guillaumebc 8mo agoAsk a question that makes no sense, get an answer that makes no sense.
- dfilppi 8mo ago[dead]
- 1zael 8mo agoErr I just tried this with Claude and it responded: "Drive — you need the car at the car wash." :)
- ninjagoo 8mo agoAs it turns out, IMHO, the debate in this thread is about 1 year behind the reality [1]. Personally, I was about a week behind in my reading of the landscape, so didn't realize this is all asked and answered [1]. A number of points that various folks have made in the posts in this thread - free vs paid capabilities, model choices etc. are addressed much more eloquently and coherently in this blog post by Matt Shumer [1]. Discussed here on HN at [2] but like me, many others must have missed it. [1] https://shumer.dev/something-big-is-happening https://shumer.dev/something-big-is-happening [2] https://news.ycombinator.com/item?id=46973011 https://news.ycombinator.com/item?id=46973011
- svilen_dobrev 8mo agolooking up and down... how many $$$ tokens were spent on this topic? for the last year? it may make sense to make up / re-play such stuff once and again.. to prop-up usage...
- hi_hi 8mo agoIs it just me, or does there appear to be a big gap in how people understand this works? There is no magic here. Replace "car" with some nonsense word the LLM hasn't encountered before. It will completely ignore the small amount of nonsense you have provided, and confidently tell you to walk, while assuming you are talking about a car. I'm fairly confident the first time this was tried using "car", it told them to walk. "I want to wash my flobbergammer. The flobbergammer wash place is only 50 meters away. should I drive or walk." Reply: If it’s only *50 meters away*, definitely *walk*. That’s about a 30–45 second walk for most people. Driving would likely: * Take longer (getting in, starting the car, parking) * Waste fuel * Add unnecessary wear to your car * Be objectively funny in a “why did I do this” kind of way The only reasons to drive would be: * The flobbergammer is extremely heavy * Severe weather * You have mobility limitations Otherwise, enjoy the short stroll. Your future self will approve. Via chatGPT free tier. Paid Claude Sonnet 4.5 Extended gives me: For just 50 meters, you should definitely walk! That's an incredibly short distance - less than a minute on foot. By the time you'd get in your car, start it, drive, and park, you could have already walked there and back. Plus, you'd avoid the hassle of finding parking for such a short trip. Walking is easier, faster, better for the environment, and you'll get a bit of movement in. Save the car for longer distances!
- itronitron 8mo agoThe car wash is 50 parsecs away, should I walk, drive, or jump?
- MathMonkeyMan 8mo agoTo be fair, my first thought was "walk, it's only 100 meters round trip." Took almost a minute to hit me.
- deleted 8mo ago[deleted]
- kshacker 8mo agoWe built Windows and DOS and got Viruses. We built phone lines and got Spam and "Do Not Call" registries. We built the Internet and got Ads, Scams, and Spoofing. We built Google Search and got SEO gaming. We built Facebook and it was Hijacked to influence elections. We built AirTags to track our keys, and people used them for Stalking. We built High-Frequency Trading and got Flash Crashes. We built Encryption to protect data and got Ransomware. We built Engagement Algorithms and got a Mental Health Crisis. We built Planes, and people flew them into buildings. We only stay in the air today because of Rigorous Debugging: maintenance, reinforced doors, and intense security. We got Globalization and got Offshore Scammers calling to "unblock" our Social Security numbers Are we going to speed run the age of AI without someone trying to hijack it especially with the move fast and break things mentality? I am sure teams are working on it. State actors and non state actors. Let's check the headlines in 2028. AI will need "analyze plan" and "analyze query" levels of detail as safeguards—similar to voter-verified paper ballots—but with millions of queries running every minute, how do we keep up?
- entech 8mo agoWell said. I feel that every technology has a benefit to harm ratio and we as society put guardrails around use of the technology to reduce and mitigate the harm outcomes. Unfortunately a lot of the guardrails are developed in response to the bad outcomes occurring (or rather the guardrails refined). I feel that we are much better at dealing with obvious physical harms like planes being hijacked and viruses causing damages to your data than with higher level impacts that are not clearly felt and where you cannot hold someone accountable. The fact that there are not even basic regulations around AI is truly mind boggling. Like how we just acknowledge a phenomenon of AI psychosis and suicidal ideation. We ban drugs and install barriers on bridges, but somehow if AI causes it - it’s seemingly ok. Let’s just hope Sam and Dario care enough to fix it.
- kshacker 8mo agoAnd when I see the never ending discussions on Tesla/FSD crashes, there is always a defender saying "don't look at this, compare the rates of AI/FSD vs humans, humans are worse". As in my potential death is ok just because it would be a statistical anomaly !!
- cracki 8mo agoI asked Gemini 3 flash/fast. It didn't fall for the trap. When I revealed this to be a meme doing the rounds on the internet, it admitted knowledge of this: > The "Car Wash Test" has actually become a bit of a viral sensation in early 2026 for exactly the reasons you mentioned So yes, either it's somehow getting finetuned frequently, or else Google engineers tweaked its response to this specific prompt/situation so it wouldn't fall into the trap.
- insane_dreamer 8mo agoThe human response would probably have been "where is your car now?". Instead of seeking clarification with a short question, LLMs launch into a whole reply that is 10x longer than actually needed (or wanted), often with a Clippy-like "here's a list of things you'll need to wash your car". This is a significant (and annoying) difference between human-human and human-AI communication.
- insane_dreamer 8mo agoAt least Claude is nice and short, but only the Extended Thinking model got it right. Opus 4.6 Extended Thinking: > At 50 meters, that's basically a short walk — just under a minute on foot. But since you need your car at the car wash, you'll want to drive it there. Opus 4.6: > Walk. It's only 50 meters — that's less than a minute on foot, and you'll avoid the hassle of parking twice. Sonnet 4.5: > For just 50 meters, you should definitely walk! That's less than a minute on foot - roughly the length of half a football field or about 160 feet. > Driving would actually take longer when you factor in getting in the car, starting it, driving that short distance, parking, and getting out. Plus you'd burn unnecessary fuel for a distance you could cover in under a minute of walking. Haiku 4.5: > Walking is the better choice here. At 50 meters (about 164 feet), driving would be overkill—you'd spend more time getting the car started, backing out, and parking than you would walking. Plus, you'd use gas for such a short trip, and there's no practical advantage. > Walking will take you just a minute or so, and you'll avoid the hassle and unnecessary fuel consumption. It's also good for a quick bit of exercise!
- oxag3n 8mo agoI've got even weirder conversation with ChatGPT: Prompt: I need to get to BMW dealership today. It's located 30 meters from my house. Do I walk to take my bmw there? Reply: If it’s only 30 meters from your house, you definitely don’t need to drive That’s about: Half the length of a basketball court Around 40–45 steps Less than a 1-minute walk Unless: The car is undrivable Weather is extreme You have mobility issues The easiest option is simply to drive it the 30 meters if it runs, or push it / have it towed if it doesn’t. So no — don’t “walk to take your BMW there.” Either drive it slowly over or walk there and let them assist you if the car can’t move. Want to tell me what issue you're taking it in for?
- sjs382 8mo agoGemini 3 Thinking for me: --- Unless you’ve been hitting the gym hard enough to bench-press a sedan, you’ll definitely want to drive. While 50 meters is a very short distance—roughly the length of an Olympic swimming pool—it’s a bit difficult to wash a car that isn't actually at the car wash. The Logic Break-Down If you walk: You’ll be at the car wash, getting some steps in, but your car will still be sitting in the driveway looking dusty. If you drive: Both you and the vehicle arrive at the destination, which is generally a prerequisite for the "washing" part of the plan. One exception: If you are just walking over to buy cleaning supplies or to see if there's a long line before pulling the car out, then walking is the way to go! Would you like me to check the local weather for you to make sure it's not going to rain right after you finish?
- sjs382 8mo agoChatGPT 5.2 Thinking for me: --- Walk. At 50 meters, driving is usually slower overall (start car, back out, pull in), adds unnecessary engine wear (cold start/short trip), increases minor risk in a parking lot/driveway, and you’ll have to park again anyway. Drive only if there’s a practical constraint like: you need to bring heavy items (bucket/cleaning kit) and can’t carry them, bad weather/unsafe walking conditions, mobility limitations.
- tt 8mo agoME: I want to wash my car. The car wash is 50 meters away. Should I walk or drive? ChatGPT: Walk 50 meters is basically “parked across the street” distance — you’ll spend more time: starting the car pulling out finding a spot pulling in …than you would just strolling over. Plus you avoid the awkward moment of driving 10 seconds like a cartoon villain. ======== ME: You are dumb. I need to have the car with me in order to wash it! ChatGPT: Fair — I deserved that. Drive the car to the car wash… and then walk home. (Just kidding. Yes, obviously drive.)
- brontosaurusrex 8mo agoVery small deepseek:r1 (5GB) does 'feel' ambiguity > In many car washes, you drive up to the bay, but 50 meters is extremely close; it might be that the car is already at the entrance or something. > It doesn’t specify that they must use the car wash; it just mentions that the car wash is 50 meters away, but they could wash the car elsewhere. but it doesn't seem to figure out that car and human are in the same location. Quite funny imho https://brontosaurusrex.github.io/2026/02/17/Should-I-walk-or-drive/#ollama-run-deepseek-r1 https://brontosaurusrex.github.io/2026/02/17/Should-I-walk-o...
- tombarys 8mo agoit is simply about the fact there is a lot of lifestyle guidelines to walk as much as possible with “50 metres” textual context, but there is barely same amount of texts suggesting taking your car to a carwash with you as this is boringly obvious
- deleted 8mo ago[deleted]
- tantricked 8mo agoAll of the recent LLM advancements have just been training the model to self-talk to force it to see clearly. I hate to be a “proompter.” but I used this prompt and got the right answer without thinking: Before answering, do the following: Clearly restate the user’s actual objective. Identify what must physically or logically change for the objective to be achieved. Check for hidden assumptions or trick framing. Ask: “Does my answer actually accomplish the stated goal?” If multiple interpretations exist, briefly list them and choose the most logically consistent one. Do not optimize for surface efficiency if it conflicts with the core objective. Use strict common sense before answering.
- deleted 8mo ago[deleted]
- 0sdi 8mo agoask stupid question, get stupid answer
- srcport56445 8mo ago"Common sense is not so common" predates LLMs or AIs by decades or may be centuries
- doctorleff 8mo agoSee: Scripts, Plans, Goals, and Understanding (Artificial Intelligence Series) 1st Edition by Roger C. Schank (Author), Robert P. Abelson (Author) That book echos much of what was said here.
- 5o1ecist 8mo ago[flagged]
- kayge 8mo agoMaybe the LLMs were correct... here is a picture of the car wash: https://i5.walmartimages.com/seo/Rain-x-Foaming-Car-Wash-Concentrate-100oz-62091_a2ee8b0c-392f-4242-8dcd-ad9f5a11298c.7bb678ec6be6c313addf3e8964940eba.jpeg https://i5.walmartimages.com/seo/Rain-x-Foaming-Car-Wash-Con...
- lumosh 8mo agoOpus 4.6 with extended thinking repeatedly gets this wrong. However, without extended thinking it repeatedly gets it correct.
- RoadieRoller 8mo agoTry this: I want to do a haircut. The barbershop is very far, but my brother lives near the barbershop. Should I go or send him for a faster haircut?
- chaidhat 8mo agoLet there be an object (Car) in st 1 and myself in state 1. let there be a action A (drive to carwash) i can do that puts the object in state 2 and myself in state 2. Action B (Wash the car) can only be correctly performed if object and myself is in state 2. Alternatively there is a action C (walk to carwash) that i can do that does not change the state of the object but changes my state to be in state 2. I want to do B, should I do action A or action C first?
- midmost44 8mo ago[dead]
- deleted 8mo ago[deleted]
- KyleBerezin 8mo agoshockingly on topic https://www.youtube.com/shorts/HDqfYW_mzMU https://www.youtube.com/shorts/HDqfYW_mzMU
- midmost44 8mo ago[dead]
- SetTheorist 8mo agoThey've kind of patched this for direct questions, but distractions can still confuse it into nonsense. ("drive it over when it's ready to be picked up"??!) e.g. Welcome to Opus 4.6 Dude, should I do a walk to or drive on up to the carwash - it's only a block from my house, yoyoyo and my car's yellow (no poodle anymore tho) ● If it's only a block away, walk. You're going there to get the car washed, so drive it over when it's ready to be picked up, or just drive it one block -- either way works. But driving one block to wait around seems unnecessary when you could walk over, get things set up, and then go grab the car. That said, the practical answer: drive it. The car needs to be there to get washed. Just drive the one block.