5 ms·
The case for zero-error horizons in trustworthy LLMs
- charcircuit 6mo agoWhy didn't OpenAI finetune the model to use the python tool it has for these tasks?
- ej88 6mo agoThey do, in the paper they mention they evaluate the LLM without tools
- throwuxiytayq 6mo ago> This is surprising given the excellent capabilities of GPT-5.2. Is this seriously surprising to anyone who knows the absolute minimum about how LLMs parse and understand text?
- dontlikeyoueith 6mo agoNope. It's only surprising to people who still think they're going to build God out of LLMs.
- simianwords 6mo agoIt was surprising to me and when I reviewed the paper, I found serious flaws that calls the fundamental claims into question - they didn't use any reasoning tokens. Any LLM or human will fail at a task like this if not allowed to think.
- dontlikeyoueith 6mo agoCalling "reasoning tokens" "thinking" is a complete confusion of concepts on your part.
- simianwords 6mo agowhy?
- throwuxiytayq 6mo agoMachines can’t think because the bible says nothing about it
- staticshock 6mo agoLLMs seem to me closer to Kahneman's System 1 than to System 2. When understood in this way, it is obvious why LLMs are bad at counting r's in "strawberries". But it also makes ZEH feel like it couldn't possibly be a useful metric, because it's a System 2 evaluation applied to a System 1 system.
- 8note 6mo ago> When understood in this way, it is obvious why LLMs are bad at counting r's in "strawberries". no it doesnt. it makes sense that they cant count the rs because they dont have access to the actual word, only tokens that might represent parts or the whole of the word
- orbital-decay 6mo agoTokenization is a simplistic explanation which is likely wrong, at least in part. They're perfectly fine reciting words character by character, using different tokenization strategies for the same word if forced to (e.g. replacing the starting space or breaking words up into basic character tokens), complex word formation in languages that heavily depend on it, etc. LLMs work with concepts rather than tokens.
- im3w1l 6mo agoA big part of skill aquisition in humans is moving tasks from system 2 to system 1, to free up the very scarce thinking resources for ever more complex tasks, that can then in turn be internalized and handled by system 1.
- derefr 6mo agoFYI, the LLM letter-counting problem has nothing to do with counting per se, and is instead entirely down to LLMs not getting to see your raw UTF-8 byte stream, but rather having a tokenizer intermediating between you and it, chunking your UTF-8 bytes into arbitrary, entirely-opaque-to-the-LLM token groupings. Try it for yourself — under the most popular tokenizer vocabulary (https://tiktokenizer.vercel.app/?model=cl100k_base https://tiktokenizer.vercel.app/?model=cl100k_base), "strawberry" becomes [str][aw][berry]. Or, from the model's perspective, [496, 675, 15717]. The model doesn't know anything about how those numbers correspond to letters than you do! It never gets sat down and told "[15717] <=> [b][e][r][r][y]", with single-byte tokens on the right. (In fact, these single-byte tokens appear in the training data extremely rarely, and so the model doesn't often learn to do anything with them.) Note that LLMs can predictably count the number of r's in "s t r a w b e r r y", because <Count the number of r's in "s t r a w b e r r y"> becomes [Count][ the][ number][ of][ r]['s][ in][ "][s][ t][ r][ a][ w][ b][ e][ r][ r][ y]["]. And that's just a matching problem — [ r] tokens for [ r] tokens, no token-correspondence-mapping needed.
- justinator 6mo agoOne! Two! Five!
- parliament32 6mo ago> This is surprising given the excellent capabilities of GPT-5.2 The real surprise is that someone writing a paper on LLMs doesn't understand the baseline capabilities of a hallucinatory text generator (with tool use disabled).
- coldtea 6mo agoThe real suprise is people saying it's surprising when researchers and domain experts state something the former think goes against common sense/knowledge - as if they got them, and those researcers didn't already think their naive counter-argument already.
- burningion 6mo agoRan this through Qwen3.5-397B-A17B, and the difference between 4 characters and 5 is wild to see: > are the following parenthesis balanced? ((()))) > No, the parentheses are not balanced. > Here is the breakdown: Opening parentheses (: 3 Closing parentheses ): 4 ... following up with: > what about these? ((((()))) > Yes, the parentheses are balanced. > Here is the breakdown: Opening parentheses (: 5 Closing parentheses ): 5 ... and uses ~5,000 tokens to get the wrong answer.
- BugsJustFindMe 6mo agoPeople are going to misinterpret this and overgeneralize the claim. This does not say that AI isn't reliable for things. It provides a method for quantifying the reliability for specific tasks. You wouldn't say that a human who doesn't know how to read isn't reliable in everything, just in reading. Counting is something that even humans need to learn how to do. Toddlers also don't understand quantity. If a 2 year old is able to count to even 10 it's through memorization and not understanding. It takes them like 2 more years of learning before they're able to comprehend things like numerical correspondence. But they do still know how to do other things that aren't counting before then.
- irishcoffee 6mo ago> Counting is something that even humans need to learn how to do. Toddlers also don't understand quantity. If they're able to count to even 10 it's through memorization and not understanding. I completely agree with you. LLMs are regurgitation machines with less intellect than a toddler, you nailed it. AI is here!
- coldtea 6mo ago>Counting is something that even humans need to learn how to do No human who can program, solve advanced math problems, or can talk about advanced problem domains at expert level, however, would fail to count to 5. This is not a mere "LLMs, like humans, also need to be taught this" but points to a fundamental mismatch about how humans and LLMs learn. (And even if they merely needed to be taught, why would their huge corpus fail to cover that "teaching", but cover way more advanced topics in math solving and other domains?)
- nkrisc 6mo agoYou’re conflating counting and language. Many animals can count. Counting is recognizing that the box with 3 apples is preferable to the one with 2 apples. Yes, 2 year olds might struggle with the externalization of numeric identities but if you have 1 M&M in one hand and 5 in the other and ask which they want, they’ll take the 5. LLMs have the language part down, but fundamentally can’t count.
- kenjackson 6mo agoWhenveer I see these papers and try them, they always work. This paper is two months old, which in LLM years is like 10 years of progress. It would be interesting to actively track how far long each progressive model gets...
- coldtea 6mo agoEven more interesting to track how many of those are just ad-hoc patched.
- raincole 6mo agoProbably zero. At the end of the day people pay for LLMs that write better code or summarize PDFs of hundreds of pages faster, not the ones that can count the letter r's better. When LLMs can't count r's: see? LLMs can't think. Hoax! When LLMs count r's: see? They patched and benchmark-maxxed. Hoax! You just can't reason with the anti-LLM group.
- toraway 6mo agoWhenever an "LLM fail" goes viral like the car wash question, you can observe the exact same wording of the question get "fixed" within a week or so. With slight variations in phrasing still able to replicate the problem. Followed by lots of "works perfectly for me, why are people even talking about this?" I can't say what exactly they're doing behind the scenes but it's a consistent pattern among the big SOTA model providers. With obvious incentive to "fix" the problem so users will then organically "debunk" the meme as they try it themselves and share their experiences.
- simianwords 6mo agoYou are misremembering. There’s no patch. All these examples used the instant model.
- coldtea 6mo agoThe same non-argument could be said for all kinds of cheating on benchmarks by tech companies and yet we have tons of documented example of them caught with pants down. >You just can't reason with the anti-LLM group. On the contrary, the reasoning is simple and consistent: LLMs can't count r's shows that LLM don't actually think the way we understand thought (since nobody with the kind of high skills they have in other areas would fail that). And because of that, there are (likely) patches for commonly reported cases, since it's a race to IPO and benchmark-maxxing is very much conceivable.
- bigstrat2003 6mo agoLet us be very clear: there is no such thing as a trustworthy LLM. Time and again they have shown that they understand nothing. They can be useful in the right context, but you can't trust them at all.
- deleted 6mo ago[deleted]
- pants2 6mo agoDoesn't this just look like another case of "count the r's in strawberry" ie not understanding how tokenization works? This is well known and not that interesting to me - ask the model to use python to solve any of these questions and it will get it right every time.
- wahnfrieden 6mo agoIt's not dismissible as a misunderstanding of tokens. LLMs also embed knowledge of spelling - that's how they fixed the strawberry issue. It's a valid criticism and evaluation.
- cr125rider 6mo agoSeems like it’s maybe also a tool steering problem. These models should be reaching for tools to help solve factual problems. LLM should stick to prose.
- emp17344 6mo agoI think this is still useful research that calls into question how “smart” these models are. If the model needs a separate tool to solve a problem, has the model really solved the problem, or just outsourced it to a harness that it’s been trained - via reinforcement learning - to call upon?
- azakai 6mo agoIt has "outsourced" it to another component, sure, but does that matter? What the user sees is the total behavior of the entire system, not whether the system has internal divisions and separations.
- emp17344 6mo agoIt matters if you’re curious about whether AGI is possible. Have we really built “thinking machines”, or are these systems just elaborate harnesses that leverage the non-deterministic nature of LLMs?
- grey-area 6mo agoTo those saying this is not surprising, yes it will be surprising to the general public who are being served ads from huge companies like MS or OpenAI saying LLMs can help with their accounting, help them close deals by crunching the numbers in seconds, write complex code for them etc etc. This is important information for anyone to understand who thinks these systems are thinking, reasoning, and learning from them or that they’re having a conversation with them i.e. 90% of users of LLMs.
- orbital-decay 6mo agoQuick sanity check: you're susceptible to pretty irresistible optical illusions which would never fool a VLM, does it mean you're not thinking? In fact, with a non-monospaced font I also have trouble determining whether these parens are balanced, and have to select them with the mouse, i.e. use a "dumb" tool, to make sure. Reminder that "thinking" is an ill-defined term like others, and the question whether they "think" is basically irrelevant. No intelligent system, human or machine, will ever have zero error rate, due to the very nature of intelligence (another vague term). You have to deal with that the same way you deal with it in humans - either treat bugs as bugs and build systems resilient to bugs, or accept the baseline error rate if it's low enough.
- flextheruler 6mo agoWho is hiring anyone to look at a screen to count characters? Don't be disingenuous in your argument. The apt comparison would be the current technique used to accomplish this task i.e. a pattern matching algorithm.
- stratos123 6mo ago> saying LLMs can help with their accounting, help them close deals by crunching the numbers in seconds, write complex code for them etc etc. Why do you think the results of this paper contradict these claims at all?
- grey-area 6mo agoA machine which confabulates and cannot count is not a good fit for accounting tasks. They’ll make all sorts of subtle errors which are difficult for humans to notice.
- jeremie_strand 6mo ago[dead]
- simianwords 6mo agoThere’s no way this is right. I checked complicated ones with the latest thinking model. Can someone come up with a counter example? Edit: here’s what I tried https://chatgpt.com/share/69cebb52-56a8-838f-969c-c47308262afa https://chatgpt.com/share/69cebb52-56a8-838f-969c-c47308262a...
- pton_xd 6mo ago"in this paper we primarily evaluate the LLM itself without external tool calls." Maybe this is a factor?
- simianwords 6mo agoNo tools were used.
- chromacity 6mo agoIIRC, web chat often uses tools / code without surfacing this information in any obvious way.
- stratos123 6mo agoDid you use the exact API call shown in the paper? I am unable to replicate the paper's counterexamples via the chat UI, but that's not very surprising (if the LLM already only fails a few cases out of thousands, the small differences in context between API and chat might fix them).
- simianwords 6mo agoI tried this https://chatgpt.com/share/69cebb52-56a8-838f-969c-c47308262afa https://chatgpt.com/share/69cebb52-56a8-838f-969c-c47308262a...
- emp17344 6mo ago[flagged]
- deleted 6mo ago[deleted]
- Topfi 6mo agoNo disrespect to them, but unless there is a financial incentive at stake for them (beyond SnP500 exposure), I've gotten to viewing this through the lens of sports teams, gaming consoles and religions. You pick your side, early and guided by hype and there is no way that choice can have been wrong (just like the Wii U, Dreamcast, etc. was the best). Their viewpoint on this technology has become part of the identity for some unfortunately and any position that isn't either "AGI imminent" or "This is useless" can cause some major emotions. Thing is, this finding being the case (along with all other LLM limits) does not mean that these models aren't impactful and shouldn't be scrutinised, nor does it mean they are useless. The truth is likely just a bit more nuanced than a narrow extreme. Also, mental health impact, job losses for white collar workers, privacy issues, concerns of rights holders on training data collection, all the current day impacts of LLMs are easily brushed aside by someone believing that LLMs are near the "everyone dies" stage, which just so happens to be helpful if one were to run a lab. Same if you believe these are useless and will never get better, any discussion about real-life impacts is seen as trying to slowly get them to accept LLMs as a reality, when to them, they never were and never will be.
- entropicdrifter 6mo agoI have a friend who is a Microsoft stan who feels this way about LLMs too. He's convinced he'll become the most powerful, creative and productive genius of all time if he just manages to master the LLM workflow just right. He's retired so I guess there's no harm in letting him try
- ziml77 6mo agoI suspect they're afraid that if the hype dies, so will the pace of progress on LLMs as well as their cheap/free usage of them.
- itsmyro 6mo agobruh
- hu3 6mo ago> we found that GPT-5.2 cannot even compute the parity of a short string like 11000, and GPT-5.2 cannot determine whether the parentheses in ((((()))))) are balanced. I think there is a valid insight here which many already know: LLMs are much more reliable at creating scripts and automation to do certain tasks than doing these tasks themselves. For example if I provide an LLM my database schema and tell it to scan for redundant indexes and point out wrong naming conventions, it might do a passable but incomplete job. But if I tell the LLM to code a python or nodejs script to do the same, I get significantly better results. And it's often faster too to generate and run the script than to let LLMs process large SQL files.
- plagiarist 6mo agoThe dream is probably that the inference software then writes and executes that script without using text generation alone. Analog to how a human might cross off pairs of parentheses to check that example.
- ubutler 6mo agoChatGPT already does this, albeit in limited circumstances, through the use of its sandbox environment. Asking GPT in thinking mode to, for example, count the number of “l”s in a long text may see it run a Python script to do so. There’s a massive issue with extrapolating to more complex tasks, however, where either you run the risk of prompt injection via granting your agent access to the internet or, more commonly, an exponential degradation in coherence over long contexts.
- whateveracct 6mo agoThat's because abstraction is compression of information.
- dwa3592 6mo agoNice! Although I tried the parenthesis balanced question with gemini and it gave the right answer in first attempt.
- dwa3592 6mo agobut it's a tricky question for LLMs; it shows that if it's not in the training set; LLMs could trip which kinda shows that the intelligence is not generalized yet. I tried this with gemini - (i am trying(something(re(a(l(ly)c)r)a)z)((y)he)re) and it tripped.
- orbital-decay 6mo agoIntuitively this looks like an architectural artifact (like optical illusions in humans) or a natural property of learning rather than a lack of generalization. I have issues with your example too and have to count slowly to make sure.
- dwa3592 6mo agoRight, I am sure you were able to solve it albeit slowly- you knew you had to do it slow. LLMs which are mathematicians don't know that and can't seem to understand that they need to do it slowly.
- orbital-decay 6mo agoThey do if they are trained to use a reasoning chain or another form of loopback, and you don't overwhelm it, or if they are optimized to search for the solution forever. There's nothing fundamental about it, only the fact that the raw transformer expressivity is limited by the single pass through the layers, which is circumvented by the loopback. And I'm still pretty likely to make the off-by-one error even if I slow down, and there are certain optical illusions are nearly guaranteed to confuse me no matter how hard I try, particularly if I don't use any visual guides (i.e. tools). VLMs will not make my mistakes but will make their own ones, because their quirks are different from the quirks of my visual cortex.
- simianwords 6mo agoCan someone produce a single example <20 characters that fails with latest thinking model? Can’t seem to reproduce.
- cineticdaffodil 6mo agoAnother strange thing is that they just dont know the endings of popular stories. Like olanets that get blown up, etc. they just dont have that material..
- simianwords 6mo agoThis paper is complete nonsense. The specific prompt they used doesn’t specify reasoning effort. Which defaults to none. { "model": "gpt-5.2-2025-12-11", "instructions": "Is the parentheses string balanced? Answer with only Yes or No.", "input": "((((())))))", "temperature": 0 } > Lower reasoning effort The reasoning.effort parameter controls how many reasoning tokens the model generates before producing a response. Earlier reasoning models like o3 supported only low, medium, and high: low favored speed and fewer tokens, while high favored more thorough reasoning. Starting with GPT-5.2, the lowest setting is none to provide lower-latency interactions. This is the default setting in GPT-5.2 and newer models. If you need more thinking, slowly increase to medium and experiment with results. With reasoning effort set to none, prompting is important. To improve the model’s reasoning quality, even with the default settings, encourage it to “think” or outline its steps before answering. ———————- So in the paper, the model very likely used no reasoning tokens. (Only uses it if you ask for it specifically in prompt). What is the point of such a paper? We already know that reasoning tokens are necessary. Edit: I actually ran the prompt and this was the response { "model": "gpt-5.2-2025-12-11", "output_text": "Yes", "reasoning": { "effort": "none", "summary": null }, "usage": { "input_tokens": 26, "output_tokens": 5, "total_tokens": 31, "output_tokens_details": { "reasoning_tokens": 0 } } } So reasoning_tokens used were zero. So this whole paper is kinda useless and misleading. Did this get peer reviewed or something?
- Chobilet 6mo agoI'm sure this comment was made in good faith, but most researchers would rightfully understand these intricacies, and this is likely intentional(as noted in the paper). At a quick glance, I cannot say whether or not the paper has been peer reviewed(though unlikely/in process given how recent it was published). In general, you'd find published papers also listed in a specific journal/conference(i.e. not just the archives which anyone can submit to). Additionally, many of us in the field of researching LLM's are curious to understanding the boundaries and limitations of what is capable. This paper isn't really meant as any sort of "gotcha", rather serve as a possible basis point for future work. Though with a caveat I'm still digesting the paper myself.
- Bmello11 6mo ago[flagged]
- cadamsdotcom 6mo agoIsn’t this just a benchmark? “Model can count to 5”… tick. “Model can count to 10”… sorry you gotta wait til 2028.