5 ms·
The type of reasoning by the OP and the linked paper obviously does not work. The observable reality is that LLMs can do mathematical reasoning. A cursory inter
by jules 1y ago
The type of reasoning by the OP and the linked paper obviously does not work. The observable reality is that LLMs can do mathematical reasoning. A cursory interaction with state of the art LLMs makes this evident, as does their IMO gold medal scored like humans are. You cannot counter observable reality with generic theoretical considerations about Markov chains or pretraining scaling laws or floating point precision. The irony is that LLMs can explain why that type of reasoning is faulty:
> Any discrete-time computation (including backtracking search) becomes Markov if you define the state as the full machine configuration. Thus “Markov ⇒ no reasoning/backtracking” is a non sequitur. Moreover, LLMs can simulate backtracking in their reasoning chains. -- GPT-5
- deleted 1y ago[deleted]
- godelski 1y ago> The observable reality is that LLMs can do mathematical reasoning I still can't get these machines to reliably perform basic subtraction[0]. The result is stochastic, so I can get the right answer, but have yet to reproduce one where the actual logic is correct[1,2]. Both [1,2] perform the same mistake and in [2] you see it just say "fuck it, skip to the answer" > You cannot counter observable reality I'd call [0,1,2] "observable". These types of errors are quite common, so maybe I'm not the one with lying eyes. [0] https://chatgpt.com/share/68b95bf5-562c-8013-8535-b61a80bada97 https://chatgpt.com/share/68b95bf5-562c-8013-8535-b61a80bada... [1] https://chatgpt.com/share/68b95c95-808c-8013-b4ae-87a3a5a42bb1 https://chatgpt.com/share/68b95c95-808c-8013-b4ae-87a3a5a42b... [2] https://chatgpt.com/share/68b95cae-0414-8013-aaf0-11acd0edeb5c https://chatgpt.com/share/68b95cae-0414-8013-aaf0-11acd0edeb...
- FergusArgyll 1y agoWhy don't you use a state of the art model? Are you scared it will get it right? Or are you just not aware of reasoning models in which case you should get to know the field
- godelski 1y agoCareful there, without a /s people might think you're being serious.
- FergusArgyll 1y agoI am being serious, why don't you use a SOTA model?
- godelski 1y agoSorry, I've just been hearing this response for years now... GPT-5 not SOTA enough for you all now? I remember when people told me to just use 3.5 - Gemini 2.5 Pro[0], the top model on LLM Arena. This SOTA enough for you? It even hallucinated Python code! - Claude Opus 4.1, sharing that chat shares my name, so here's a screenshot[1]. I'll leave that one for you to check. - Grok4 getting the right answer but using bad logic[2] - Kimi K2[3] - Mistral[4] I'm sorry, but you can fuck off with your goal post moving. They all do it. Check yourself. > I am being serious Don't lie to yourself, you never were People like you have been using that copy-paste piss-poor logic since the GPT-3 days. The same exact error existed since those days on all those models just as it does today. You all were highly disingenuous then, and still are now. I know this comment isn't going to change your mind because you never cared about the evidence. You could have checked yourself! So you and your paperclip cult can just fuck off [0] https://g.co/gemini/share/259b33fb64cc https://g.co/gemini/share/259b33fb64cc [1] https://0x0.st/KXWf.png https://0x0.st/KXWf.png [2] https://grok.com/s/c2hhcmQtNA%3D%3D_e15bb008-d252-4b4d-8233-5679c0a1789a https://grok.com/s/c2hhcmQtNA%3D%3D_e15bb008-d252-4b4d-8233-... [3] http://0x0.st/KXWv.png http://0x0.st/KXWv.png [4] https://chat.mistral.ai/chat/8e94be15-61f4-4f74-be26-3a4289d6956b https://chat.mistral.ai/chat/8e94be15-61f4-4f74-be26-3a4289d...
- FergusArgyll 1y agoThat's very weird, before I wrote my comment I asked gpt5-thinking (yes, once) and it nailed it. I just assumed the rest would get it as well, gemini-2.5 is shocking (the code!) I hereby give you leave to be a curmudgeon for another year...