3 ms·
This AI generated post (100% on Pangram) is pretty out of date. >On SimpleQA, a benchmark of factual recall with no tools allowed, the current leader is Gemini
by COAGULOPATH 2mo ago
This AI generated post (100% on Pangram) is pretty out of date.
>On SimpleQA, a benchmark of factual recall with no tools allowed, the current leader is Gemini 2.5 Pro at 53%, so the best recall money can buy still misses half the questions.
SimpleQA hasn't been updated in a long time. Gemini 2.5 Pro is a sixteen-month-old model, not "the best recall money can buy".
>The part I find most promising is what this does to hallucination. When a fact lives in weights, a wrong fact is unfindable and unfixable.
This seems confused. LLM hallucinations don't come from the weights containing "wrong facts", they are artifacts that appear at runtime.
>When the fact lives outside the model, a wrong answer has an address. The model cites a document, so you can open the document. If the document is wrong, you edit the document
You can make any modern LLM explain its reasoning and find sources for its claims. None of this has anything to do with facts needing to exist in weights or in harnesses.
The internet is full of wrong information and I cannot magically edit it to make it all correct, so this doesn't help me.
>if a model is factually wrong a claim with a source is checkable and a claim from weights isn't.
Why? If a model's weights claim that Bart Simpson became President in 2020, why does this fact suddenly become uncheckable?
- margalabargala 2mo agoI agree with everything you say except this: > You can make any modern LLM explain its reasoning You can make any modern LLM create a plausible, self-consistent explanation that looks like reasoning, but it's not "the reasoning it used to arrive at that answer".
- dist-epoch 2mo agoThis was beautifully shown by asking a model to explain how it added two numbers together (something like 45+21), and it told a plausible story, when in fact they showed it was some rotation on a helix living in some internal manifold. Like asking a human "how did you catch that fast ball coming at you?"
- markasoftware 2mo agoIt could be that the rotation in the helix manifold whatever is a low level representation of the logical steps (carry the 2, add the next column,...) it's describing. The point stands that the explanation it generates doesn't necessarily in all cases reflect what it "actually did" but your counterexample doesn't hold.
- dist-epoch 2mo agoThere have been multiple studies on how LLMs add numbers. They use the "Clock" algorithm, the "Pizza" algorithm, a few other ones. > All networks we study implement the same simple neuron model in their first-layer MLPs: degree-1 sinusoidal fits in layer 1, with deeper layers combining into degree-2 sinusoidal interactions. https://neurips.cc/virtual/2025/loc/san-diego/133808 https://neurips.cc/virtual/2025/loc/san-diego/133808 https://arxiv.org/abs/2502.00873 https://arxiv.org/abs/2502.00873 When you ask them they don't mention these at all, they give you high-school math: https://chatgpt.com/share/6a82afdd-872c-83eb-aad7-622d27f2dcdb https://chatgpt.com/share/6a82afdd-872c-83eb-aad7-622d27f2dc... > did you use the "cos" or "sin" function at all during this addition computation? > No. There is no need for trigonometric functions like sin or cos. The computation only uses basic arithmetic and place-value reasoning. Of course, if someone were implementing arithmetic in a computer, it is theoretically possible to express addition using extremely complicated formulas involving sin and cos. But in the reasoning I described, no trigonometric functions were involved at all. I simply decomposed the numbers into hundreds and smaller parts and added them.
- kentonv 2mo agoTangent: This is often true of humans as well. We often make a decision based on a gut feeling, and then backfill a logical reason supporting our feeling, without even realizing we're doing it -- rationalization.
- angry_octet 2mo agoAnd we all know some people that rationalise poor choices and misbehavior, hide their mistakes, etc, to an unacceptable degree. Sometimes the individual knows they are rationalising but continues anyway, other times they seem incapable of seeing that. When you ask people who are rationalising poor behaviour about the scenario, but it is someone else doing it, they may arrive at a better answer. Can we use multiple LLMs to achieve self criticism and critical thinking?
- efilife 2mo agoisn't this happening already? There's the concept called "thinking" where the models talks with itself before giving you the final answer
- angry_octet 2mo agoIt does multiple rounds of feedback which is called 'thinking' but whether that is critical thinking is unknown.
- glenstein 2mo agoTangent on the tangent: I think that's true in a minority of cases and in a majority of AI cases. Though in principle I think it should be possible for an LLM to have access to and faithfully represent its own reasoning.
- margalabargala 2mo agoOn the contrary, I would argue that it's true in a totality of AI cases. To your point, I agree that nominally there should be a way to give conceptual names to paths of weights, and when answering a question, notice which weights were and were not applied and retrospect on that. That's not what reasoning traces as they currently exist are, though.
- malfist 2mo agoSeriously, anybody with a passing knowledge of LLMs knows thats not how they function. You can't encode logic in them because that's not how they work. It's a statistical model with useful emergent properties. It doesn't think, it doesn't reason, it isn't aware of facts or the rules of logic.
- andai 2mo agoA true camper doesn't need to check Pangram, Jimbo. He goes by pure animal instinct!
- andai 2mo ago>>You can make any modern LLM explain its reasoning and find sources for its claims. >The internet is full of wrong information and I cannot magically edit it to make it all correct, so this doesn't help me. My favorite RAG experience was asking Bart (or whatever they were calling Gemini back then) an answer to a question I knew. It gave me the opposite of the truth (as was common with LLMs at the time). But weirdly, it had cited sources for this "fact." I checked the sources. Two of them, both AI SEO slop. In this moment, andai was enlightened...
- Gander5739 2mo ago>>if a model is factually wrong a claim with a source is checkable and a claim from weights isn't. >Why? If a model's weights claim that Bart Simpson became President in 2020, why does this fact suddenly become uncheckable? Because in one case you have a source you can use to validate the fact, and in the other you don't. Though, as you explain earlier in your comment, the premise is misguided/hallucinated.
- claiir 2mo agoYea the "When the fact lives outside the model, a wrong answer has an address" sentence seems aggressively AI written. Saw that and my senses went off.
- RealWed6 2mo agoSenses of what? LOL. The whole Internet is AI generated by now and we all contribute to that on daily basis. get used to it or dull your senses ...
- RealWed6 2mo agoAh Ah Ah – how we did't like this. Downvoting me will surely help! Can't you see how much slope is already around? Don't we – you and me – contribute to that, especially at work? Isn't "dull your senses" standard answers of most expensive shrinks? So what did you disagree with? Or you simply didn't like the truth? Ok, I got it. No problem.
- nojs 2mo ago> This AI generated post (100% on Pangram) is pretty out of date. Quite ironic given the topic. It seems that the author’s model indeed contained too much knowledge about old Gemini releases, and did not do enough tool calling.