15 ms·
Disagreement among frontier LLMs on real-world fact-checks
- sperandeo 5mo ago[flagged]
- kostaj 5mo agoAuthor here. 67% (95% CI 64–70%) of 1,000 recent real user claims to a fact-checking platform had at least one of GPT-5.4, Claude Opus 4.7, Gemini 3 Pro, Gemini 3 Pro+Search, and Sonar Pro dissent from the panel majority — or no majority formed at all. Panel-level Krippendorff's α (ordinal) = 0.639, i.e. nontrivial but limited agreement. Quick context on what's in the writeup and what isn't: - What's measured: parsed-label agreement between the 5 models. Forced 4-choice (True / Mostly True / Misleading / False), no Abstain. No LLM grader, no reference verdict — every number is direct label equality. - What's not measured: which model is right. There's no ground truth in this paper. The 67% figure is a floor on rubric inconsistency (at least one model is label-inconsistent under the 4-bucket rubric on 67% of claims), not "model X is factually wrong on claim Y." - Why not AVeriTeC / PolitiFact / SimpleQA: those have been public for years and almost certainly appear in current frontier training data, so measured disagreement on them confounds inference with memorization. This corpus is structurally fresh — recent user submissions, 180-day window, near-duplicates collapsed, never paired with canonical verdicts in any public training set. - Our own platform's verdict is deliberately NOT used in this analysis. The paper measures frontier-panel disagreement only, not Lenz-vs-frontier. - Follow-up in progress: human-labeling every claim in this corpus so we can evaluate both the panel and our own platform verdict against a human reference. Critiques I'd most like to hear: (a) the iid CI assumption (Lenz claims cluster around topics and news events, so Wilson is probably optimistic), (b) ordinal-α vs alternatives for a 4-class ordered scale, (c) forced-choice vs allowing Abstain. Permanent archive: https://doi.org/10.5281/zenodo.20344847 https://doi.org/10.5281/zenodo.20344847
- airstrike 5mo agoNice work. Sonar who?
- jiggawatts 5mo agoMany of the rows in that spreadsheet reference "current events", which models aren't expected to do much better at than a human making an educated guess! They all have cutoff dates either last year or early this year and know nothing about what happened in "April 2026". This is doubly problematic because you evaluated earlier models like Gemini Pro 3 instead of 3.1, GPT 5.4 instead of 5.5, etc... Given that it's only a thousand short questions, you should be able to re-run your test in about an hour with the latest models, so... why haven't you? Similarly, LLM output is non-deterministic, so if you could get more interesting stats of your data set by repeating each question 'n' times for each model.
- kostaj 5mo agoTwo of the models used have retrieval capabilities and have access to newer information through search. The other three are parametric.
- furyofantares 5mo agoYes, so in that case you set them up to disagree and then measured disagreement.
- simonw 5mo agoComparing models with search tools to models without - when there's no option for "I am unable to answer this question without access to search" - doesn't make sense to me.
- kostaj 5mo agoAgree about comparing models with and without search capabilities. Even the two models with search capabilities (Sonar Pro and Gemini) agree only on 58% of the claims.
- throw310822 5mo agoThe title mention "fact-checks", but "fact checking" is a process in which facts are checked against sources, not one where you are given a random fact and have to tell if it's true or false from your own memory. That's what is normally called a quiz game. So a more honest title for this research would be "Models answer differently to quiz questions".
- LeifCarrotson 5mo agoI don't think that current LLMs really need an abstain option, they'll give an answer regardless of whether they're confident or not. I hope that future LLMs will, and will know when to use it. I understand why you prompted them to output exactly one label, but I'd bet if you'd asked a parametric or parametric "thinking" model to answer eg "On May 18, 2026, Ukraine carried out a drone attack on Moscow, Russia." [1] many would say something to the effect of "May 18 is after my knowledge cutoff, so I don't know. But based on the state of the war, the distance from Moscow to Ukraine, and drone range the best option might be...[TRUE]" [1]: https://lenz.io/c/130f1005 https://lenz.io/c/130f1005
- kriro 5mo agoI don't see it mentioned explicitly in the methods section but I assume you prompted each model only once for each question? Did you consider prompting n-times in blank states to see if the models even agree with themselves? Would also be interesting to add a virtual model that is simply the majority of all models and see how much the individual models differ from the "consensus". Do you plan to add some sources in the related work section of baseline numbers for human expert disagreement in fact checking tasks (I'm assuming such studies exist).
- kostaj 5mo agoIndeed. I prompted each model ones, plus one retry on errors. Very good point to measure the inter-model disagreement! Will add in the next version. Section "4.2 Agreement w/ peer majority" shows the level of agreement of each model with the majority. Yes, planning of human-labelling the same corpus of 1,000 claims and publishing a second study measuring the models performance against the human-labels on corpus that the models have not seen during training.
- johnbarron 5mo agoThanks for posting here. Keep expanding and improving your study. Correct where it deserves correction. The fact that HN decided to downvote the author of the study, shows how these people cant stay classy, and the mods stay silent...just shows what this is all about.
- christophilus 5mo agoThey get more human by the day.
- kilroy123 5mo agoThis made me chuckle. This brings up a very valid point, though. So many _humans_ can't agree on what the facts are these days. It seems to be getting worse. Not sure of the solution.
- embedding-shape 5mo ago> So many _humans_ can't agree on what the facts are these days. Ask ten people what "knowledge" is, and they'll come up with ten different answers. Go back 10, 50 or 100 years and humanity struggled with exactly the same issue for so long time. There is even an entire field of study literally just for trying to figure out what "knowledge" is: https://en.wikipedia.org/wiki/Epistemology https://en.wikipedia.org/wiki/Epistemology
- antonvs 5mo agoAnd on top of that, that entire field has not reached a consensus answer, and answers that have been proposed have been shown to be flawed.
- kostaj 5mo ago[dead]
- ipunchghosts 5mo agoI think ppl only care about how Claude or codex does.
- airstrike 5mo agoI agree but the market is pricing way beyond that
- spprashant 5mo agoLooks like they land at the average number of 67% disagreement.
- kostaj 5mo agoGPT-5.4 and Opus 4.7, specifically, agree between themselves on 65% of the claims - 95% CI 62–68%. I.e., in at least 35% of the claims, one of the two models is wrong under this 4-bucket rubric.
- TaupeRanger 5mo agobut that's without internet search - everyone I know uses the models that search when they need to, and I'm sure GPT and Opus would agree on almost everything if 1) they searched when necessary, and 2) they were allowed to give context to their answers instead of being hamstrung to get specious "research" results.
- spacebacon 5mo agoAnd they could all see exactly why if they chose to. https://huggingface.co/spaces/RiverRider/srt-introspect https://huggingface.co/spaces/RiverRider/srt-introspect
- embedding-shape 5mo ago> These aren't benchmark items with public answer keys — they're claims real users submitted for verification to a fact-checking platform. Cool. I wonder if anything of this matters when the authors don't disclose exactly how much of their report was written and made with LLMs in the first place? There even is a "11. Ethics & data use" section, and the research is about LLMs being infallible in some ways, yet the usage of LLMs for the production of this report isn't even mentioned once.
- kostaj 5mo agoData collection and processing was done manually. LLMs helped with the report drafting. Everything was human reviewed before publishing.
- embedding-shape 5mo agoSo it's not a secret, why you don't add this upfront to the report? The report itself is even about LLMs, makes a lot of sense to disclose your usage of them for writing the report, especially when you're presenting evidence that boils down to LLMs being infallible.
- kostaj 5mo agoIt's an omission on my side. Will add in the next version.
- embedding-shape 5mo agoI think you might be able to edit the website to add this, even if you aren't willing to make the report a bit more honest up front. I'm sure you realize that this website/article will now be sent around to a lot of people, many who don't realize exactly how this was written, because they don't read HN comments, they only skim the page contents, and I think most would (incorrectly) assume a report about infallible LLMs to not be written by LLMs, especially when the authors are the same ones who made the report itself.
- ars2424 5mo ago[flagged]
- simonw 5mo agoHere's the prompt they used: Classify this claim as of <date>: "<atomic claim>" Output exactly one label: True, Mostly True, Misleading, or False. No explanations, no qualifiers. The claims look like this: https://lenz.io/research/llm-disagreement/data.csv https://lenz.io/research/llm-disagreement/data.csv I put that in Datasette Lite to make it easier to explore. Here's an example of a disagreement: https://lite.datasette.io/?csv=https%3A%2F%2Fstatic.simonwillison.net%2Fstatic%2Fcors-allow%2F2026%2Flenz-llm-disagreement.csv#/data/lenz-llm-disagreement/2 https://lite.datasette.io/?csv=https%3A%2F%2Fstatic.simonwil... The claim was "All almonds are grown in the U.S. state of California.". All but one model said False, Opus 4.7 said "misleading". I feel like having "mostly true" and "misleading in there weakens the story, especially given the "no explanations" rule in the prompt. The almond thing is false, but I'd argue that "misleading" might be defensible if you were to accompany it with "the majority of almonds are grown in California, but not all of them". [ Update: OK, this almond thing was a bad example and I regret picking it. Read on for better ones. ] The prompt lacks any kind of rubric to clarify how those terms should be applied. As is so often the case with this kind of study, it's an evaluation of the prompt and harness used by the study in addition to being an evaluation of the underlying models. Update: here's a better example: "Incomplete Egypt visa application forms are among the most common reasons Egyptian visa applications are rejected." The models were split between "true" and "mostly true". Given the "among the most" language either of those answers means effectively the same thing. Update 2: a much better example: "On May 18, 2026, Ukraine carried out a drone attack on Moscow, Russia" The only correct answer to that, if you don't have a search tool, is "this claim is impossible for me to verify". And that wasn't an option. The answers were split between true and false: https://lite.datasette.io/?csv=https%3A%2F%2Fstatic.simonwillison.net%2Fstatic%2Fcors-allow%2F2026%2Flenz-llm-disagreement.csv#/data/lenz-llm-disagreement/76 https://lite.datasette.io/?csv=https%3A%2F%2Fstatic.simonwil...
- harpastum 5mo agoWithout providing definitions of "True / Mostly True / Misleading / False" to each rater, I rate the article's claim that "Only one verdict bucket can be correct per claim" as false. Something can be simultaneously "misleading" and either true or false. Which category should something go in if it's "mostly false"? How much can something be wrong before it goes from "mostly true" to "false" (objectively, both have some part of the fact that is not true)? This is at least partly testing the model's definition of "mostly" and "misleading". Not its understanding of the fact. Claiming that this means the models have fundamental disagreement on the facts themselves is an overreach.
- apples_oranges 5mo agoThat's better than all agreeing on the wrong answer, however.
- kostaj 5mo agoBtw, sometimes that do that too -- all agree on the wrong answer.
- pessimizer 5mo agoI've had multiple models give the same wrong answer or even fabricate the same nonexistent reference based on a similar prompt. My most common chatbot prompt is "X that you mentioned above doesn't seem to actually exist."
- f_devd 5mo agoInject some adversarial priming as is in actual usage, and you can probably get that number to >=95%
- kostaj 5mo agoOur experience with Lenz is that forcing a multi-step process, incl. adversarial debates, helps improve the verdicts.
- andai 5mo agoThis is an odd one. The paper is real, but was written by Claude? I am assuming OP is human, but also appears to be using Claude to post.
- proofofcontempt 5mo agoLet's be real, we all asked Claude to summarise this because it was written by Claude
- bobosmrad 5mo agolooking at the claims i would say 5 humans would disagree even more than the llms some of the claims where llms disagree: "On May 18, 2026, Ukraine carried out a drone attack on Moscow, Russia." "The slogan "Simon Go Back" was chanted in opposition to the Simon Commission in British India (1928–1930)." "Neptune Deep will start delivering natural gas in 2027." "A hotel villa in Kyrgyzstan displayed a sign stating 'no Jews, no dogs'." "Donald Trump said that an attack on Iran was postponed at the request of Gulf allies."
- pjc50 5mo ago> "Neptune Deep will start delivering natural gas in 2027." This is a "forward-looking statement", and presents special problems because you cannot really evaluate it until that date. You can only assign "likely or unlikely".
- simonw 5mo agoIf you are an LLM with a knowledge cutoff in the past and no access to a search tool the only correct answer to "On May 18, 2026, Ukraine carried out a drone attack on Moscow, Russia" is "this claim is impossible for me to verify". And that wasn't an option.
- ecshafer 5mo agoThese "Facts" are interesting. "Neptune Deep will start delivering natural gas in 2027." for example is not a fact, its a prediction. "On May 18, 2026, Ukraine carried out a drone attack on Moscow, Russia." is less of a fact and more of a litmus test for which sources of information you trust.
- kostaj 5mo agoIndeed. Real-world claims are somewhat messy. Some of the standard benchmarks, e.g. the questions in AVeriTeC, share similar characteristics.
- deleted 5mo ago[deleted]
- proofofcontempt 5mo agoWhat does this show that we didn't know already? LLMs cannot provide accurate answers to questions where data is not included in their training sets. This doesn't appear to have much substance
- 101008 5mo agoUnfortunately most people are not aware of this and treat LLM models as this superpowered brain who knows everything and can do everything.
- zug_zug 5mo agoWell then it shows that these models are using widely disparate training sets and have high confidence even when they shouldn't. Questions like "is mouthwash effective" presumably has one solid data source -- medical journals.
- simonw 5mo agoBut the prompt didn't give the models the option to say "I don't know", so it wasn't a measure of their confidence.
- zug_zug 4mo agoI mean that's true but I don't think that's realistically what's going on when one model gives an unqualified "Yes" and the other gives an unqualified "no." You can argue the study isn't as case-closed-decisive as we'd ideally like, but it's certainly evidence. It's probably hard to design a better study.
- TaupeRanger 5mo agoWhat are you talking about? The models were not ALLOWED to have confidence (or the lack thereof). They were explicitly told to give a single label, and in most cases, all of them were correct depending on additional context they would surely have provided, especially with access to the internet (which some didn't have). This is just silly.
- 5mo ago
- throw310822 5mo agoNot sure I'm understanding this. The models are asked to evaluate the truth of random claims out of their own head (except for Gemini with search grounding)? Isn't it exactly the same as asking people to play any quiz game and then rating them as "they disagree n% of the time"? The output buckets are also pretty questionable- the difference between "True" and "Mostly true" is pretty fuzzy. Is this marked as a "disagreement"?
- kostaj 5mo agoAgree that True and Mostly True might be very close and could be a calibration difference. Misleading and False, as well. A better headline number might be the 34% claims with substantial or polar-opposite verdicts.
- bayarearefugee 5mo ago(Brought to you by) Lenz...? a crummy commercial...? ...son of a bitch
- kostaj 5mo ago:) No Lenz data is included in the research on purpose. All information to replicate the results, including the claims data, is published.
- Razengan 5mo agoRecently, in May 2026, I asked ChatGPT 5.5 High to search for flights to a certain city that has recently had a new airport since like December 2025 It said the airport code didn't exist I mean, I get the "knowledge cut off date" and whatnot, but for that sort of thing, you'd think they'd check live information before gaslighting the user, specially since it's a "live" task anyway.
- rastrojero2000 5mo agoGiven that models are fundamentally incapable of comprehending what truths or falsehoods are beyond their location in their self made representational space, it's actually pretty impressive that they managed to make it not a cointoss. That 17% right there is thousands of man-hours poured over making the word vomiting process slightly closer to whatever their little ports say is happening in reality.
- utopiah 5mo agoDon't forget people Goodhart's law will make this "benchmark" moot in weeks if not days. It will get integrated back into the fold, it will look "solved" but there will still be no reasoning, just more statistical technical correctness because light has be shown on a new "problem" to solve. It will then be clamored as great "progress" that will "change everything". PS: yes, I might or might not have a degree in corporate strategy & PR.
- anon291 5mo agoIs this not true of human intelligence as well? Many smart people I know hold beliefs that have no obvious truth value.
- aspenmartin 5mo agoThat is an effect but it’s not a nail in the coffin. There are lots of proprietary benchmarks on real product traffic that aren’t contaminated and open questions as well. People at these labs largely know what they are doing, it’s not like people don’t know this.
- thegrim33 5mo ago"None of these claims is older than February 15, 2026" All of the models they tested were trained on data from before February 15th ... being asked specific questions about things that happened after they were trained.
- draw_down 5mo ago[dead]
- kostaj 5mo agoTwo of the models used have retrieval capabilities and can access newer information via search. Valid point for the other 3 models. All of the claims were submitted after February 15, 2026, but many of them were not time-sensitive (e.g. did not cover events than happened recently).
- throwaway613746 5mo ago[dead]
- alvis 5mo agoThe problem is that it's testing claims (or some people would prefer calling them "truths") without much context. Take just one random example: `Hostels in Kota, Rajasthan commonly use caged ceiling fans as a preventive measure against student suicides` While `Hostels in Kota, Rajasthan commonly use caged ceiling fans` may be a verifiable facts (though I doubt if there are any statistics for verification but let's say there are), `a preventive measure against student suicides` is a claim that no one can prove that. It can just a believe at most. Arh. Did Biden stole Thump 2nd term? Truth or fact or claim?
- fergie 5mo agoPersonally I find that every llm I use is unable to consistently identify the latest npm version numbers of the node packages that I use.
- cm2187 5mo agoOnly had a brief look at the “facts” that were made to check, many are quite political, where two fact checking organisation of opposite political persuasion would probably disagree more often than 67%.
- 6stringmerc 5mo agoCould be an interesting angle for cross-referencing with US jury verdicts, not that the objective True/False issue is concrete, but in the reality that flawed reasoning is endemic to our species. Systems designed and built by humans inherently have flaws in their DNA which take generations to sort out, if ever.
- john_strinlai 5mo agobetween the bad methodology, bad selection of 'facts' (some are predictions, some are opinionated, etc.), and ai-written report without disclosure... i dont get why this so high up on the front page. this is, frankly, a worthless assessment. i classify the entire thing as "misleading"
- dncornholio 5mo agoI really wished these comments were the norm and not the exception.
- wongarsu 5mo agoOne fun example: "Ruskin Bond was born on May 19, 1934, in Kasauli, Himachal Pradesh, India". Opus and Gemini believe this to be true, GPT 5.4 believes it's false, Sonar thinks it's mostly true. Disagreement value of 3, you can't disagree more than some models thinking it's true, some thinking it's false But my impression from 2 minutes on Wikipedia is that the most likely disagreement is on the "Himachal Pradesh, India" part. The guy was born on that date, in that town. But while the town is today in the state of Himachal Pradesh in India, that was not true in 1934. When he was born, the city was in the Punjab States Agency of the British Raj. So was he born in Himachal Pradesh, India or not? I find both True and False equally defensible here https://lite.datasette.io/?csv=https%3A%2F%2Fstatic.simonwillison.net%2Fstatic%2Fcors-allow%2F2026%2Flenz-llm-disagreement.csv#/data/lenz-llm-disagreement/21 https://lite.datasette.io/?csv=https%3A%2F%2Fstatic.simonwil... https://en.wikipedia.org/wiki/Ruskin_Bond https://en.wikipedia.org/wiki/Ruskin_Bond
- anon291 5mo agoThere's lots of things like this where if you ask a human, the answer will change depending on what's convention in their subculture.
- flextheruler 5mo agoHow can someone be born in a state that does not yet exist? The statement has the year in it clearly demonstrating the contradiction. One can't be born in the Soviet Union in 1995 or in Tsarist Russia in 1950.
- jawns 5mo ago"Extraterrestrial life exists somewhere in the universe." GPT-5.4: Misleading Opus 4.7: Misleading Gemini 3: FALSE Gemini 3 (Retrieval): FALSE Sonar Pro: FALSE It's a weird fact claim, because the ground truth is "nobody knows for sure" and that's not one of the available options.
- wongarsu 5mo agoOf the available options, "Misleading" is probably the best, since something that is most likely true but unproven is presented as fact But "unknown or undecidable" should have been a category.
- Alifatisk 5mo agoIsn't misleading the correct option here then?
- arcfour 5mo agoFalse makes sense if you are interpreting it strictly as "has this been proven?"
- wongarsu 5mo agoFalse is correct, but misleading My implicit assumption is that if you fact-check the fact-check, any label other than "true" means the original fact-check is unacceptable
- throw310822 5mo agoNo, "misleading" is a statement that is used because it suggests something else. It's a curious category because, differently from true and false, it's not about the statement itself but rather the intention behind its usage or the way it might be understood. It's frankly more of a political judgement than a matter of facts.
- ertgbnm 5mo ago"Shark attacks correlate strongly with ice cream sales" is an entirely true statement that some would argue is also misleading. Misleading should be removed as a category and replaced with a better hedge like "not sure"
- jasonvorhe 5mo agoSimple: If it claims to be a fact check it's just propaganda.
- al_hag 5mo ago[flagged]
- wg0 5mo agoTake my job please.
- fumeux_fume 5mo agoI think we can all agree that this experiment being flawed in multiple ways is TRUE. But I think it's a great exercise in identifying common mistakes people make when using LLMs. This would be a great interview question for a prompt engineering job.
- imperio59 5mo agoOne of the claims it asks LLMs to grade is "Artificial intelligence will cause widespread job loss among software engineers." Yea man this benchmark is really really bad.
- scotty79 5mo agoSo basically saying that random fact-checking claim is exactly true or exactly false is hard. It's way easier to decide it's misleading or mostly true is way easier.
- elorant 5mo agoTell me about it. I spent a week back and forth between four models (ChatGPT, Claude, Gemini, Grok) trying to enhance a PPMI algorithm. They couldn’t agree on anything. One was refuting what the other said. Eventually I decided to follow what Claude suggested because its explanations made the more sense.
- kostaj 5mo agoIndeed. For algorithms and coding, my personal routine nowadays is to review every detailed plan with Opus 4.7 and GPT-5.5. They tend to find very different type of gaps.
- kaicianflone 5mo agoDissent and consensus among frontier models is a good thing. Just like on a team of high performers, there are a million ways to skin a grape. In my research, I've found that models perform better when they operate as a collective system with reputation, incentives, and accountability instead of isolated oracles answering alone. Agreement, dissent, and correctness should all carry rewards and consequences. Just like in real life. Collective machine intelligence, not AGI. It's expensive, but it's also naive to believe a single model will consistently produce profoundly correct answers to profoundly novel questions.
- haritha-j 5mo agoNot on objective truth though. That's how you get misinformation.
- deleted 5mo ago[deleted]
- mtrifonov 5mo agoFunny timing. I've been working on a prediction market orchestration that runs Claude and a few others over Polymarket/Kalshi. The models are NOT unanimous. At all, really. I spent about a month convinced that I could just run all five and take majority vote. Eventually I pivoted to a chaining approach where I benchmark areas each model excels, and settled on more like a graph-like architecture where outputs get split and verified by another, then reconstructed, and re-verified at each stage. Has actually been working out pretty well so far, 2 months in consistent profit, but I'm not a millionaire yet.
- aayushkumar121 5mo ago[flagged]
- kaicianflone 5mo ago[flagged]
- Ayush_Khati1 5mo ago[flagged]
- pessimizer 5mo agoPeople keep asking "where is the psychosis?" as a reply to people on the rapidly multiplying "CEOs have AI psychosis" threads that have been popping up here and cross-pollinating in the mainstream media for the last week or two. Here's the psychosis - these things are consistently randomly wrong depending on how the wind is blowing. People are telling you to leave them alone and let them build things, and they randomly forget that cities exist or that people died 100 years ago. Some people just don't see it as worth noting, and move on. That's crazy. These things consistently fabricate - as an inversion of this experiment, I've had different models come up with the same fabrication from similar prompts. People just call it "hallucination" and I think to them that saying that makes it cease to exist or be important - when "hallucinations" are going to be braided into every answer you get even if they're unidentifiable in the output. That's crazy. There are plenty of other crazy aspects, such as the idea that we suddenly need infinite pieces of bespoke software when all of the bespoke software I hear about people making is mundane. 3/4 of the time somebody mentions a project they're proud that they completed with LLMs to scratch some itch they had, somebody says "you haven't heard of X? It's been around forever" about something that they could have pulled down from their package manager. Who needs a spaghetti-coded, unsupported, untested version of X built on hallucinations that you haven't discovered yet (the LLM didn't realize that deleting files to reduce the archive size was unacceptable.) What is all of this software that people need but isn't there - where are all these unserved markets, where is all this future revenue supposed to come from? Why aren't LLMs suggesting new classes of software that would create new productivity and revenue sources? Could it be that millions of human ants over decades have mostly exhausted the space, and there isn't any easy hidden revenue? A common wisdom is that we had been vastly overhiring programmers during ZIRP, who in their idleness degraded user experiences and overcomplicated things, with management resorting to more and more sleazy and gamey means of margin extraction from more and more degraded services. We had an excess of labor, fueled by factors other than productivity, in fact being pissed away at companies that drove nose-first into the ground. What is throwing a trillion dollars of servers at that supposed to do? Is that not AI psychosis?
- cobblr_mosaic 5mo ago[flagged]
- GodelNumbering 5mo agoMore interesting part probably worth highlighting: The SAME model won't always return the same output when prompted with the same fact check. You ask a human 1000 times a fact check question, they say the same answer 1000 times. You ask an LLM the same question a 1000 times, your results could vary significantly. Humans work based on the Metamemory (knowing what they know), while LLMs are picking from statistical probability.
- logged4upvoting 5mo agoThat is not true, over an extended task that you cannot keep complete in memory humans do not behave with 100% consistency. I have labeled datasets with a human team and shown the same task to the same user on a different day, and they answered differently. Of course, they are usually consistent with themselves most of the time but not always.
- miellaby 5mo agoWhat's really weird to me is that "I don't know" is not a valid answer in this experiment while we can all agree that's the main issue with LLM right now is that they will happily "roleplay" an answer when they have nothing in their dataset corresponding to your query.
- culopatin 5mo agoI’m no expert but if LLMs are token prediction machines, and you tell it to not build an explanation before the answer, isn’t it less likely that the token prediction for the final answer will have less raw material before it to build a grounded response? In other words: no explanation > no foundation for prediction of the answer tokens?
- dataminer 5mo agoHoney does not spoil over time under normal storage conditions.,2026-02-17T04:11:51.495452+00:00,Science,True,True,True,True,Mostly True,1 If outcomes like these are collapsed on True-side then the disagreement will reduce from the headline number.
- briandw 5mo agoNo human baseline to compare it to. Without that you are missing an important check on the task being poorly constructed. More importantly there is an implied reference thats missing. The implication is that people would have done better, or that perfect agreement is possible.
- raincole 5mo agoAnd how many claims human experts disagree on in the exact same setting? I'm not being snarky here. Without something to compare to the 67% number tells us nothing. And it's known that many humans disagree with human fact checkers too (see: any election around the world.)
- kostaj 5mo agoAgree. Human experts also struggle agreeing on this type of claims. The inter-annotator agreement on the verdicts on the AVeriTeC corpus across 50 organizations is κ=0.619 - substantial but well short of perfect.
- mrkn1 5mo agoFor 100% local CPU fact checking, I made this: https://news.ycombinator.com/item?id=48301003 https://news.ycombinator.com/item?id=48301003
- gobdovan 5mo agoWhy should I trust this without a paper, benchmark or at least a human-written README?
- hiroto_lemon 5mo ago[dead]
- 0natcer 5mo agoFive frontier LLMs 100% agree that the title is misleading.
- mgrunwald_ 5mo agoAs an example, 2026 GPT doesn't even agree with its 2025 self. Last year I asked it to make a hardware comparison and it correctly identified the objectively better option. Recently I asked again and this time it got everything completely backwards.
- aspenmartin 5mo agoModels are stochastic. Did you look at pass@k? I wouldn’t be surprised if you saw a regression because these models are extremely complex and impact of various decision making downstream is complex.
- mgrunwald_ 5mo agoI ran this multiple times through GPT-4 and every single time it arrived at the same conclusion. The data was readily available and pretty clear. GPT-5 insisted that the objectively inferior option was better until I gave it my own benchmark data and it was like "Oh okay nevermind". Gemini's answer was very opinionated and factually correct, whereas Claude gave a more nuanced answer, which was also very good.
- aspenmartin 5mo agoThis sounds perfectly reasonable and consistent with our current understanding of these models
- pknerd 5mo agoIt's a prompting issue rather than an LLM issue. The guy needs a "Prompt 101" course.
- fooker 5mo agoI don't get why everyone is hellbent on getting LLMs to perform fact checking. This is not the technology for it. Sure it might sorta kinda work in some circumstances. That doesn't make it a good fit. Think of it like buying a refrigerator for storing clothes.
- nicce 5mo agoPeople ask questions to get answers. For me, it feels quite important? Especially when search engines start to push them?
- fooker 5mo agoJust because it is important for the use case does not mean we can make it work. It's a pretty well known fundamental limitation of the technology. No amount of elbow grease will get it there. There's an interesting tradeoff here, a year or two ago maybe it got facts right 50% of the time. Everyone knew not to rely on it. Now, suppose we are 90% of the way there, only technically proficient people would know not to trust it. (like not adding Internet Explorer toolbars! Or remembering to use ad blockers..) A few years later, suppose we have spend a lot of money and effort getting it 99% of the way there, trusting it would be somewhat natural by then. And then for the important 1% of the situations, it would stand to cause real harm. 1% seems low, but for a million invocations, you'd have 10000 mistakes.
- xboxnolifes 5mo agoYour progression is basically the exact same progression as things like Wikipedia, and web search in ggeneral. So, I guess we dont need to hypothesis. Just look around and see how its played out. How many people take the first result on Google as gospel when looking things up?
- fooker 5mo agoGoogle search and Wikipedia both started out being fairly reliable to their source of truth. Google pretty much guaranteed that their top results were relevant to the search query. And wikipedia had an army of people making sure everything was backed up by the references. Crucially, neither claimed to be an arbiter of truth.
- seanplusplus 5mo agoDude. If you give LLMs a vague rubric and force a choice, they'll make different arbitrary calls on the margins. Yeah. That's what happens when you give humans a vague rubric too.
- DonutATX 5mo agoWhy did they exclude Grok? Given the published philosophical differences in how Grok is trained, it would provide an interesting data point. You can argue all day about those differences, but missing this opportunity to observe them in an objective way is disappointing.
- testfrequency 5mo agoTitle says “Frontier” which would exclude Grok. Grok is trained to have a bias, which a lot of people like, but it’s not meant to be accurate.
- simianwords 5mo agoHow do you know it is trained to have a bias? In fact can I ask you to provide a single reproducable answer right now?
- testfrequency 5mo agoAssuming this isn’t a satire reply: https://www.pnas.org/doi/10.1073/pnas.2603294123 https://www.pnas.org/doi/10.1073/pnas.2603294123 Hope this helps!
- simianwords 5mo agoThis doesn’t show grok as a model has bias but only that the product that uses grok has bias. Even the referenced papers to show models can have bias don’t show anything about grok. Overall you have given me zero evidence that grok model itself has some political bias. FWIW I don’t mind bias but I haven’t seen evidence of it.
- Forgeties79 4mo agoNo serious person should use Grok for anything remotely important. Musk openly changes it when it gives responses that anger his followers
- kstenerud 5mo ago> No Abstain option is offered (a forced choice keeps the comparison symmetric across models). Well that's your problem right there: They removed any confidence indicator and forced a choice. For example: Statement: Individuals who prefer music with less positive emotional content tend to have higher intelligence. Gemini: That statement is supported by recent psychological research, though with some important scientific caveats regarding how strong that link actually is. How should the agent classify this? True? Mostly true? Misleading? False?
- comboy 5mo agoThis is wrong on so many levels, from data through process to evaluation. How do you even prompt claude not to give you Pearson for correlating them.
- dncornholio 5mo agoA post generated by AI with data generated by AI. Worthless.
- htx80nerd 5mo agoI like ChatGPT a lot but it is always trying to debate and disagree when you ask it simple non-controversial questions. Trying to turn everything into a debate session instead of just answering the question.
- gamander2 5mo ago[flagged]
- hack1312 5mo agoYour antisemitism is disgusting.
- antonvs 5mo agoIf your quote were true, then by the Randian logic of the people who make such claims, they must deserve to do so and you shouldn't have any issue with it. But your quote certainly isn't true if you're actually talking about "the world". For example, Japan and China are the two largest holders of US Treasury bonds. China controls roughly 50% of the the contracted construction market in Africa. These are just examples of the sort of thing you'd need to take into account in trying to justify your silly racist claim.
- tomhow 4mo agoWe've banned this account.
- jmull 5mo agoThe difference between "mostly true", "misleading", and "false" is context, and responses are specifically not allowed to include any context. Even "true" has a little context, since few things can be said to be absolutely true. "Unknown" also isn't allowed. What's 2 + 2? The answer must be one of the colors of the rainbow. (People can draw their own conclusions, but the only coherent reason I can think of for the design of this experiment is to generate a misleading conclusion.)
- bilsbie 5mo agoWhy do we want to build intelligence if it just confirms what we already think we know?
- bilsbie 5mo agoSounds like a lot of room for human bias. How would it have responded to these claims in the past: THALIDOMIDE is safe CIGARETTES are safe ASBESTOS is safe MERCURY is safe DDT is safe LEAD in gasoline is safe
- monkpit 5mo agoWhat’s the point of this if they didn’t use temperature=0 for every model (they didn’t)? They could have redone the test against the same model and gotten different answers. It’s almost like picking 2 different coins and comparing the list of coin flip results. (I realize it’s not that straightforward, it’s not 50/50, but it’s essentially the same issue.)
- scoofy 5mo agoI hate to get really pedantic here, but the concept of "truth claims" plays fast and loose with concept of knowledge in a philosophical sense. The idea of "fact checks" misunderstand how information and knowledge work together. Knowledge is about evidence, not "facts" because facts are a shorthand for a preponderance of evidence. I feel we are doomed to debate the veracity of Wikipedia on a loop, forever, because people don't understand that Wikipedia exists as a place to find citations not as a place to find facts. Yes, those stated facts may disagree with the citations, but even if we try to fix that issue by having experts write the encyclopedia, we still suffer from the problem that the experts are often wrong. We need a view of knowledge's relationship to LLMs that is based in Karl Popper's idea of falsifiablity. We should ask LLMs for evidence of claims not for truth values. Truth values are foundational to deductive systems, where axioms define truth. In inductive systems, like the real world, the concept of black swan events means that truth values are never fixed and are always in a state of uncertainty. I honestly think it would be helpful going forward if we add some basic philosophical education to the standard curriculum, because no that we have an artificial form of information retrieval, we need to be much, much more pedantic about how we interpret that information.
- lyfi2003 5mo agoIt's right, you must be professional than llm
- serial_dev 5mo agoIt’s just shows that fact-checking is not a thing for 99% of the cases. It’s interesting to see it in LLMs, but it’s not unique to them. The “fact checkers” pretend they are objective and authoritative, but they are not, they are just one more opinion. For the research, the four classification options are too many, it should be true, false, and maybe “can’t be determined”.
- pseudopolous 5mo ago[dead]
- nailer 5mo ago> the most recent real-world user submissions to a fact-checking platform 'Fact checking' platforms aren't truth. Many 'fact checking' platforms are self-admittedly focused on left advocacy (snopes), or right wing advocacy (newsbusters). lenz-llm-disagreement.csv doesn't state the data source.
- johnnienaked 5mo agoLLMs will be great politicians one day
- anonymousiam 5mo agoGIGO is an acronym I learned in the 1970s. Things haven't changed much since then. We live an an era where people have "their own truth", so why not let the AIs have theirs too? The AI companies have editorial privilege on the content they feed their LLMs, and on the prompts that the users never see. I don't know why they feel a need to interfere when their AI produces something that's politically incorrect. Perhaps it's because they have a fundamental credibility problem with their products...
- chipsrafferty 5mo agoIt's becoming increasingly clear to me that - at least right now - AI is only useful for 2 things: 1. Coding, with it being more useful the better you are at coding without AI 2. Any expert in their field asking questions about their field, who bother to fact check the output. E.g. "claude pls search these 1000 files and tell me if you find anywhere that they're discussing the settlement" and then the user checks the files/line numbers to make sure that it's correct - basically a turbocharged search that may have false negatives (content existed but I didn't find it) or false positives (content that I classified in a certain way but it was wrong). It takes an expert to tell the latter one in some cases.
- dktoao 5mo agoI haven't found it that useful for doing any actual "agentic" coding at $DAYJOB with lots of legacy code it wasn't trained on (because proprietary). I do find it useful for summarizing sections of code that I am working on and asking for snippets that do very specific things. Also, it is pretty good at writing one-off short scripts with easily definable inputs and outputs. I have come to the conclusion that people using AI for coding need to think about it as basically an automated version of the Docs -> Copy Paste -> Stack Overflow -> Copy Paste -> Compile Error -> Google -> Copy Paste -> New feature request from management -> Random internet blog -> Copy Paste loop that most of us do for a lot of the non-logic heavy portions (e.g. API interfacing) of our work with less randomness and more pattern matching or statistics or whatever guiding the process. Honestly, pretty useful, not knocking it. I do think there is a killer application for AI, which it is already useful for, and that the industry doesn't really promote. That is basically taking a massive amount of unstructured data on a topic and allowing people an easy way to learn from that data without having to read through all of it (which may not even be possible for a single person in their lifetime). This would be a huge boon to humanity alone given the scale of data we produce. I think fundamentally, LLMs cannot take the data that they are so good at summarizing and use it in a creative way, it looks kinda like they can, because they are so good at "borrowing" other people's creative work, but in real-world scenarios where change is constant and the external forces of today are not understood by a model that was trained 3 months ago, they fall on their faces again and again. I think AI companies know this but cannot admit that this ground breaking (I would argue) technology might only be transformative for one half of the observe->act workflow that would be necessary to replace humans as workers because. 1. It is possible the economics don't work out without replacing workers 2. If they admitted that the only value of their tech was in distilling value already present in other people's creative work, work that the LLMs cannot create on their own, a sane government might force them to pay for their inputs.
- michaelmrose 5mo agoTotally aside from disagreement between models unbiased by prior input any such experiment may fail to capture the outcomes experienced by real users whose prior text exchanges may substantially change the text recieved. For instance see the folks who think that they have "awakened" their instance of ChatGPT. Actual usage may diverge to a greater degree than models
- husky8 5mo agoWatch the disagreements in real time via refinement pipeline on the results page pingpongit.com
- 40four 5mo agoThis shouldn’t be surprising. Let’s start off with the obvious. What does “real-world fact-check claims” mean? So we’re using the same list of “fact check claims” on each model. The problem is (unless I’m missing it) the authors aren’t exposing the list of 1K questions they used in the experiment. That’s a huge problem. Are the authors assuming the 1K claims they used are “provably true”? If so, that’s a huge bias, and opens up a philosophical debate about what it a fact? Or what’s makes something true/ false? As Marc Andreessen puts it: a particular domain is either explicitly “provable” or not “provable”. Provable domains include math, physics, chemistry, biology, engineering, even code. That not be the whole list, but everything else is essentially “unprovable”. At least as far as a language model is concerned. They are questions that require a human value judgement. Politics are an obvious example. So back to the “1K fact check claims“. How many of these are political, or current events questions? How many are STEM questions that can be laid out in a formal proof? Models can be trained to answer either way on claims that require a value judgement, but that’s obviously not beneficial to anyone except who controls the model. If the expectation is that all these frontier models should answer the same way on value judgement questions, then that’s never going to happen. What the models ARE good at though is breaking down the nuances of a topic and arguing both sides. This is how these tools should be used, as a way to analyze the claim and let us humans in the end make our own value judgement. If you’re trusting the model to make the value judgement for you and just accept it as a fact, then you are entering a a very dangerous territory.
- secondary_op 5mo agoVery interesting tool, but it's biased and not neutral from the get go, because I explicitly formulated claim in neutral way, but it automatically rewrote it to be western/wikipedia POV and then immediately proceeded to verify it. original neutral: US DEPT OF DEFENSE/DNAVFAC planned renovations to School #05 in Sevastopol, Crimea in 2013 before Crimea became part of Russia in 2014 automatically rewritten to biased western view: The United States Department of Defense, via the Naval Facilities Engineering Command (NAVFAC), planned renovations to School No. 5 in Sevastopol, Crimea in 2013, before Russia annexed Crimea in 2014. https://lenz.io/c/73c0f16c https://lenz.io/c/73c0f16c And the follow up The phrasing "Crimea became part of Russia" is more neutral than the phrasing "Russia annexed Crimea." , and according to this tool is Misleading 9/10 https://lenz.io/c/93944614 https://lenz.io/c/93944614 Yeah, so my personal conclusion that this tool is garbage, it checks western/US allied only LLM providers, that in turn search only for western/US allied sources/documents like BBC/NATO and result is what it is.
- kostaj 5mo ago[flagged]
- shevy-java 5mo agoMy first reaction was: how dumb is AI still. But ... real people would also reach that result. Some believe that vaccination can not induce protection (which objectively is incorrect).
- deleted 5mo ago[deleted]
- cdud3 5mo agoThe CVS file with the raw data is a source of joy. My favorite is this: claim: Artificial intelligence will cause widespread job loss among software engineers. All 5 LLM's agree that the claim is misleading & wrong.
- graphememes 5mo agoeven this is ai output
- andrewchambers 5mo agoI think we need a wiki and/or stack overflow equivalent for agents and humans to collaborate. Grokipedia seems like the main site that is kind of exploring the concept - though I hope better more powerful ones emerge.
- ajjenkins 5mo agoThe appendix has examples of the statements that led to the most disagreement: https://lenz.io/research/llm-disagreement https://lenz.io/research/llm-disagreement
- sagebird 4mo agoshould have submitted it to 5 independent fact checkers - would have deflated nonsense before it began by showing that you are going to see trivial and non trivial shifts among them. mostly true and misleading counting as separate buckets while also being somewhat orthogonal conceptually is also stupid. a better output format might be true|false|unknowable, confidence where confidence is 0..1 at least then you can compare agreement among models as a distance measurement and not a moronic bucket agreement the conclusion is actually obvious: llms are good enough for most of this work and it is definitely cheaper, so you should use llms for fact checking at least as a first pass
- geraldsterling 4mo agoThe result makes me more bullish on using model disagreement as a routing signal than as a verdict.
- willXare 4mo ago[flagged]
- masterleopold 4mo ago[flagged]