Y
HN Search
Hacker News Search
new
|
comments
|
top
|
jobs
kostaj
searching PlanetScale…
1.
▲
2.
▲
3.
▲
4.
▲
5.
▲
6.
▲
4 ms
·
31.
▲
by
kostaj
4mo ago
:) No Lenz data is included in the research on purpose. All information to replicate the results, including the claims data, is published.
32.
▲
by
kostaj
4mo ago
Indeed. Real-world claims are somewhat messy. Some of the standard benchmarks, e.g. the questions in AVeriTeC, share similar characteristics.
33.
▲
by
kostaj
4mo ago
Yes, they are much closer verdicts. True and Mostly True are also close. Used Krippendorff's α (ordinal) to not penalize much closer disagreements. 21% of the claims have models that are on the polar opposite sides - at least one True,
34.
▲
by
kostaj
4mo ago
sonar-pro for the retrieval capabilities
35.
▲
by
kostaj
4mo ago
Two of the models used have retrieval capabilities and have access to newer information through search. The other three are parametric.
36.
▲
by
kostaj
4mo ago
It's an omission on my side. Will add in the next version.
37.
▲
by
kostaj
4mo ago
Used "No explanations, no qualifiers." to force the models to answer only with one of the four labels. It's worth running a separate test with more explanation in the prompt on how to classify between the four buckets.
38.
▲
by
kostaj
4mo ago
Data collection and processing was done manually. LLMs helped with the report drafting. Everything was human reviewed before publishing.
39.
▲
by
kostaj
4mo ago
Our experience with Lenz is that forcing a multi-step process, incl. adversarial debates, helps improve the verdicts.
40.
▲
by
kostaj
4mo ago
Author here. 67% (95% CI 64–70%) of 1,000 recent real user claims to a fact-checking platform had at least one of GPT-5.4, Claude Opus 4.7, Gemini 3 Pro, Gemini 3 Pro+Search, and Sonar Pro dissent from the panel majority — or no majority fo
41.
▲
Disagreement among frontier LLMs on real-world fact-checks
(lenz.io)
505 points
by
kostaj
4mo ago
|
347 comments