3 ms·
Obscene levels of hallucinations, the worst of LLMs, unfortunately. Deepseek v4 pro 94% Deepseek v4 flash - 96% https://artificialanalysis.ai/evaluations/omn
by scrollop 5mo ago
Obscene levels of hallucinations, the worst of LLMs, unfortunately.
Deepseek v4 pro 94%
Deepseek v4 flash - 96%
https://artificialanalysis.ai/evaluations/omniscience?models=gemini-3-1-pro-preview%2Cgpt-5-5%2Cgrok-4-3%2Cclaude-sonnet-4-6-adaptive%2Cgemini-3-flash-reasoning%2Cqwen3-6-max%2Ckimi-k2-6%2Cgpt-5-4%2Cmimo-v2-5-pro%2Cglm-5-1%2Cminimax-m2-7%2Cclaude-4-5-haiku-reasoning%2Cdeepseek-v4-pro%2Cgpt-5-4-mini%2Cdeepseek-v3-2-reasoning%2Cdeepseek-v4-flash%2Cqwen3-5-397b-a17b%2Cmistral-small-4%2Cnvidia-nemotron-3-super-120b-a12b%2Cnova-2-0-pro-reasoning-medium%2Cgpt-oss-120b%2Cgpt-oss-20b#omniscience-hallucination-rate-tabs https://artificialanalysis.ai/evaluations/omniscience?models...
- _0ffh 5mo agoPersonally, I'm not bothered very much by LLM confabulation, as long as it's the result of missing context. In most practical tasks, we either give context to the model, or tell it to find it itself using the internet. What I am concerned with is confabulation that contradicts available in-context information, but that doesn't seem to be what is measured here.
- dust42 5mo agoThe output of any LLM is always 100% hallucination by principle. On top of that, most benchmarks are at best an approximation of LLM quality. Your use case decides which one to use. That said, I haven't tested v4 yet but the old 3.2 is still a decent model. And concerning use cases, I had coding problems that Opus couldn't solve but a local 35B model did. All the talk about frontier and SOTA is do dig deeper and deeper into the pockets of VCs and finally do an IPO.
- UlisesAC4 5mo agoThis must be easily benchmaxed because I have never gotten an "idk like" answer for the western frontier models. All my personal "real world" use cases will always resort to hallucinations.