6 ms·
> What you get is a beautiful animation that is 100% accurate and free of hallucinations. I'm not sure I follow how this is actually guaranteed? The fact-check
by wxw 2mo ago
> What you get is a beautiful animation that is 100% accurate and free of hallucinations.
I'm not sure I follow how this is actually guaranteed? The fact-checking process mentioned just seems to involve asking AI to review its own work.
- dozerly 2mo agoAll these LLM-as-review hype pieces don’t acknowledge that it’s turtles all the way down
- enraged_camel 2mo agoWhat do you mean by this?
- Gander5739 2mo agoIf the output can't be trusted, and you use another llm whose output can't be trusted to check the untrusted output of the first llm, then you're back where you started.
- duncangh 2mo agoYeah this seems to me similar to how the mortgage backed security risk concentration occurred leading up to the global financial crisis. Whereby the risk from exposure to low grade / risky single mortgages was eliminated via diversification but the diversification was simply packaging all of the risky MBS’s together and in no way diversified or de-risked the entire portfolio
- Terr_ 2mo agoI'm hoping that the Big Horrible Realization comes sooner rather than later, when we have less collective damage and pain riding on it. (Plus I'd feel personally vindicated.)
- cromka 2mo agoI don't see it. To me it's like having e.g. 3 drunk PhDs arguing between each other to settle on truthful answers to questions.
- DrewADesign 2mo agoBut the problem with LLMs is that they get the facts wrong. PhDs are PhDs because they’d look it up in an authoritative source, or actually find out through research and experimentation. The whole point is that facts aren’t a matter of opinion. The only people that argue over documented, findable facts are idiots that nobody should listen to.
- daishi55 2mo agoNot really. Take hallucinations for example. If they are 1 in 100 (actually they are much rarer, but for the sake of argument), then the chances that 2 LLMs or even just 2 runs of the same LLM have the same hallucination is, well, a lot less than 1 in 100.
- deleted 2mo ago[deleted]
- Terr_ 2mo agoThat rests on a false-assumption that the errors are statistically independent events, and have nothing to do with the shared nature of the judges.
- daishi55 2mo agoAre there any reproducible hallucinations on any of the currently available OAI/Anthropic models? I’m not aware of any. And even if they are related - if Opus 4.8 always has a 1:100 chance of a specific hallucination - then running the same model twice does indeed dramatically reduce the odds of an error in the final output.
- Terr_ 2mo agoIf simply running things thrice-over was enough to stop "hallucinations" (and not incur other problems) we wouldn't be here talking about it today, it'd have been "solved" months or years ago.
- daishi55 2mo agoI mean they have been solved practically I think. I use these things all day every day and haven’t seen a hallucination in a long time.
- djhn 2mo agoConstant hallucinations. OpenAI:s latest on max settings. If you are to naively feed say, a short corpus of text to turn it into a parallel corpus in a few different languages, the original text gets subtly mangled and no longer matches the original. Say you have several hundred annotated sentences. Without hand-coding some regex to make sure that each sentence in the source column occurs in the original corpus you’re bound to get hallucinated sentences with an error rate that exceeds 1:100. Whatever you use as the output, JSON or XML, you will end up with columns that just repeat the original instead of translating it, especially for languages that are very close to each other or represent the same language. Yes, LLMs can be SOTA for NLP, but you’re going to have to use them to write software or workflows that are more deterministic.
- ex-aws-dude 2mo agoNo you don’t get it, I asked it specifically to make sure it’s accurate
- northern-lights 2mo agoUnfortunately, I don;t think that can ever be fixed. For an LLM to know that it is not hallucinating about something, it must know that that statement(s) is/are true. Which it cannot infer due to Godel's Incompleteness theorems.
- phatskat 2mo agoI read it as “if your LLM is being checked by another LLM, well then you need another LLM to check the checker. And can you really trust _that_ LLM? Probably should have an LLM to check the third one, and…”
- lordnacho 2mo ago[dead]
- furyofantares 2mo ago"Turtles all the way down" is a phrase of, I think, unknown origin (https://en.wikipedia.org/wiki/Turtles_all_the_way_down https://en.wikipedia.org/wiki/Turtles_all_the_way_down) about infinite regress or trying to patch up some bad theory by appealing to itself. Someone claims that what holds the Earth in place is that it sits atop a giant turtle, and a skeptic asks what holds the turtle up, and the response is that it's turtles all the way down. Personally I think this is a bad characterization of using LLMs to fix up LLMs because while you can never guarantee results this way (as the quoted line claims here, which is worthy of criticism), it is, in practice, useful to use LLMs on top of LLMs. And there's no infinite regress. Auto-mode in Claude Code, for example, seems to me like it's been successful at making the system more safe than --dangerously-bypass-permissions without prompting the user for permissions constantly.
- dozerly 2mo agoThere are certainly uses where it’s good enough, but you can never be 100% certain of correctness in the way that people claim you can by stacking N layers of these models. What triggered my response was the “just review the output with another LLM and it’s perfectly correct”
- TooSmugToFail 2mo agoI would have bet it’s a Terry Pratchett quote, and it kinda is: “"The turtle moves," said Didactylos. "The turtle is a giant reptile that swims through space. It doesn't have to stand on anything. Swimming is what turtles do. The idea that it has to stand on another turtle, and that turtle has to stand on another turtle, is just silly. It's turtles all the way down, and that's a logical absurdity." Small Gods, 1992
- furyofantares 2mo agoHe's using it but if you CTRL+F the linked wikipedia page for 1967 you get someone using that exact phrase, or for 1882 above you get a pretty close version.
- tclancy 2mo agoIt’s a reference to Bernard Shaw, who once said that if we ever created a truly artificial mind it would be inside a turtle’s shell. Sturgill Simpson covered the track on his seminal work, Xeno’s Paradox.
- eyepea2007 2mo agoan LLM tells me there is no evidence that Bernard Shaw ever said any such thing :)
- arvid-lind 2mo agoIt's Sturgill Simpson all the way down: https://en.wikipedia.org/wiki/Turtles_All_the_Way_Down_(song) https://en.wikipedia.org/wiki/Turtles_All_the_Way_Down_(song...
- tclancy 2mo agoIn my defense, I was suffering from white line fever.
- tclancy 2mo agoWell sure, it was George Bernard Shaw, not the old CNN anchor. Also, turtles are tight lipped by nature.
- Dumblydorr 2mo agoMight be a reference to the story at the beginning of A Brief History of Time, attributed to Bertrand Russell’s audience member. Full story in the book
- tclancy 2mo agoThis guy remembers what I thought I was saying. In my defense I stole the whole thing from Stephen King’s It which I read … forty years ago, that can’t be accurate. Let me sort out my instruments and get back to you.
- latexr 2mo agoOther comments have answered in the concrete what was meant, but to answer in the abstract, in case you’re unfamiliar with the expression: https://en.wikipedia.org/wiki/Turtles_all_the_way_down https://en.wikipedia.org/wiki/Turtles_all_the_way_down
- bryanrasmussen 2mo agoI could certainly envision a scenario whereby review would increase reliability but not how it would every guarantee 100%, there is a pretty big logical gap there.
- laurentiurad 2mo agoIn my experience it depends on how much in detail you want to go. Chip manufacturing is a really opaque industry, so in this particular case LLMs might not even have the training data. However, using it for a high-level introduction into something is usually pretty safe from hallucinations.
- hn_throwaway_99 2mo agoCompletely agree. While I didn't set things up to have AI review its output in a loop, my experience trying to build a specific acoustic testing rig with Opus 5 also aligns with the other "it's turtles all the way down" comment. Opus 5 first built me a detailed plan, but a couple important details were either obviously wrong or felt unnecessary. I went back and forth asking for sources and more information probably like 4 times and every time it did the "in looking at things in more detail it appears my previous advice was incorrect" spiel. It just became exhausting at some point because it feels like it really lays bare how LLMs are just minimizing that loss function but don't actually "understand" anything. It was really useful as a search engine (it correlated some highly relevant source docs), but I just couldn't trust it to believe it was actually done at any step.
- knollimar 2mo agoA second agent reviewing it adversarially resolves some context rot. Whatever they trained these LLMs on will just infinitely double down so I break it with 1 layer of checking and then a judge who looks at facts, since the checker is adversarial. I know it sounds silly but 1 layer ends uo being way worse than 2.
- tempacct2cmmnt 2mo agoAgreed. Given how many significant errors LLMs make in my topic of expertise, despite my taking multiple error checking steps, the idea of catching 100% of hallucinations because you told the LLM to check itself is hilarious. It’s just a wild lack of insight: “I’m using the LLM to teach me something I don’t know about, I definitely have the knowledge base to spot any errors that might remain!”
- DrewADesign 2mo agoYeah. People with technical and/or tech business bonafides claiming that AI granted them expertise are so often taken at face value when they really shouldn’t be. Who told them that they were proficient — a chatbot? Someone who knows even less about the topic, so any expertise seems impressive? I’ll bet it wasn’t someone that actually knew what they were talking about. Even some tech reporters are tripping over themselves to be amazed, but don’t bother checking if they should be. People don’t even have to be lying to be wrong about this stuff. Someone can learn enough about a topic to be halfway up Mt. Stupid in no time flat, and in doing so, think they not only truly understand the topic at hand, but might be particularly adept because they were such quick studies. People that know less are impressed, because why wouldn’t they be? Anybody that knows more than them sounds like an expert. And people that know what they’re talking about cringe at the overconfidence, and probably try not to engage: who wants to have to prove that someone’s boundless confidence is entirely baseless? Most of the time, they think the actual expert is full of shit because they think they’re the expert. It’s incredible how many times I’ve had people in tech confidently, even smugly “explain” design concepts and strategies to me that they did not actually understand, knowing I was an experienced, degree-holding designer… and they didn’t even have a chatbot’s lips on their ass telling them how smart and insightful they were.
- latexr 2mo ago[dead]
- a3w 2mo agoI thought after reading the title that the text was about learning something, yet the actual text seems to be about having a system do something for me.
- spike021 2mo agoeven if you say use RAG or something to a source you can trust, there's no guarantee the agent will still use exactly what the source has. i can't even get agents to remember core instructions like "use jq instead of writing a python script to parse some json"..
- deleted 2mo ago[deleted]
- 63stack 2mo agoYou can add "make no mistakes" to the end of the prompt and achieve the same result while burning less tokens.
- lowsong 2mo agoYep, that's impossible. The hard truth is that most people this lost to LLM psychosis cannot understand that fact. It's better to treat it like someone in a cult, arguing the facts isn't going to help if they refuse to accept them.
- zahrevsky 2mo agoI don't do animations, but I have an answer. You research a topic well enough to be able to understand if the result is OK or not. Usually it means figuring out some sort of testing. I'm researching causal inference right now, and my main goal was to make sure I understand how to test estimation on synthetic data. Basically, it's the same way it works with people. If you delegate a task that you don't understand, and you can't have a credibility proof (i.e. doctors, lawyers), then you research a topic well enough to be able to (1) define the task and (2) verify the end result.
- SwtCyber 2mo agoThe "100% accurate and free of hallucinations" claim should probably be replaced with something much weaker
- nedt 2mo agoI'm just looking at the rocker engine piece. The belt is running backwards. That's the first thing I'd expect it to get right or flag if it doesn't. Human in the loop is also the quality control. If that fails the rest might have similar issues.
- rootlocus 2mo agoOP posted his rollercoaster tycoon inspired simulator that explained how chips are fabricated. Comments were full of people pointing out inaccuracies and hallucinations. The problem with LLM explanations of unknown topics is that you literally cannot determine how right or wrong it is. I usually ask LLMs to bring references and they almost always admit they pulled random shit out of their ass and quickly appologize when evidence to the contrary surfaces.