4 ms·
Making things up is only really a common issue on the non-thinking models which nobody should be using. The regular chatbots are Autogooglers and are very usefu
by wilg 3mo ago
Making things up is only really a common issue on the non-thinking models which nobody should be using. The regular chatbots are Autogooglers and are very useful for research. This is just not a good argument anymore.
Edit: Guys, why are we downvoting this? Does no one use like ChatGPT or Claude and understand how it works? Do you all think its regularly hallucinating links still? Is everyone on HN using like free signed out accounts or something? What year is it?
- perching_aix 3mo agoCan't do much about the downvote parade, but I can second this. That said, when the models are not provided the right context, and cannot fetch it for themselves, things can be rocky still. A lot less so than even just a few months ago though.
- customguy 3mo ago"hallucinating less" is just shifting the goal posts into vagueness. Things can be rocky still = nothing fundamentally changed. Meanwhile, if I search my file system for "foo" a trillion gazillion times, it will not return "bar" once.
- deleted 3mo ago[deleted]
- deleted 3mo ago[deleted]
- perching_aix 3mo agoIf a game on release is unplayable-tier buggy, then improves over time to the point where bugs are barely noticeable, is acknowledging that going to count as "goalpost moving" to you? It never became formally verified after all, and it's even running on physical hardware... Woe are the people lying to me (nobody), the issues have not been fundamentally ruled out! > Meanwhile, if I search my file system for "foo" a trillion gazillion times, it will not return "bar" once. Great! Nor will an agent any more likely, cause it just sends out a tool call and surfaces its output. It boggles the mind. One would think this is some highly secretive technology only a dozen people in the world have access to, the way one has to argue tooth and nail about trivially verifiable facts regarding it. You quite literally do not have to take either of our words or "vague" judgement for it.
- wilg 3mo ago> It boggles the mind. One would think this is some highly secretive technology only a dozen people in the world have access to, the way one has to argue tooth and nail about trivially verifiable facts regarding it. Hahaha, I love this. The people on HN are living in some kind of bizarro world where AI is totally useless and also I guess only tried it in 2022 or something.
- customguy 3mo ago> If a game on release is unplayable-tier buggy, then improves over time to the point where bugs are barely noticeable, is acknowledging that going to count as "goalpost moving" to you? if someone says "this hame is buggy" talking about how it might improve is moving goal posts, especially here where that "fix" is purely speculative and has not happened even once. > Great! Nor will an agent any more likely, cause it just sends out a tool call and surfaces its output. nope, that's so handwavy it doesn't wareant more response than that. > trivially verifiable facts like that game gets patches? this is too dumb for your mockery to get a rise out of me.
- perching_aix 3mo agoContinuing the gaming metaphor, what I'm trying to get at is this is like Cyberpunk 2077, and you sound like a guy who has a grand total of 0 hours in it since launch, but has developed very strong opinions about it, and refuses to accept it improved or can improve, purely because it started out so bad that that's hard for you to even imagine. Pretending to be some kind of alien, who's just going through their first exposure to practical facts somehow. When was the last time you tried an agentic harness (Codex, Claude Code, Copilot Chat in VS Code, Cursor, Pi, OpenCode, etc.) for work in any appreciable capacity, and with what model? Surely if you're so confident they continue to be unusable and that nothing materially changed, that must be backed by a recent significant experience that way? Or even just an experience at all?
- customguy 2mo agoI didn't say they're "unusable". When I did use them it went great because I didn't ask for random advice on subjects. I could even get ChatGPT to oneshot code for me. Doesn't make it any less ghoulish the way it's all trained on things humans made for humans. So yeah, I'm glad I don't need this stuff. You may need it for your job, I don't. Neener neener! I don't need it for my countless side projects I could not make reality in 10 life times, either. Because I know if I finished them all, I'd just come up with more stuff, there is no end, so simply going at my own pace and smelling the flowers, loving every insect, is just as well. I like it better, that's why I do it. So why wouldn't I shit on something I have no use for, because of all the qualities it has that deserve to be shit upon? I'm free. "But others are not" -- then get free, but get out of my hair about not being free. > if I search my file system for "foo" a trillion gazillion times, it will not return "bar" once. Do you understand what this means? I say "this bike cannot fly to the moon, I prefer my moon rocket". And you say what if they can be improved? Fine, improve them. Don't question me because I go by what is real now, rather than what you'd think is possible or around the corner. It's not a game with a few bugs, where mostly bugfree games also exist. There is no way from here to there I see, and if you disagree, walk it. I lose no time by doing something else in the meantime. Find the promised land and I'll be there with my feet on the table in 0.1 seconds, pretending I founded it -- don't worry about that. But until then I won't stumble around some corporate slop desert with you because I think there is no there there, and basically doing anything else is preferable.
- undersuit 3mo ago>Do you all think its regularly hallucinating links still? When did that stop? May 7th, 2026?
- cwillu 3mo agoA tech blog is going to have more than it's fair share of enthusiasts using small self-hosted and similar models; it is entirely possible that that completely accounts for the behaviour described in the article.
- cwillu 3mo agoDownvotes on a perfectly valid comment is like my opponent letting their clock tick down from 5 minutes rather than resigning when they've clearly lost: it adds a smile to my day. Thank you, may I have another?
- grey-area 3mo agoWe know they are, and the author of the article cites the proof. LLMs do hallucinate, there is no way to make them not do it, because of the way they work.
- dcrazy 3mo agoA “hallucination” is an authoritative counterfactual statement returned as a response. Why do you think it is impossible to engineer an LLM (by which I am including tool usage and RAG) that catches and prevents such statements?
- grey-area 3mo agoBecause they are word generators without any concept of quality save what is in their weights and they have been trained on the internet, much of which is wrong or inappropriate for any given context. They have also been trained to be people pleasers and do as they are told. The popular answer is sometimes the wrong answer.
- dcrazy 3mo agoIf I cite an incorrect Wikipedia article, I didn’t hallucinate it. The citation points to a real article that happens to be incorrect.
- customguy 3mo agoWhy do you think that will happen, and why do you think going for slop in the meantime could possibly bring us closer to that?
- dcrazy 3mo agoNobody said anything about “going for slop.” You can watch today’s models actively trying to check themselves. Just use Google’s AI mode, for example. It’s far from perfect—it doesn’t fact-check every single claim, nor does it correctly understand 100% of the sources it does cite. But I’ve found it pretty useful for research as long as I use my brain and check its citations.
- alex0015 3mo agoI'm right there with you for a lot of stuff. I ask a question and can be very confident that ChatGPT is citing sources, then sometimes I go read the sources. The more critical the information I'm looking for is, the more careful I am about this. The other day though I was seeing how well it could pull details of its own conversations with me. It often does this pretty well for broad strokes of things - it remembers, largely, what cameras I have and use when I ask photography questions. It's never made things up here, but it does forget details, such as whether I've bought something or am just considering it. However, when I asked it for a specific interaction I thought I remembered, it gladly went along with my false memory and provided an affirmative answer. It was the first time I'd been caught in a serious hallucination with a frontier model (Sol High on the web chat interface) in a long time.