12 ms·
> Learning something online 5 years ago often involved trawling incorrect, outdated or hostile content and attempting to piece together mental models without th
by romaniitedomum 1y ago
> Learning something online 5 years ago often involved trawling incorrect, outdated or hostile content and attempting to piece together mental models without the chance to receive immediate feedback on intuition or ask follow up questions. This is leaps and bounds ahead of that experience.
But now, you're wondering if the answer the AI gave you is correct or something it hallucinated. Every time I find myself putting factual questions to AIs, it doesn't take long for it to give me a wrong answer. And inevitably, when one raises this, one is told that the newest, super-duper, just released model addresses this, for the low-low cost of $EYEWATERINGSUM per month.
But worse than this, if you push back on an AI, it will fold faster than a used tissue in a puddle. It won't defend an answer it gave. This isn't a quality that you want in a teacher.
So, while AIs are useful tools in guiding learning, they're not magical, and a healthy dose of scepticism is essential. Arguably, that applies to traditional learning methods too, but that's another story.
- QuantumGood 1y agoI often ask first, "discuss what it is you think I am asking" after formulating my query. Very helpful for getting greater clarity and leads to fewer hallucinations.
- deleted 1y ago[deleted]
- gronglo 1y ago> But now, you're wondering if the answer the AI gave you is correct or something it hallucinated. Every time I find myself putting factual questions to AIs, it doesn't take long for it to give me a wrong answer. I know you'll probably think I'm being facetious, but have you tried Claude 4 Opus? It really is a game changer.
- physix 1y agoA game changer in which respect? Anyway, this makes me wonder if LLMs can be appropriately prompted to indicate whether the information given is speculative, inferred or factual. Whether they have the means to gauge the validity/reliability of their response and filter their response accordingly. I've seen prompts that instruct the LLM to make this transparent via annotations to their response, and of course they comply, but I strongly suspect that's just another form of hallucination.
- cvoss 1y ago> But now, you're wondering if the answer the AI gave you is correct > a healthy dose of scepticism is essential. Arguably, that applies to traditional learning methods too, but that's another story. I don't think that is another story. This is the story of learning, no matter whether your teacher is a person or an AI. My high school science teacher routinely mispoke inadvertently while lecturing. The students who were tracking could spot the issue and, usually, could correct for it. Sometimes asking a clarifying question was necessary. And we learned quickly that that should only be done if you absolutely could not guess the correction yourself, and you had to phrase the question in a very non-accusatory way, because she had a really defensive temper about being corrected that would rear its head in that situation. And as a reader of math textbooks, both in college and afterward, I can tell you you should absolutely expect errors. The errata are typically published online later, as the reports come in from readers. And they're not just typos. Sometimes it can be as bad as missing terms in equations, missing premises in theorems, missing cases in proofs. A student of an AI teacher should be as engaged in spotting errors as a student of a human teacher. Part of the learning process is reaching the point where you can and do find fault with the teacher. If you can't do that, your trust in the teacher may be unfounded, whether they are human or not.
- tekno45 1y agoHow are you supposed to spot errors if you don't know the material? You're telling people to be experts before they know anything.
- ToucanLoucan 1y ago> You're telling people to be experts before they know anything. I mean, that's absolutely my experience with heavy LLM users. Incredibly well versed in every topic imaginable, apart from all the basic errors they make.
- maksimur 1y agoThey have the advantage to be able to rectify their errors and have a big leg up if they ever decide to specialize.
- yieldcrv 1y agoIf LLMs of today's quality were what was initially introduced, nobody would even know what your rebuttals are even about. So "risk of hallucination" as a rebuttal to anybody admitting to relying on AI is just not insightful. like, yeah ok we all heard of that and aren't changing our habits at all. Most of our teachers and books said objectively incorrect things too, and we are all carrying factually questionable knowledge we are completely blind to. Which makes LLMs "good enough" at the same standard as anything else. Don't let it cite case law? Most things don't need this stringent level of review
- kristofferR 1y agoAgree, "hallucination" as an argument to not use LLMs for curiosity and other non-important situations is starting to seem more and more like tech luddism, similar to the people who told you to not read Wikipedia 5+ years after the rest of us realized it is a really useful resource despite occasional inaccuracies.
- majormajor 1y agoFun thing about wikipedia is that if one person notices, they can correct it. [And someone's gonna bring up edit wars and blah blah blah disputed topics, but let's just focus on straightforward factual stuff here.] Meanwhile in LLM-land, if an expert five thousand miles a way asked the same question you did last month, and noticed an error... it ain't getting fixed. LLMs get RL'd into things that look plausible for out-of-distribution questions. Not things that are correct. Looking plausible but non-factual is in some ways more insidious than a stupid-looking hallucination.
- johnnyanmac 1y ago> to not use LLMs for curiosity and other non-important situations is starting to seem more and more like tech luddism We're on a topic talking about using an LLM to study. I don't particularly care if someone wants an AI boyfriend to whisper sweet nothings into their ear. I do care when people will claim to have AI doctors and lawyers.
- deleted 1y ago[deleted]
- ramraj07 1y agoWhat exactly did 2025 AI hallucinate for you? The last time I've seen a hallucination from these things was a year ago. For questions that a kid or a student is going to answer im not sure any reasonable person should be worried about this.
- tekno45 1y agoHow do you know? its literally non-deterministic.
- r3trohack3r 1y agoMost (all?) AI models I work with are literally deterministic. If you give it the same exact input, you get the same exact output every single time. What most people call “non-deterministic” in AI is that one of those inputs is a _seed_ that is sourced from a PRNG because getting a different answer every time is considered a feature for most use cases. Edit: I’m trying to imagine how you could get a non-deterministic AI and I’m struggling because the entire thing is built on a series of deterministic steps. The only way you can make it look non-deterministic is to hide part of the input from the user.
- throwaway31131 1y ago> I’m trying to imagine how you could get a non-deterministic AI Depends on the machine that implements the algorithm. For example, it’s possible to make ALUs such that 1+1=2 most of the time, but not all the time. … Just ask Intel. (Sorry, I couldn’t resist)
- tekno45 1y agoSo by default. Its non-deterministic for all non power users.
- sumeno 1y agoThis is an incredibly pedantic argument. The common interfaces for LLMs set their temperature value to non-zero, so they are effectively non-deterministic.
- teleforce 1y agoPlease check this excellent LLM-RAG AI-driven course assistant at UIUC for an example of university course [1]. It provide citations and references mainly for the course notes so the students can verify the answers and further study the course materials. [1] AI-driven chat assistant for ECE 120 course at UIUC (only 1 comment by the website creator): https://news.ycombinator.com/item?id=41431164 https://news.ycombinator.com/item?id=41431164
- ink_13 1y agoGiven the propensity of LLMs to hallucinate references, I'm not sure that really solves anything
- ok123456 1y agoRAG means it injects the source material in and knows the hash of it and can link you right to the source document.
- PaulRobinson 1y agoI've worked on systems where we get clickable links to source documents also added to the RAG store. It is perfectly possible to use LLMs to provide accurate context. It's just asking a SaaS product to do that purely on data it was trained on, is not how to do that.
- mvieira38 1y agoI haven't seen it happen at all with RAG systems. I've built one too at work to search internal stuff, and it's pretty easy to make it spit out accurate references with hyperlinks
- ay 1y agoMy favourite story of that involved attempting to use LLM to figure out whether it was true or my hallucination that the tidal waves were higher in Canary Islands than in Caribbean, and why; it spewed several paragraphs of plausibly sounding prose, and finished with “because Canary Islands are to the west of the equator”. This phrase is now an inner joke used as a reply to someone quoting LLMs info as “facts”.
- ricardobeat 1y agoThis is meaningless without knowing which model, size, version and if they had access to search tools. Results and reliability vary wildly. In my case I can’t even remember last time Claude 3.7/4 has given me wrong info as it seems very intent on always doing a web search to verify.
- flir 1y agoThere's something darkly funny about that - I remember when the web wasn't considered reliable either. There's certainly echoes of that previous furore in this one.
- otabdeveloper4 1y ago> I remember when the web wasn't considered reliable either It still isn't.
- commakozzi 1y agoYes, it still isn't, we all know that. But we all also know that it was MUCH more unreliable then. Everyone's just being dishonest to try to make a point on this.
- flir 1y agoI'm more talking about the conversation around it, rather than its absolute unreliability, so I think they're missing the point a bit. It's the same as the "never use your real name on the internet" -> facebook transition. Things get normalized. "This too shall pass."
- jimmaswell 1y ago> you're wondering if the answer the AI gave you is correct or something it hallucinated Regular research has the same problem finding bad forum posts and other bad sources by people who don't know what they're talking about, albeit usually to a far lesser degree depending on the subject.
- bradleyjg 1y agoThe difference is that llms mess with our heuristics. They certainly aren’t infallible but over time we develop a sense for when someone is full of shit. The mix and match nature of llms hides that.
- Paradigma11 1y agoYou need different heuristics for LLMs. If the answer is extremely likely/consistent and not embedded in known facts alarm bells should go off. A bit like the tropes in movies where the protagonists get suspicious because the antagonists agree to every notion during negotiations because they will betray them anyway. The LLM will hallucinate a most likely scenario that conforms to your input/wishes. I do not claim any P(detect | hallucination) but my P(hallucination | detect) is pretty good.
- y1n0 1y agoYes but that is generally public, with other people able to weigh in through various means like blog posts or their own paper. Results from the LLM are your eyes only.
- dyauspitr 1y agoNo you’re not, it’s right the vast, vast majority of the time. More than I would expect the average physics or chemistry teacher to be.
- m463 1y agoYou should practice healthy skepticism with rubber ducks as well: https://en.wikipedia.org/wiki/Rubber_duck_debugging https://en.wikipedia.org/wiki/Rubber_duck_debugging
- reactordev 1y agoWhile true, trial and error is a great learning tool as well. I think in time we’ll get to an LLM model that is definitive in its answer.
- Paradigma11 1y agoJust have a second (cheap) model check if it can find any hallucinations. That should catch nearly all of them in my experience.
- apparent 1y agoWhat is an efficient process for doing this? For each output from LLM1, you paste it into LLM2 and say "does this sound right?"? If it's that simple, is there a third system that can coordinate these two (and let you choose which two/three/n you want to use?
- Paradigma11 1y agoMarkdown files are everything. I use LLMs to create .md files to create and refine other .md files and then somewhere down the road I let another LLM write the code. It can also do fancy mermaid diagrams. Have it create a .md and then run another one to check that .md for hallucinations.
- niemandhier 1y agoYou can use existing guardrails software to implement this efficiently: NVIDIA NeMo offers a nice bundle of tools for this, among others an interface to Cleanlabs API to check for thruthfullness in RAG apps.
- Paradigma11 1y agoI realized that this is something that someone with Claude Code could reasonably easily test (at least exploratively). Generate 100 prompts of "Famous (random name) did (random act) in the year (random). Research online and elaborate on (random name) historical significance in (randomName)historicalSignificance.md. Dont forget to list all your online references". Then create another 100 LLMs with some hallucination Checker claude.md that checks their corresponding md for hallucinations and write a report.md.
- UltraSane 1y agoI had teachers tell me all kinds of wrong things also. LLMs are amazing at the Socratic method because they never get bored.
- PaulRobinson 1y agoDespite the name of "Generative" AI, when you ask LLMs to generate things, they're dumb as bricks. You can test this by asking them anything you're an expert at - it would dazzle a novice, but you can see the gaps. What they are amazing at though is summarisation and rephrasing of content. Give them a long document and ask "where does this document assert X, Y and Z", and it can tell you without hallucinating. Try it. Not only does it make for an interesting time if you're in the World of intelligent document processing, it makes them perfect as teaching assistants.
- nathias 1y agodid you trust everything you read online before?
- johnnyanmac 1y agoDid you get to see more than one source calling out or disagreeing with potential untrustworthy content? You don't get that here.
- nathias 1y agoof course you do, you have links to sources
- johnnyanmac 1y ago> for the low-low cost of $EYEWATERINGSUM per month. This part is the 2nd (or maybe 3rd) most annoying one to me. Did we learn absolutely nothing from the last few years of enshittification? Or Netflix? Do we want to run into a crisis in the 2030's where billionaires hold knowledge itself hostage as they jack up costs? Regardless of your stance, I'm surprised how little people are bringing this up.
- deleted 1y ago[deleted]
- demarq 1y agoTo be honest I now see more hallucinations from humans on online forums than I do from LLMs. A really great example of this is on twitter Grok constantly debunking human “hallucinations” all day.
- littlestymaar 1y agoAh yes, like when Grok hallucinated Obama and Biden in a picture with two drunk dudes (both white, BTW).
- Wilder7977 1y agoLets not forget also the ecological impact and energy consumption.
- wing-_-nuts 1y agoHonestly, I think AI will eventually be a good thing for the environment. If ai companies are trying to expand renewables and nuclear to power their datacenters for training, well, that massive amount of renewables and battery storage becomes available when training is done and the main workload is inference. I know they are consistently training new stuff on small scale but from what I've read the big training batches only happen when they've proven out what works at small scale. Also, one has to imagine that all this compute will help us run bigger / more powerful climate models, and google's ai is already helping them identify changes to be more energy efficient. The need for more renewable power generation is also going to help us optimize the deployment process. I.e. modular nuclear reactors, in situ geothermal taking over old stranded coal power plants, etc
- Wilder7977 1y agoI find this take overly optimistic. First, it's bases on the assumption that the training will stop, and that energy will be available for other, more useful, purposes. This is not guarantees. Besides this, it completely disregards the fact that today, tomorrow, energy will be utilized. We will keep emitting co2 for sure, and maybe, in the future, this will cause a surplus of energy? It's a bet I wouldn't take, even because LLMs need lots of energy to run as well as for training. But in any case, I wouldn't want Microsoft, Google, Amazon and OpenAI to be the ones owning the energetic infrastructure in the future, and if we realize, collectively, that building renewable sources is what er need, we should simply tax them and use that wealth to build collective resources.
- ThatMedicIsASpy 1y agoI ask: What time is {unix timestamp} ChatGPT: a month in the future Deepseek: Today at 1:00 What time is {unix timestamp2} ChatGPT: a month in the future +1min Deepseek: Today at 1:01, this time is 5min after your previous timestamp Sure let me trust these results...
- ThatMedicIsASpy 1y agoAlso since I was testing a weather API I was suspicious of ChatGPTs result. I would not expect weather data from a month in the future. That is why I asked Deepseek in the first place.
- tim333 1y ago>But now, you're wondering if ... hallucinated A simple solution is just to take <answer> and cut and paste it into Google and see if articles confirm it.
- threecheese 1y ago> you're wondering if the answer the AI gave you is correct or something it hallucinated Worse, more insidious, and much more likely is the model is trained on or retrieves an answer that is incorrect, biased, or only conditionally correct for some seemingly relevant but different scenario. A nontrivial amount of content online is marketing material, that is designed to appear authoritative and which may read like (a real example) “basswood is renowned for its tonal qualities in guitars”, from a company making cheap guitars. If we were worried about a post-truth era before, at least we had human discernment. These new capabilities abstract away our discernment.
- ijk 1y agoThe sneaky thing is that the things we used to rely on as signals of verification and credibility can easily be imitated. This was always possible--an academic paper can already cite anything until someone tries to check it [1]. Now, something looking convincing can be generated more easily than something that was properly verified. The social conventions evaporate and we're left to check every reference individually. In academic publishing, this may lead to a revision of how citations are handled. That's changed before and might certainly change again. But for the moment, it is very easy to create something that looks like it has been verified but has not been. [1] And you can put anything you like in footnotes.
- berkes 1y agoIs this a fundamental issue with any LLM, or is it an artifact of how a model is trained, tuned and then configured or constrained? A model that I call through e.g. langchain with constraints, system prompts, embeddings and whatnot, will react very different from when I pose the same question through the AI-providers' public chat interface. Or, putting the question differently: could OpenAI not train, constrain, configure and tune models and combine them into a UI that then acts different from what you describe for another use case?
- roudaki 1y agoThe joke is on you, I was raised in Eastern Europe, where most of what history teachers told us was wrong That being said. as someone who worked in a library and bookstore 90% of workbooks and technical books are identical. NotebookLM's mindmap feature is such a time saver