4 ms·
> My point is that the hallucination rate of modern LLMs, while not zero, is so low that this is no longer an issue in practice. Ahahahahahahahahahahahahahahah
by sjsdaiuasgdia 3mo ago
> My point is that the hallucination rate of modern LLMs, while not zero, is so low that this is no longer an issue in practice.
Ahahahahahahahahahahahahahahahahaha
Do you actually use these things?
- Marha01 3mo ago> Do you actually use these things? I use them every day for SWE. I cannot remember any hallucination from the last 6 months. It pretty much stopped being an issue from the beginning of this year, if you use top paid models. Free models or lower-tier paid models can still hallucinate, though.
- sjsdaiuasgdia 3mo agoI can use Claude as much as I want at work, and I get hallucinations every day. I'm not even doing anything all that fancy with it, but I get frequent hallucinatory behavior regardless of the model or how I go about things. Every single time I have asked Claude or another "top tier" offering a detailed question about something I'm an expert in, there's something wrong in the answer. Usually many things. Do you actually look at what comes out?
- wizzwizz4 3mo agoI would suggest that your ability to discern when the model is wrong might be the thing that has changed.
- duskdozer 3mo agoThis is what it seems to be. I've had to review the LLM output from coworkers who have insisted that the "old" models are terrible but the new ones are very good. They are always discussing the latest models and their tweaks and things. Their output is just.... not very good. For example, it may not fail tests, but it does so for wrong reasons and will fail under different conditions. Or it checks or guards for things in trivially redundant ways, like doing something like `if (x > 5 && x > 3)` that betray its lack of "knowing" what it's doing. And I can't even get answers from them anymore because they just feed anything I say into the LLM and copy/paste the response. I'm basically being forced to code with LLMs via review through a person proxy. Or maybe they just have it hooked up to read it directly. It's maddening. It's like having a realistic enough chatbot just shuts people's brains off.
- 27183 3mo agoIt would be extremely valuable to be able to reliably identify this kind of behavioral tendency in an interview setting. Some kind of reliable test that screens for likelihood to go AI psychotic.
- hnisfulomrons 3mo ago[dead]