5 ms·
GPT-4o is consistently hallucinating
- freesam 2y agoWe found that gpt-4o is consistently hallucinating. Here is an example: [ https://platform.openai.com/playground/chat?models=gpt-4o&preset=yEwza9Ibnw3RnQxGNnTwzOqd https://platform.openai.com/playground/chat?models=gpt-4o&pr... ] In this case, we never mentioned any store location in Ringsted. However , gpt-4o still fakes a store in the address: Adresse: Ringstedet, Klosterparks Allé 10, 4100 Ringsted Åbningstider: Mandag-fredag: 9.30-19, Lørdag: 10-17, Søndag: 11-16 Tlf: 50603398 Email: ringsted4100@gmail.com This hallucination is quite consistent. The gpt-4-turbo or even gpt-3.5 model doesn't have the same issue and correctly acknolwedged that there is no store in this location.
- andrei512 2y agoThis isn't front-page material...
- grrowl 2y agoTruly, (for those who don't want to log in) it's just a shared playground conversation with GPT-4, in Danish, where the author got a result they're obviously unhappy with. Author, do some more robust tests, write a blogpost about it, and submit that instead.
- bhaney 2y ago> Please log in to access this page No thanks. Not like an LLM hallucinating is particularly surprising or newsworthy anyway.
- imdsm 2y agoIt's a link to a playground session that's why btw
- yencabulator 2y agoIt's a link to a playground session, sure, but that doesn't explain in any way why one would have to log in to see it.
- threeseed 2y agoNot sure what you are expecting. The models are non-deterministic and there is no way to predict hallucinations.
- gmerc 2y agoto be fair, a standard model is deterministic at temp==0. And there is a way to predict the presence of hallucinations, but it’s expensive (SelfCheckGPT)
- lionkor 2y agoI found that, compared to GPT-3.5, it refuses to shut up when told to shut up. In the middle of a conversation, try going "SHUT UP, STOP TALKING ALREADY". For me, it just keeps repeating the last output. Very cool.
- dailykoder 2y agoIt feels like the answers are getting longer and longer too. Even for the most basic questions, which could be answered with 2 sentences. Does it have ADHD? Who wants to read all these wall of text?
- gmerc 2y agoWell it’s paid by token
- dailykoder 2y agoBut even the gpt3.5 answers were getting longer and longer. I don't know, I don't pay for it myself, we just have 4o at cagie and I don't know how that's different in terms of the tuning compared to the "normal" 4o
- emsign 2y agoOh. Oooh! Yes. And "I don't know." aren't a lot of tokens. So where's the incentive there, lol.
- deleted 2y ago[deleted]
- mlyle 2y agoYes: GPT-4 turbo could receive a meaningful correction and generally change its answer in that direction. GPT-4o is very, very resistant to doing this and will tend to parrot the previous answer, even after admitting it was in error. I routinely fix this by toggling from GPT-4o to GPT-4t.
- 2y ago
- pietz 2y agoI haven't noticed that GPT-4o hallucinates a lot more than the previous version but I noticed 2 other things of which especially the latter seems relevant here. 1) it's insanely chatty, to a point where it ignores instructions about not doing certain things. I think this behavior is heavily favoured by benchmarks but as somehow who expects concise answers, this model annoys me. Custom instructions don't fully fix this for me. 2) It likes repetitive answers a lot more than the previous version. Meaning that it will try its hardest to generate the followup answer in the same format as the first one. I think this is also the problem in your example. To my understanding, this is a measure against laziness, where the model would exclude information from the first answer that haven't changed in the followup. I always liked this behavior but maybe you remember the time from a few months ago where many people complained about the laziness of (I believe) 0125. Btw, while I type this, I notice that this is probably the highest level of first world problems I've ever complained about. There is this amazing almost free tool that answers all my questions and does most of my coding and I dislike it because it provides me with thorough context.
- emsign 2y agoGiven that this is highly inefficient from a user's point of view and every request costs energy and someone else's money, I wouldn't call it a first world problem but a design flaw with more serious implications than just repercussions for annoyed customers. It stacks up.
- me_vinayakakv 2y agoI am using GPT 4o and have observed both 1 and 2. Based on a tweet in X[1], I had to add "I REPEAT" to the instruction to get the model to not to ignore instructions. [1]: https://x.com/btibor91/status/1796077902959640893 https://x.com/btibor91/status/1796077902959640893
- laborcontract 2y agoThe instruction ignoring piece is noticeable to the extent that it sometimes reminds me of 3.5-turbo. I just wonder if it’s a side effect of their training, or whatever they did to make the model more efficient.