5 ms·
I find ChatGPT to be great at research too-but there are pathological failure modes where it is biased to shallow answers that are subtly wrong, even when defin
by larsiusprime 1y ago
I find ChatGPT to be great at research too-but there are pathological failure modes where it is biased to shallow answers that are subtly wrong, even when definitive primary sources are readily available online:
https://www.fortressofdoors.com/researchers-beware-of-chatgpts-wikipedia-brain/ https://www.fortressofdoors.com/researchers-beware-of-chatgp...
- jbm 1y agoYes, this is very much my experience too. Switching to GPT5 Thinking helps a little, but it often misses things that it wouldn't when I was using o3 or o1. As an example, I asked it if there were any incidents involving Botchan in an Onsen. This is a text that is readily available and must have been trained on; in the book, Botchan goes swimming in the onsen, and then is humiliated when the next time he comes back, there is a sign saying "No swimming in the Onsen". According to GPT5 it gives me this, which is subtly wrong. > In the novel, when Botchan goes to Dōgo Onsen, he notes the posted rules of the bath. One of them forbids things like: > “No swimming in the bath.” (泳ぐべからず) > “No roughhousing / rowdy behavior.” (無闇に騒ぐべからず) > Botchan finds these signs funny because he’s exactly the sort of hot-headed, restless character who might be tempted to splash around or make noise. He jokes in his narration that it seems as though the rules were written specifically to keep people like him out. Incidentally, Dogo Onsen still has the "No swimming sign", or it did when I went 10 years ago.
- black_knight 1y agoI feel like the value of my plus subscription went down when they released GPT-5, it feels like a downgrade from o3. But of course OpenAI being not open, there is no way for me to know now.
- jbm 1y agoLikewise. I'll play devil's advocate and say that I think the Codex-cli included with the plus subscription is pretty good (quality wise). However, after using it, it suddenly told me I couldn't use it for a week out without warning. Claude is a bit more reasonable there.
- ants_everywhere 1y agoThis isn't really how you described. You have an opinion that conflicts with the research literature. You published a blog about that opinion, and you want ChatGPT to say you're to accept your view. Your view is grinding a political axe and I don't think you're in a position to objectively assess whether ChatGPT failed in this case.
- typpilol 1y agoYea this isn't really a chat gpt problem as a source credibility problem no?
- larsiusprime 1y agoIt’s mostly that it was not citing verifiable - and available online - primary source documents, the way I would expect an actual researcher investigating this question would. This is relevant when it is billed as "Research Grade" or "PhD" level intelligence. I expect a PhD level researcher to find the German-language primary sources.
- eru 1y agoEspecially since ChatGPT speaks fluent German.
- larsiusprime 1y agoWhat are you talking about? There are verifiable primary sources that ChatGPT was not citing. There are direct primary historical sources that lay out the full budget of the historical German colony in extreme detail, that directly contradict assertions made in the Silagi paper, that’s not a matter of opinion that’s a matter of verifiable fact. Also what “axe” am I grinding? The findings are specifically inconvenient for my political beliefs, not confirming my priors! My priors would be flattered if Silagi was correct about everything but the primary sources definitively prove he’s exaggerating. > You published a blog about that opinion, and you want ChatGPT to say you're to accept your view. False, and I address this multiple times in the piece. I don’t want ChatGPT to mindlessly agree with me, I want it to discover the primary source documents.
- simianwords 1y agoI found your article interesting and it is relevant to the discussion. To be honest, while I think GPT could have performed better here, I think there is something to be said about this: There is value in pruning the search tree because the deeper nodes are usually not reputable. I know you have cause to believe that "Wilhelm Matzat" is reputable but I don't think it can be assumed generally. If you were to force GPT to blindly accept counter points from people - the debate would never end. And there has to be a pruning point at which GPT would accept this tradeoff: maybe the less reputable or well known sources may have a correct point at the cost of being incorrect more often due to taking an incorrect analysis from a not well known source. You could go infinitely deep into any analysis and you will always have seemingly correct points on both sides. I think it is valid for GPT to prune the search at a point where it converges to what society at large believes. I'm okay with this tradeoff.
- larsiusprime 1y agoMy contention is if it’s going to just give me a Wikipedia summary, I can do that myself. I just have greater expectations of “PhD” level intelligence. If we’re going to claim to it is PhD level it should be able to do “deep” research AND think critically about source credibility, just as a PhD would. If it can’t do that they shouldn’t brand it that way. Also it’s not like I’m taking Matzat’s word for anything. I can read the primary source documents myself! He’s also hardly an obscure source, he’s just not listed on Wikipedia.
- simonw 1y agoI suggest ignoring the "PhD level intelligence" marketing hype.
- magicalist 1y agoA couple of times when I've gotten an answer sourced basically only from wikipedia and stackoverflow, I've thrown in a comment about its "PhD level intelligence" when I tell it to dig deeper, and it's taken it pretty well ("fair jab :)"), which is amusing. I guess that marketing term has been around long enough to be in gpt5's training data.
- Helmut10001 1y agoMore recently, I find ChatGPT to become increasingly unreliable. It makes up almost every second answer, forgets context, or is just downright wrong. Maybe I am used these days more and more to dump huge texts for context into the prompt, as aistudio allows me. Maybe ChatGPT isn't as good as with such information. Gemini/Aistudio will stay on track even with 300k tokens consumed, it just needs a little nudge here and there.
- herewegohawks 1y agoFWIW, I found things improved greatly once I turned off the memory feature of ChatGPT. My guess is that a lot of tokens were going towards trying to follow instructions from past conversations.
- kmijyiyxfbklao 1y agoThis doesn't tell us much. I don't know why you would expect ChatGPT to do original PhD research. It's a general product that will trust already published research. That doesn't meat that GPT-5 can't do PhD research, when given the right sources.