4 ms·
While it might be good for a non-technical essay, for technical matters, it has a bad habit of spewing nonsense, both in answers and citations. Professor Alex W
by nukeman 4y ago
While it might be good for a non-technical essay, for technical matters, it has a bad habit of spewing nonsense, both in answers and citations. Professor Alex Wellerstein (of Nukemap fame) gives two anecdotes highlighting their issues.
1. https://old.reddit.com/r/AskHistorians/comments/11u21ie/the_consensus_from_a_brief_search_of_previous/jcn3aee/ https://old.reddit.com/r/AskHistorians/comments/11u21ie/the_...
An anecdote, but I recently was asked to review the essay of a student who I had not taught. I became highly suspicious it had been generated by ChatGPT, because it had the "feel" of its output. The clincher was that it had an entire page of references... all of which were fake. They all looked plausible, and even had URLs. But not one of them was accurate, and all of the URLs were dead, and all investigation made it clear there references had never existed. I was somewhat amazed, both at the gall of a chatbot inventing fake references, and for the student who clearly did not click on even one of the generated links, yet had still asked for an essay re-grade!!
2. https://old.reddit.com/r/AskHistorians/comments/11u21ie/the_consensus_from_a_brief_search_of_previous/jcn3w2q/ https://old.reddit.com/r/AskHistorians/comments/11u21ie/the_...
One experiment I ran with it recently was to ask it about the RIPPLE, which is a nuclear weapon design that was tested in the 1960s. The details of the RIPPLE are not public, but the fact of its existence, who invented it, and its testing are, as well as the some very broad pieces of information about it. Anyway, I repeatedly asked ChatGPT how the RIPPLE worked, and why it was called the RIPPLE, and every time it gave me a totally new and contradictory answer, freely making it up each time. After giving me maybe 6 different answers in a row it then noticed it was giving me contradictions, and from that point onward claimed that the most recent answer was correct. I was impressed at how inconsistent it was, that you could just ask it the same thing over and over again and it would just make new things up each time. The only consistency it gave me was wrong: it repeatedly emphasized that the design was entire hypothetical and never tested, which is false (it was tested at least four times).
In a separate exchange, I asked it to ask me a question, and when (for whatever reason) I told it I was interested in nuclear weapons, it began to lecture me on how this was a topic that should be left to experts. I then told it I was an expert, and it then started lecturing me on how an expert on this topic ought to behave and think. It almost seemed defensive. I thought it was pretty rich — an impressive mansplaining simulator, indeed.
3. The full discussion outlined in (2): https://old.reddit.com/r/nuclearweapons/comments/117hssn/chatgpt_makes_up_shit_about_ripple/ https://old.reddit.com/r/nuclearweapons/comments/117hssn/cha...
- tuatoru 4y agoI like the characterization of chatGPT in this comment in the second link: > [–]righthandofdog 46 points 2 days ago Mansplaining as a service is the best description of GPT. > The reason it CAN be right about more general info is because people trained it away from lies. No one has trained out the lies on more rarified knowledge, so it makes shit up. > Trusting its answers to be correct is flatly stupid when it is literally designed to make shit up that sounds good instead of saying "I don't know".
- nukeman 4y agoAnd the reality is “I don’t know” is a valid response! But it isn’t satisfying, I remember five year-old me would get very upset with my mom if she said something like that, unless it was followed up with “but let’s go look it up together.” Unfortunately for OpenAI, that isn’t the kind of response their customers probably want to see; they would likely get annoyed and go back to Google or DDG or something else.
- nullc 4y ago> that you could just ask it the same thing over and over again and it would just make new things up each time Why is that described as bad or surprising? It doesn't know and so these varrious answers seem about equally likely to it.
- doctor_eval 4y agoI always think this is a bit like some kind of interrogation where the subject is trying to give the desired answer. If you are compelled to give and answer - any answer - then that’s what you’re gonna do.
- nukeman 4y agoI suspect the fact that it’s mostly all over the place, it doesn’t narrow toward the correct answer. Here’s his full test: https://old.reddit.com/r/nuclearweapons/comments/117hssn/chatgpt_makes_up_shit_about_ripple/ https://old.reddit.com/r/nuclearweapons/comments/117hssn/cha...
- 4y ago