4 ms·
Those numbers are arbitrary and fictional, and the more relevant made-up quantity would be the variance rather than the mean. It doesn't really matter if the "a
by nicklecompte 2y ago
Those numbers are arbitrary and fictional, and the more relevant made-up quantity would be the variance rather than the mean. It doesn't really matter if the "average user" saves time over 10,000 queries. I am much more concerned about the numerous edge cases, especially if those cases might be "edge fields" like animal cognition (see below).
In my experience it takes quite a bit longer to falsify GPT-4's incorrect answers than it does to a Google search and get the right answer. It might take 30 seconds to check a correct answer (jump to the relevant paragraph and check), but 30 minutes to determine where an incorrect answer actually went wrong (you have to read the whole paper in close detail, and maybe even relevant citations). More specifically, it is somewhat quick to falsify something if it is directly contradicted by the text. It is much harder to falsify unsupported generalizations or summaries.
As a specific example, I recently asked GPT for information on arithmetic abilities in amphibians. It made up a study - that was easy to check - but it also made up a bunch of results without citing specific studies. That was not easy to check[1]: each paragraph of text GPT generated needed to be cross-checked with Google Scholar to try and find a relevant paper. It turned out that everything GPT said, over 1000 words of output, was contradicted by actual published research. But I had to read three papers to figure that out. I would have been much better off with Google Scholar. But I am concerned that a large minority of cynical, lazy people will say "90% is good enough, I don't want to read all these papers and nobody's gonna check the citations anyway" and further drag down the reliability of published research.
[1] This was a test of GPT. If I were actually using it for work, obviously I would have stopped at the fake citation.
- throwup238 2y agoThis is impacting online discourse too. It used to be that when someone is wildly wrong it was relatively easy to identify why: ideology, common urban myth, outdated research, whatever. Now? I’ve seen people argue positions that are demonstrably wildly wrong in unusually creative and often subtle ways and there’s no way to figure out where they went off the rails. Since the LLM is responsive, they can use it to come up plausibly sounding nonsense to answer any criticisms collapsing the debate into a black hole of bullshit.
- jamesbrady 2y agoElician here! Thanks for your comment. I'm not sure I agree that those rule-of-thumb statistics are "arbitrary" or "fictional"… I guess it depends on what you mean by that. I can say that on our part they're a good faith attempt to help users calibrate how best to use the tool, using evaluations of Elicit based on real usage. Definitely accept that the tool can work better or worse depending on your domain or workflow though! One way we do try to distinguish ourselves from vanilla LLMs is that we provide sources for all of the claims made. I mention this because we hope our users can approach the falsification process you mention for Google. We want to show people where particular claims come from such that we earn their trust. Walking citation trails and verifying transitive claims is something we've talked about but need more people to implement! (https://elicit.com/careers https://elicit.com/careers)
- nicklecompte 2y ago> I'm not sure I agree that those rule-of-thumb statistics are "arbitrary" or "fictional"… I guess it depends on what you mean by that. Sorry for the confusion: I meant that fragmede's comment was arbitrary and fictional, not the 90% figure. I was talking about these numbers: if it takes 1 hour to get one answer by hand, but only 20 minutes for the machine, and 20 minutes to check the answer, the user still comes out ahead
- jamesbrady 2y agoOh, my bad—I misunderstood, thanks for the clarification