4 ms·
I think it's appropriate to be extremely critical. The paper is basically useless. The thing that they actually measured is "can GPT-4, when given a 'question'
by mquander 3y ago
I think it's appropriate to be extremely critical. The paper is basically useless. The thing that they actually measured is "can GPT-4, when given a 'question' with lots of additional information and many tries with small permutations to produce an 'answer', at some point produce an 'answer' that GPT-4 will then claim is a 5 out of 5 answer, on a dataset of extremely messy 'questions' and 'answers' from MIT coursework."
That's not an interesting thing to measure. The paper talks about it in terms that make it sound like it's a close proxy for whether GPT-4 "knows" how to do things in MIT coursework, by writing misleadingly about "fulfilling the graduation requirements" and having a "perfect solve rate." But in fact it's totally different. The result is that a bunch of people hear about this paper and get fooled into thinking that there is new interesting evidence about GPT-4's capabilities, unless they manage to read closely enough to see what actually happened.
It's not a matter of whether the results would get weaker if repeated, it's a matter of the results being totally disconnected from any useful real-world information about what GPT-4 can do, or how it can do it.