3 ms·
The link to the "complete list of the experiments" is actually much more than that. It is a description of their methodology, and it's very revealing. >These e
by ppod 6y ago
The link to the "complete list of the experiments" is actually much more than that. It is a description of their methodology, and it's very revealing.
>These experiments are not, by any means, either a representative or a systematic sample of anything. We designed them explicitly to be difficult for current natural language processing technology. Moreover, we pre-tested them on the "AI Dungeon" game which is powered by some version of GPT-3, and we excluded those for which "AI Dungeon" gave reasonable answers. (We did not keep any record of those.) The pre-testing on AI Dungeon is the reason that many of them are in the second person; AI Dungeon prefers that. Also, as noted above, the experiments included some near duplicates. Therefore, though we note that, of the 157 examples below, 71 are successes, 70 are failures and 16 are flawed, these numbers are essentially meaningless.
https://cs.nyu.edu/faculty/davise/papers/GPT3CompleteTests.html https://cs.nyu.edu/faculty/davise/papers/GPT3CompleteTests.h...