4 ms·
Most of the GPT-4 benchmarks from their report were things like AP tests or leer code scores. Which aren’t benchmarks that can be compared by a different set of
by tempusalaria 3y ago
Most of the GPT-4 benchmarks from their report were things like AP tests or leer code scores. Which aren’t benchmarks that can be compared by a different set of researchers as you don’t know the constituent parts of the test to run
- pama 3y agoTable 2 of the OpenAI report had 7 public benchmarks and figure 5 had another 27.
- YetAnotherNick 3y agoGPT-4 report has MMLU score, which is believed to one of the most important metric for question answering task. GPT-4 MMLU score is slightly higher than PaLM 2(86 vs 81). Google didn't compare it in with PaLM 2 in this paper.
- in3d 3y ago86 vs 81 is not slightly higher. It’s 26% fewer wrong answer.