3 ms·
Confidence intervals too wide to draw that conclusion.
by hackerlight 2y ago
Confidence intervals too wide to draw that conclusion.
- oersted 2y agoNot really, right now it's within +14/-15 ELO with 95% probability, which barely changes the ranking. It's well ahead of Gemini Pro and squarely within the top 5 with Claude 3 Opus and a few versions of GPT-4-Turbo. More impressively even, Llama 3 8B is approximately tied with GPT-4 (depending on the version), as well as Mistral-Large and Mixtral 8x22B, which is mad for the size.