3 ms·
Slightly better on the NYT Connections benchmark (27.9) than Claude 3 Opus (27.3) but massively improved over Claude 3 Sonnet (7.8). GPT-4o 30.7 Claude 3.5 So
by zone411 2y ago
Slightly better on the NYT Connections benchmark (27.9) than Claude 3 Opus (27.3) but massively improved over Claude 3 Sonnet (7.8).
GPT-4o 30.7
Claude 3.5 Sonnet 27.9
Claude 3 Opus 27.3
Llama 3 Instruct 70B 24.0
Gemini Pro 1.5 0514 22.3
Mistral Large 17.7
Qwen 2 Instruct 72B 15.6
- zsmizzle 2y agoIt still fails to be the moderator of a WORDLE board. That is always the first test I do of these new models.