3 ms·
It ranks between Mistral Small and Mistral Medium on my NYT Connections benchmark and is indeed better than Command R Plus and Qwen 1.5 Chat 72B, which were the
by zone411 2y ago
It ranks between Mistral Small and Mistral Medium on my NYT Connections benchmark and is indeed better than Command R Plus and Qwen 1.5 Chat 72B, which were the top two open weights models. Grok 1.0 is not an instruct model, so it cannot be compared fairly.
- J_Shelby_J 2y agoCan you share the details about the benchmark?
- zone411 2y agoUses an archive of 267 NYT Connections puzzles (try them yourself if unfamiliar). Three different 0-shot prompts, words in both lowercase and uppercase. One attempt per puzzle. Partial credit is awarded if not all lines are solved correctly. Top humans get near 100. Most other benchmarks don't clearly show the difference between the top models and the rest. This may be because they are older and have been over-optimized or perhaps because they are just easier.