4 ms·
I've just created a new benchmark to see how top LLMs do on NYT Connections (https://www.nytimes.com/games/connections https://www.nytimes.com/games/connections
by zone411 3y ago
I've just created a new benchmark to see how top LLMs do on NYT Connections (https://www.nytimes.com/games/connections https://www.nytimes.com/games/connections). 267 puzzles, 3 prompts for each, uppercase and lowercase.
GPT-4 Turbo: 31.0
Claude 3 Opus: 27.3
Mistral Large: 17.7
Mistral Medium: 15.3
Gemini Pro: 14.2
Qwen 1.5 72B Chat: 10.7
Claude 3 Sonnet: 7.6
GPT-3.5 Turbo: 4.2
Mixtral 8x7B Instruct: 4.2
Llama 2 70B Chat: 3.5
Qwen 1.5 14B: 3.1
Nous Hermes 2 Yi 34B: 1.5
Notes: 0-shot. Maximum possible is 100. Partial credit is given if the puzzle is not fully solved. There is only one attempt allowed per puzzle. In contrast, humans players get 4 attempts and a hint when they are one step away from solving a group. Gemini Advanced is not yet available through the API.
What I found interesting is how this benchmark reveals a large capabilities gap between the top, large models and the rest, in contrast to existing over-optimized benchmarks.
- d--b 3y agoAlso these puzzles can be _really_ hard. As a French person who's lived 10+ years in English-speaking countries, I am often completely baffled. I am not sure humans would do a lot better with 0-shot.
- Vecr 3y agoIt's probably somewhat g loaded. I don't know how much, but someone could look at the curves (if they have access?) for the similar sub-section of an IQ test.
- mike986 3y agoDo you know the average or top human score / SD? Not sure if that data is available on the link or elsewhere. Just to make sense of your result, can you show your prompt? When you say 3 prompts and one attempt, what does that mean? Also regarding 0-shot, did you give the LLM the instruction that is given to human by the game? If yes, I would count that as one shot as an example of how to properly solve one example puzzle is given. ``` How to Play Find groups of four items that share something in common. Select four items and tap 'Submit' to check if your guess is correct. Find the groups without making 4 mistakes! Category Examples FISH: Bass, Flounder, Salmon, Trout FIRE ___: Ant, Drill, Island, Opal Categories will always be more specific than "5-LETTER-WORDS," "NAMES" or "VERBS." Each puzzle has exactly one solution. Watch out for words that seem to belong to multiple categories! Each group is assigned a color, which will be revealed as you solve ``` Thanks
- QuantumGood 3y agoCan you break out the partial credit score from the "solved" score?