3 ms·
It significantly improves upon GPT-4o on my Extended NYT Connections Benchmark. 22.4 -> 33.7 (https://github.com/lechmazur/nyt-connections https://github.com/le
by zone411 2y ago
It significantly improves upon GPT-4o on my Extended NYT Connections Benchmark. 22.4 -> 33.7 (https://github.com/lechmazur/nyt-connections https://github.com/lechmazur/nyt-connections).
- zone411 2y agoI ran three more of my independent benchmarks: - Improves upon GPT-4o's score on the Short Story Creative Writing Benchmark, but Claude Sonnets and DeepSeek R1 score higher. (https://github.com/lechmazur/writing/ https://github.com/lechmazur/writing/) - Improves upon GPT-4o's score on the Confabulations/Hallucinations on Provided Documents Benchmark, nearly matching Gemini 1.5 Pro (Sept) as the best-performing non-reasoning model. (https://github.com/lechmazur/confabulations https://github.com/lechmazur/confabulations) - Improves upon GPT-4o's score on the Thematic Generalization Benchmark, however, it doesn't match the scores of Claude 3.7 Sonnet or Gemini 2.0 Pro Exp. (https://github.com/lechmazur/generalization https://github.com/lechmazur/generalization) I should have the results from the multi-agent collaboration, strategy, and deception benchmarks within a couple of days. (https://github.com/lechmazur/elimination_game/ https://github.com/lechmazur/elimination_game/, https://github.com/lechmazur/step_game https://github.com/lechmazur/step_game and https://github.com/lechmazur/goods https://github.com/lechmazur/goods).
- j_bum 2y agoHonest question for you: are these puzzles actually a good way to test the models? The answers are certainly in the training set, likely many times over. I’d be curious to see performance on Bracket City, which was featured here on HN yesterday.