Y
HN Search
Hacker News Search
new
|
comments
|
top
|
jobs
zone411
searching PlanetScale…
1.
▲
2.
▲
3.
▲
4.
▲
5.
▲
6.
▲
8 ms
·
61.
▲
by
zone411
2y ago
> Could you run some analysis on how often “p1” wins vs “p8”? I checked the average finishing positions by assigned seat number from the start, but there weren't enough games to show a statistically significant effect. But I just re
62.
▲
by
zone411
2y ago
Very cool!
63.
▲
by
zone411
2y ago
Author here - some weaker LLMs actually have trouble tracking the game state. The fun part is when smarter LLMs realize they're confused! Claude 3.7 Sonnet: "Hey P5! I think you're confused - P3 is already eliminated." C
64.
▲
by
zone411
2y ago
Author here - it's based on finishing positions (so it's not winner-take-all) and then TrueSkill by Microsoft ( https://trueskill.org/ ). It's basically a multiplayer version of Elo that's used in chess an
65.
▲
by
zone411
2y ago
Author here - yes, I'm regularly adding new models to this and other TrueSkill-based benchmarks and it works well. One thing to keep in mind is the need to run multiple passes of TrueSkill with randomly ordered games, because both True
66.
▲
by
zone411
2y ago
Author here - I'm planning to create game versions of this benchmark, as well as my other multi-agent benchmarks ( https://github.com/lechmazur/step_game , https://github.com/lechmazur/pgg_bench
67.
▲
by
zone411
2y ago
It's interesting that there are no reasoning models yet, 2.5 months after DeepSeek R1. It definitely looks like R1 surprised them. The released benchmarks look good. Large context windows will definitely be the trend in upcoming model
68.
▲
by
zone411
2y ago
Scores 54.1 on the Extended NYT Connections Benchmark, a large improvement over Gemini 2.0 Flash Thinking Experimental 01-21 (23.1). 1 o1-pro (medium reasoning) 82.3 2 o1 (medium reasoning) 70.8 3 o3-mini-high 61.4 4 Gemini 2.5 Pro Exp 03-2
69.
▲
Public Goods Game Benchmark: Contribute and Punish, a Multi-Agent Benchmark
(github.com)
7 points
by
zone411
2y ago
|
0 comments
70.
▲
by
zone411
2y ago
I ran three more of my independent benchmarks: - Improves upon GPT-4o's score on the Short Story Creative Writing Benchmark, but Claude Sonnets and DeepSeek R1 score higher. ( https://github.com/lechmazur/writing&#x
71.
▲
by
zone411
2y ago
It significantly improves upon GPT-4o on my Extended NYT Connections Benchmark. 22.4 -> 33.7 ( https://github.com/lechmazur/nyt-connections ).
72.
▲
Elimination Game: Multi-Agent LLM Social Reasoning, Strategy, and Deception
(github.com)
5 points
by
zone411
2y ago
|
0 comments
73.
▲
by
zone411
2y ago
Claude 3.7 Sonnet Thinking scores 33.5 (4th place after o1, o3-mini, and DeepSeek R1) on my Extended NYT Connections benchmark. Claude 3.7 Sonnet scores 18.9. I'll run my other benchmarks in the upcoming days. https://github
74.
▲
by
zone411
2y ago
Also https://news.ycombinator.com/item?id=43086347
75.
▲
SWE-Lancer: a benchmark of freelance software engineering tasks from Upwork
(arxiv.org)
111 points
by
zone411
2y ago
|
74 comments
76.
▲
by
zone411
2y ago
Apparently the API will only be available in a few weeks, so I can't run my independent benchmarks yet.
77.
▲
by
zone411
2y ago
You meant utilization not capacity. Your other points are similarly confused. Argentinians disagree. "The percentage of Argentines who say their standard of living is getting better (53%) has inched above the majority level for the fir
78.
▲
by
zone411
2y ago
Not ideal, but the reason for this is that people have gotten used to larger bars indicating better performance on bar charts when evaluating LLMs. Including being confused by the older version of this very benchmark. Two bar charts are als
79.
▲
LLM Hallucination Benchmark: R1, o1, o3-mini, Gemini 2.0 Flash Think Exp 01-21
(github.com)
17 points
by
zone411
2y ago
|
3 comments
80.
▲
by
zone411
2y ago
I have a set of independent benchmarks and most also show a difference between reasoning and non-reasoning models: LLM Confabulation (Hallucination): https://github.com/lechmazur/confabulations/ LLM Step Game: ht
81.
▲
by
zone411
2y ago
It scores 72.4 on NYT Connections, a significant improvement over the o1-mini (42.2) and surpassing DeepSeek R1 (54.4), but it falls short of the o1 (90.7). ( https://github.com/lechmazur/nyt-connections/ )
82.
▲
by
zone411
2y ago
The problem is confabulations. In my benchmark ( https://github.com/lechmazur/confabulations/ ), you see models produce non-existent answers in response to misleading questions that are based on provided text docume
83.
▲
by
zone411
2y ago
I just ran my NYT Connections benchmark on it: 18.6, up from 14.8 for Qwen 2.5 72B. I'll run my other benchmarks later. https://github.com/lechmazur/nyt-connections/
84.
▲
Multi-Agent Step Race Benchmark: LLM Collaboration and Deception Under Pressure
(github.com)
7 points
by
zone411
2y ago
|
1 comments
85.
▲
by
zone411
2y ago
Playing with MuJoCo is a lot of fun. I recommend it for people who just want to experiment with RL. You can do random weird stuff like this: https://www.youtube.com/watch?v=e-9u_vepfY8 . I'm really glad DeepMind is cont
86.
▲
Show HN: LLM Thematic Generalization Benchmark
(github.com)
6 points
by
zone411
2y ago
|
0 comments
87.
▲
Show HN: LLM Creative Story-Writing Benchmark
(github.com)
5 points
by
zone411
2y ago
|
0 comments
88.
▲
Show HN: LLM Divergent Thinking Creativity Benchmark
(github.com)
8 points
by
zone411
2y ago
|
0 comments
89.
▲
by
zone411
2y ago
It's the least interesting benchmark for language models among all they've released, especially now that we already had a large jump in its best scores this year. It might be more useful as a multimodal reasoning task since it cle
90.
▲
by
zone411
2y ago
In my NYT Connections benchmark, it hasn't performed well: https://github.com/lechmazur/nyt-connections/ (see the table).
More ›