Y
HN Search
Hacker News Search
new
|
comments
|
top
|
jobs
zone411
searching PlanetScale…
1.
▲
2.
▲
3.
▲
4.
▲
5.
▲
6.
▲
6 ms
·
31.
▲
by
zone411
10mo ago
I did some searches when I posted this project, but I didn't find any at the time.
32.
▲
by
zone411
10mo ago
Without monitoring, you can definitely end up with rule-breaking behavior. I ran this experiment: https://github.com/lechmazur/emergent_collusion/ . An agent running like this would break the law. "In a simula
33.
▲
by
zone411
11mo ago
Sets a new record on the Extended NYT Connections: 96.8. Gemini 2.5 Pro scored only 57.6. https://github.com/lechmazur/nyt-connections/
34.
▲
by
zone411
11mo ago
Sets a new record on the Extended NYT Connections benchmark: 96.8 ( https://github.com/lechmazur/nyt-connections/ ). Grok 4 is at 92.1, GPT-5 Pro at 83.9, Claude Opus 4.1 Thinking 16K at 58.8. Gemini 2.5 Pro scored
35.
▲
by
zone411
1y ago
You got many answers already, but a couple more points: Poker doesn't require lying or table talk. Bluffing is rule-legal strategic deception expressed through betting. More like a feint in sports than cheating. If "sitting at a t
36.
▲
by
zone411
1y ago
I've benchmarked it on the Extended NYT Connections ( https://github.com/lechmazur/nyt-connections/ ). It scores 20.0 compared to 10.0 for Haiku 3.5, 19.2 for Sonnet 3.7, 26.6 for Sonnet 4.0, and 46.1 for Sonne
37.
▲
by
zone411
1y ago
Matches Grok 4 at the top of the Extended NYT Connections leaderboard: https://github.com/lechmazur/nyt-connections/
38.
▲
Show HN: LLM Round‑Trip Translation Benchmark
(github.com)
6 points
by
zone411
1y ago
|
0 comments
39.
▲
Show HN: LLM Creative Story‑Writing Benchmark V3
(github.com)
8 points
by
zone411
1y ago
|
0 comments
40.
▲
Show HN: Mapping LLM Style and Range in Flash Fiction
(github.com)
7 points
by
zone411
1y ago
|
0 comments
41.
▲
Pact: Head-to-head negotiation benchmark for LLMs
(github.com)
6 points
by
zone411
1y ago
|
0 comments
42.
▲
by
zone411
1y ago
It is the new leader on my Short Story Creative Writing benchmark: https://github.com/lechmazur/writing/
43.
▲
by
zone411
1y ago
GPT-5 set a new record on my Confabulations on Provided Texts benchmark: https://github.com/lechmazur/confabulations/
44.
▲
by
zone411
1y ago
On the Extended NYT Connections benchmark, GPT-5 Medium Reasoning scores close to o3 Medium Reasoning, and GPT-5 Mini Medium Reasoning scores close to o4-Mini Medium Reasoning: https://github.com/lechmazur/nyt-connectio
45.
▲
by
zone411
1y ago
I benchmarked the 120B version on the Extended NYT Connections (759 questions, https://github.com/lechmazur/nyt-connections ) and on 120B and 20B on Thematic Generalization (810 questions, https://github.com&
46.
▲
Show HN: Bazaar – a new LLM benchmark for economic reasoning under uncertainty
(github.com)
8 points
by
zone411
1y ago
|
1 comments
47.
▲
AI Comes Up with Physics Experiments. But They Work
(quantamagazine.org)
4 points
by
zone411
1y ago
|
0 comments
48.
▲
Emergent Price-Fixing by LLM Auction Agents
(github.com)
7 points
by
zone411
1y ago
|
0 comments
49.
▲
by
zone411
1y ago
The exact questions are almost certainly not in the training data, since extra words are added to each puzzle, and I don't publish these along with the original words (though there's a slight chance they used my previous API reque
50.
▲
by
zone411
1y ago
Grok 4 sets a new high score on my Extended NYT Connections benchmark (92.4), beating o3-pro (87.3): https://github.com/lechmazur/nyt-connections/ . Grok 4 Heavy is not in the API.
51.
▲
by
zone411
1y ago
This is the speculation, but then it wouldn't have to take much longer to answer than o3.
52.
▲
by
zone411
1y ago
It's a non-profit, they can't just replace their management. They've created a new library board charged with determining how to provide library services in the county: https://www.nvdaily.com/nvdaily/war
53.
▲
by
zone411
1y ago
It never went anywhere because of the politicians. The Boring Company is opening new tunnels in Vegas without spending public money.
54.
▲
by
zone411
1y ago
I benchmarked it on four of my benchmarks so far. Got first place in two of them: https://github.com/lechmazur/confabulations https://github.com/lechmazur/nyt-connections https://github
55.
▲
by
zone411
1y ago
Yes, the name should be changed ASAP. You don't want to rename it after it becomes established, which you'll inevitably have to do.
56.
▲
by
zone411
1y ago
Omproves on the Extended NYT Connections benchmark compared to both Gemini 2.5 Pro Exp (03-25) and Gemini 2.5 Pro Preview (05-06), scoring 58.7. The decline observed between 03-25 and 05-06 has been reversed - https://github.com&
57.
▲
by
zone411
1y ago
If anyone is interested in a larger sample size comparing how often LLMs confabulate answers based on provided texts, I have a benchmark at https://github.com/lechmazur/confabulations/ . It's always interestin
58.
▲
by
zone411
1y ago
On my Thematic Generalization Benchmark ( https://github.com/lechmazur/generalization , 810 questions), the Claude 4 models are the new champions.
59.
▲
by
zone411
1y ago
On the extended version of NYT Connections - https://github.com/lechmazur/nyt-connections/ : Claude Opus 4 Thinking 16K: 52.7. Claude Opus 4 No Reasoning: 34.8. Claude Sonnet 4 Thinking 64K: 39.6. Claude Sonnet 4 T
60.
▲
by
zone411
1y ago
Yes, predefined strategies are very interesting to examine. I have two simple ones in another multi-agent benchmark, https://github.com/lechmazur/step_game (SilentGreedyPlayer and SilentRandomPlayer), and it's fas
More ›