Y
HN Search
Hacker News Search
new
|
comments
|
top
|
jobs
S1M0N38-hn
searching PlanetScale…
1.
▲
2.
▲
3.
▲
4.
▲
5.
▲
6.
▲
3 ms
·
1.
▲
by
S1M0N38-hn
8mo ago
Im pretty open to feedback and contribution (also regarding the default strategy). So feel free to open Issues on GH. However I'd like to collect a bunch of them (including bugs) before re-running the whole benchmark (balatrobench v2).
2.
▲
by
S1M0N38-hn
8mo ago
not really. I've downloaded balatro. I saw that it was moddable. I wrote a mod API to interact programmatically. I was just curious if, from text only game state representation, a LLM would be able to make some decent play. the benchma
3.
▲
by
S1M0N38-hn
8mo ago
Hi, BalatroBench creator here. Yeah, Google models perform well (I guess the long context + world knowledge capabilities). Opus 4.6 looks good on preliminary results (on par with Gemini 3 Pro). I'll add more models and report soon. Tbh
4.
▲
BalatroBench – Benchmarking LLMs' Strategic Performance Through Games
(balatrobench.com)
3 points
by
S1M0N38-hn
8mo ago
|
0 comments