Y
HN Search
Hacker News Search
new
|
comments
|
top
|
jobs
CallumFerg
searching PlanetScale…
1.
▲
2.
▲
3.
▲
4.
▲
5.
▲
6.
▲
3 ms
·
1.
▲
by
CallumFerg
4mo ago
Since the library tools are just an MCP server, I did some testing on ChatGPT and Claude where I don't have to pay for api credits. With maximum thinking and web search to look up magic rules, I didn't ever see it make a mistake.
2.
▲
by
CallumFerg
4mo ago
The scoring is just based on a simple prompt which is given the game state at the start and end of the turn and the log of tool calls and the final turn summary. The prompt asks it to evaluate the quality of the simulation from 0 to 10, and
3.
▲
by
CallumFerg
4mo ago
Admittedly, the mulligan phase system prompt is the weakest part of the project. I had to add heuristics to stop the LLMs from mulliganing down to just a few cards looking for a perfect hand. The scoring for the benchmark is mostly based on
4.
▲
by
CallumFerg
4mo ago
No, I was not aware of that project when I made this. I'll have to look into that project, but I also have an RTX 5090 and did a lot of testing with Qwen3.6 27B and Gemma 4 31B. I was not able to get it to play legal turns consistently
5.
▲
by
CallumFerg
4mo ago
I actually considered using card forge when I started this. I mostly didn't end up using it because of how much more work it would have been. But also with a rules engine, you have to manually go though every step, and pass priority af
6.
▲
MTG Bench: Testing how well LLMs can play Magic
(mtgautodeck.com)
68 points
by
CallumFerg
4mo ago
|
34 comments