4 ms·
I've run some evals on my puzzle game https://redactle.net/llm-leaderboard https://redactle.net/llm-leaderboard Deepseek v4.1 flash is able to solve it some of
by pampas 16d ago
I've run some evals on my puzzle game https://redactle.net/llm-leaderboard https://redactle.net/llm-leaderboard
Deepseek v4.1 flash is able to solve it some of the time. I've found it burns through more reasoning tokens than any other model. Google models like Gemini 3.8 Flash are still dominating and is able to one-shot most evals while being the cheapest.
I'm curious what other unique evals people are running.
- mordae 16d agoSince it has low activated parameter count but huge total parameter count it needs more tokens to move the relevant information into the context.
- pampas 15d agoThanks for the help. I ran it on high and it did pretty well and got a lot of one-shots in. The reasoning makes a much bigger difference than some other models.
- cbg0 16d agoDoes it move the needle on high reasoning?
- pampas 15d agoYes. I've just run it on high and it did a lot better.
- gandreani 16d agoIt's so bizarre having a low score be GOOD. It's like reverse intuition. Shouldn't it be called `score error` or something along those lines?
- pampas 15d agoGreat point. I've changed the naming.
- pimeys 16d agoA colleague of mine has a strategy game to compare language models, 4.1 scores pretty high in this: https://clankerbattle.com/ https://clankerbattle.com/