4 ms·
Baba Is Solved by Fable 5 and GPT-5.6 Sol, but at what cost?
- stared 3mo agoOpen-source repo with the benchmark (requires a copy of the lovely game): https://github.com/stared/baba-is-harbor https://github.com/stared/baba-is-harbor
- dezgeg 3mo agoGood stuff, I've been trying this myself with Geminis: https://www.youtube.com/watch?v=brkP58pZ23w&list=PL8C_UWcLmvGlv7dWY-oSowUfgFnDOeDg2 https://www.youtube.com/watch?v=brkP58pZ23w&list=PL8C_UWcLmv... Will be interesting to compare the results.
- dezgeg 3mo agoSeems like I have much better results with Gemini <= 3.5 Flash being able to solve 10 Lake levels. Of course there may be methodology differences (like how many times the level is attempted from scratch). I wonder if the reasoning tokens are still being accidentally discarded between turns in those tests. It seems to be much more important to preserve those for Gemini, as it likes to stay quiet and and just make tool calls, unlike Claude that yaps a lot what it's planning to do.