7 ms·
Gemini 3 Pro vs. 2.5 Pro in Pokemon Crystal
- jwrallie 9mo agoBeing through the game recently, I am not surprised Goldenrod Underground was a challenge, it is very confusing and even though I solved it through trial and error, I still don't know what I did. Olivine Lighthouse is the real surprise, as it felt quite obvious to me.
- MrCheeze 9mo agoThis writeup on the underground puzzle is worth reading, it's a pretty baffling "puzzle" design. https://pokemow.com/Gen2/ShutterPuzzle/ https://pokemow.com/Gen2/ShutterPuzzle/ That said, it's definitely Gem's fault that it struggled so long, considering it ignored the NPCs that give clues.
- wild_pointer 9mo agoI wonder how much of it is due to the model being familiar with the game or parts of it, be it due to training of the game itself, or reading/watching walkthroughs online.
- andrepd 9mo agoThere was a well-publicised "Claude plays Pokémon" stream where Claude failed to complete Pokemon Blue in spectacular fashion, despite weeks of trying. I think only a very gullible person would assume that future LLMs didn't specifically bake this into their training, as they do for popular benchmarks or for penguins riding a bike.
- criley2 9mo agoWhile it is true that model makers are increasingly trying to game benchmarks, it's also true that benchmark-chasing is lowering model quality. GPT 5, 5.1 and 5.2 have been nearly universally panned by almost every class of user, despite being a benchmark monster. In fact, the more OpenAI tries to benchmark-max, the worse their models seem to get.
- astrange 9mo agoHm? 5.1 Thinking is much better than 4o or o3. Just don't use the instant model.
- deleted 9mo ago[deleted]
- malnourish 9mo ago5.2 is a solid model and I'm actually impressed with M365 copilot when using it.
- ctoth 9mo ago> as they do for popular benchmarks or for penguins riding a bike. Citation?
- deleted 9mo ago[deleted]
- dwaltrip 9mo agoIf they game the pelican benchmark, it’d be pretty obvious. Just try other random, non-realistic things like “a giraffe walking a tightrope”, “a car sitting at a cafe eating a pizza”, etc. If the results are dramatically different, then they gamed it. If they are similar in quality, then they probably didn’t.
- deleted 9mo ago[deleted]
- oceansky 9mo ago"Crucially, it tells the agent not to rely on its internal training data (which might be hallucinated or refer to a different version of the game) but to ground its knowledge in what it observes. " Does this even have any effect?
- blibble 9mo agoI very much doubt it
- tootyskooty 9mo agoI'm wondering about this too. Would be nice to see an ablation here, or at least see some analysis on the reasoning traces. It definitely doesn't wipe its internal knowledge of Crystal clean (that's not how LLMs work). My guess is that it slightly encourages the model to explore more and second-guess it's likely very-strong Crystal game knowledge but that's about it.
- Workaccount2 9mo agoThe model probably recognizes the need for a grassroots effort to solve the problem, to "show it's work".
- ragibson 9mo agoYes, at least to some extent. The author mentions that the base model knows the answer to the switch puzzle but does not execute it properly here. "It is worth noting that the instruction to "ignore internal knowledge" played a role here. In cases like the shutters puzzle, the model did seem to suppress its training data. I verified this by chatting with the model separately on AI Studio; when asked directly multiple times, it gave the correct solution significantly more often than not. This suggests that the system prompt can indeed mask pre-trained knowledge to facilitate genuine discovery."
- hypron 9mo agoMy issue with this is that the LLM could just be roleplaying that it doesn't know.
- soulofmischief 9mo agoNice writeup! I need to start blogging about my antics. I rigged up several cutting edge small local models to an emulator all in-browser and unsuccessfully tried to get them to play different Pokémon games. They just weren't as sharp as the frontier models. This was a good while back but I'm sure a lot of people might find the process and code interesting even if it didn't succeed. Might resurrect that project.
- giancarlostoro 9mo agoI have to think they need to know enough of the guides for the game for it to work out, how do they know whats on screen?
- soulofmischief 9mo agoIn my project I rigged up an in-browser emulator and directly fed captured images of the screen to local multimodal models. So it just looks right at what's going on, writes a description for refinement, and uses all of that to create and manage goals, write to a scratchpad and submit input. It's minimal scaffolding because I wanted to see what these raw models are capable of. Kind of a benchmark.
- giancarlostoro 9mo agoI have a feeling if you gave them access to GameFAQ guides they might be able to play better, but it depends on how you can feed them the data.
- soulofmischief 9mo agoIt turns out that cutting edge super small (3b param etc) models that fit in the browser are not great at playing Pokémon on an even basic level, even navigation is difficult when only providing raw visual information, and object recognition of the low-resolution sprites is not great. So I lost interest before even getting to the point of providing specific strategy. But, it runs in browser and works with any supplied ROM, none of it is Pokémon-specific so I should set aside time to serve it and make the code available
- bbondo 9mo ago1.88 billion tokens * $12 / 1M tokens (output) suggests a total cost of $22,560 to solve the game with Gemini 3 Pro?
- brianwawok 9mo agoTrue though I bet the $200 a month plan could do it, maybe a few extra days of downtime when quota was maxed
- AstroBen 9mo agoFor how long would it stay $200 of you can rack up 5 figures if usage..
- manmal 9mo agoThat is the reason they severely limited Claude Max subscriptions. Some users racked up 1k+ in API equivalent cost per day.
- jchw 9mo agoThis is exactly why I upgrade to the Pixel 10 Pro. On Black Friday, you could get a Pixel 10 Pro for about $450 on the U.S. Google Fi store (which sells unlocked phones)... which is also about how much a Pixel 9 Pro goes for on eBay; minus eBay fees and accounting for shipping, that's an upgrade for < $100. But, it's even better, the Pixel 10 Pro comes with a year of their "AI Pro" plan (which I believe costs around $240/year.) There is really, really no point in upgrading to a Pixel 10 Pro from a Pixel 9 Pro, and environmentally it pains me to be the person upgrading my phone on an annual basis (this is the fastest I've ever upgraded a phone, ever) but it's hard to turn down when Google is selling $800~ish for $400~ish. And yeah, it's not the insanely priced AI Ultra plan, but if there are any hard limits on Gemini Pro usage I haven't found them. I have played a lot with really long Antigravity sessions to try to figure out what this thing is good for, and it seems like it will pretty much sit there and run all day. (And I can't really blame anyone for still remaining mad about AI to be completely honest, but the technology is too neat by this point to just completely ignore it.) Seeing as Google is still giving away a bunch of free access, I'm guessing they're still in the ultra-cash-burning phase of things. My hope (hopium, realistically) is that by the time all of the cash burning is over, there will be open-weight local models that are striking near where Gemini 3 Pro strikes today. It doesn't have to be as good, getting nearby on hardware consumers can afford would be awesome. But I'm not holding my breath, so let's hope the cash burning continues for a few years. (There is, of course, the other way to look at it, which is that looking at the pricing per token may not tell the whole story. Given that Google is running their own data centers, it's possible the economic proposition isn't as bad as it looks. OTOH, it's also possible it is worse than it looks, if they happen to be selling tokens at a loss... but I quite doubt it, given they are currently SOTA and can charge a premium.)
- squimmy26 9mo agoHow certain can we be that these improvements aren't just a result of Gemini 3 Pro pre-training on endless internet writeups of where 2.5 has struggled (and almost certainly what a human would have done instead)? In other words, how much of this improvement is true generalization vs memorization?
- zurfer 9mo agoYou're too kind. Even the CEO of Google retweeted how well Gemini 2.5 did on Pokemon. There is a high chance that now it's explicitly part of the training regime. We kind of need a different kind of game to know how well it generalizes.
- kqr 9mo agoI have a draft doing this with text adventures: https://entropicthoughts.com/updated-llm-benchmark https://entropicthoughts.com/updated-llm-benchmark
- prmoustache 9mo agoIsn't that the point of a new model anyway?
- DANmode 9mo agoYes. Sort of. Just don’t confuse it with a random benchmark!
- MrCheeze 9mo agoThere were no such writeups, 99% of the discussion about difficulties in Crystal were in twitch and discord chats where Google doesn't scrape. (It hadn't yet gotten the public attention that Claude and Gemini's runs of Pokemon Red and Blue have gotten.) That said, this writeup itself will probably be scraped and influence Gemini 4.
- cg5280 9mo agoI like the inclusion of the graph at the end to compare progress. It would be cool to compare this directly to competing models (Claude, GPT, etc).
- kqr 9mo agoIt would unfortunately also need several runs of each to be reliable. There's nothing in TFA to indicate the results shown aren't to a large degree affected by random chance! (I do think from personal benchmarks that Gemini 3 is better for the reasons stated by the author, but a single run from each is not strong evidence.)
- sussmannbaka 9mo agoSo after years of being gleefully told that AI will replace all jobs an omniscient state of the art model, with heavy assistance, takes more than two weeks and thousands of dollars in tokens to do what child me did in a few days? Huh.
- murukesh_s 9mo agoI used to think the same until latest agents started adding perfectly fine features to a large existing react app with just basic input (in English) . Most of the jobs require levels of intelligence below that. It's just a matter of time before agents get to that.
- blauditore 9mo agoIt's about the complexity of the task. Front end apps tend do be much less complex and boilerplate-y than backends, hence AI tends to work better.
- etse 9mo agoIsn’t frontend more complex? If my task starts with a Figma UI design, how well does a code agent do at generating working code that looks right, and iterate on it (presuming some browser MCP)? Some automated tests seem enough for an genetic loop on backend.
- murukesh_s 9mo ago>Isn’t frontend more complex? If my task starts with a Figma UI design, how well does a code agent do at generating working code that looks right, and iterate on it (presuming some browser MCP)? Some automated tests seem enough for an genetic loop on backend. Haven't tried a Figma design, but i built an internal tool entirely via instructions to agent. The kind of work I could easily quote 3 weeks previously.
- murukesh_s 9mo agoI disagree - having worked on backends most of the time, I find modern frontend much more complex (and difficult to test) than pure backend. When I say modern frontend - its mostly React, state management like Redux, Zustand, Router framework like React Router, a CSS framework like Tailwind and component framework like Shadcn. Not to mention different versions of React, different ways of managing state, animation/transitions etc. And on top of that the ever increasing complex quirks in the codebase still needed to be compatible with all the modern browsers and device sizes/orientation out there.
- elif 9mo agoGive it the gameFAQ next time
- orbital-decay 9mo agoThe baked-in assumptions observation is basically the opposite of the impression I get after watching Gemini 3's CoT. With the maximum reasoning effort it's able to break out of the wrong route by rethinking the strategy. For example I gave it an onion address without the .onion part, and told it to figure out what this string means. All reasoning models including Gemini 2.5 and 3 assume it's a puzzle or a cipher (because they're trained on those) and start endlessly applying different algorithms to no avail. Gemini 3 Pro is the only model that can break the initial assumption after running out of ideas ("Wait, the user said it's just a string, what if it's NOT obfuscated"), and correctly identify the string as an onion address. My guess is they trained it on simulations to enforce the anti-jailbreaking commands injected by the Model Armor, as its CoT is incredibly paranoid at times. I could be wrong, of course.
- jug 9mo agoI've had some weird "thinking outside the box" behavior like this. I once asked 3 Pro what Ozzy Osbourne is up to. The CoT was a journey, I can tell you! It's not in its training data that he actually passed away. It did know he was planning a tour though. It had a real struggle trying to consolidate "suspicious search results" and even questioned whether it was fake news, or running against a simulation!, determining it wasn't going to fall for my "test". It did ultimately decide Ozzy was alive. I pushed back on that, and it instantly corrected itself and partially blamed my query "what is he up to" for being formulated as if he was alive.
- Wowfunhappy 9mo agoOdd, mine didn't do anything interesting.
- reilly3000 9mo agoI’d love to see how the new flash-3 model would fare.
- dash2 9mo ago> it often makes early assumptions and fails to validate them, which can waste a lot of time Is this baked into how the models are built? A model outputs a bunch of tokens, then reads them back and treats them as the existing "state" which has to be built on. So if the model has earlier said (or acted like) a given assumption is true, then it is going to assume "oh, I said that, it must be the case". Presumably one reason that hacks like "Wait..." exist is to work around this problem.
- topaz0 9mo agoWho do I have to talk to to get somebody to pay me thousands of dollars to beat a game from the 90s?
- krige 9mo agoAs a fun comparison, Gemini 3 Pro took 17 days to beat the game. Twitch Plays Pokemon, which was frequently random, chaotic, even malicious, took 13 days to clear Crystal.
- dpedu 9mo agoIs the code behind this available?