13 ms·
I play competitive Geoguessr at a fairly high level, and I wanted to test this out to see how it compares. It's astonishingly good. It will use information it
by SamPatt 1y ago
I play competitive Geoguessr at a fairly high level, and I wanted to test this out to see how it compares.
It's astonishingly good.
It will use information it knows about you to arrive at the answer - it gave me the exact trailhead of a photo I took locally, and when I asked it how, it mentioned that it knows I live nearby.
However, I've given it vacation photos from ages ago, and not only in tourist destinations either. It got them all as good or better than a pro human player would. Various European, Central American, and US locations.
The process for how it arrives at the conclusion is somewhat similar to humans. It looks at vegetation, terrain, architecture, road infrastructure, signage, and it just knows seemingly everything about all of them.
Humans can do this too, but it takes many thousands of games or serious study, and the results won't be as broad. I have a flashcard deck with hundreds of entries to help me remember road lines, power poles, bollards, architecture, license plates, etc. These models have more than an individual mind could conceivably memorize.
- simonw 1y agoIs that flashcard deck a commercial/community project or is it something you assembled yourself? Sounds fascinating!
- SamPatt 1y agoI made it myself. I use Obsidian and the Spaced Repetition plugin, which I highly recommend if you want a super simple markdown format for flashcards and use Obsidian: https://www.stephenmwangi.com/obsidian-spaced-repetition/ https://www.stephenmwangi.com/obsidian-spaced-repetition/ There are pre-made Geoguessr decks for Anki. However, I wouldn't recommend using them. In my experience, a fundamental part of spaced repetition's efficacy is in creating the flashcards yourself. For example I have a random location flashcard section where I will screenshot a location which is very unique looking, and I missed in game. When I later review my deck I'm way more likely to properly recall it because I remember the context of making the card. And when that location shows up in game, I will 100% remember it, which has won me several games. If there's interest I can write a post about this.
- dr_dshiv 1y agoI’m interested from a learning science perspective. It’s a nice finding even if anecdotal
- simonw 1y agoI'd be fascinated to read more about this. I'd love to see a sample screenshot of a few of your cards too.
- SamPatt 1y agoSure, I'll write something up later. I'll give you two samples now. One reason I love the Obsidian + Markdown + Spaced Repetition plugin combo is how simple it is to make a card. This is all it takes: https://cdn.jsdelivr.net/gh/sampatt/media@main/posts/2025-04-26-geoguessr/image/2025-04-26-11-45.png https://cdn.jsdelivr.net/gh/sampatt/media@main/posts/2025-04... The top image is a screenshot from a game, and the bottom image is another screenshot from the game when it showed me the proper location. All I need to do is separate them with a question mark, and the plugin recognizes them as the Q + A sides of a flashcard. Notice the data at the bottom: <!--SR:!2025-04-28,30,245--> That is all the plugin needs to know when to reintroduce cards into your deck review. That image is a good example because it looks nothing like the vast majority of Google Street View coverage in the rest of Kenya. Very people people would guess Kenya on that image, unless they have already seen this rare coverage, so when I memorize locations like this and get lucky by having them show up in game, I can often outright win the game with a close guess. I also do flashcards that aren't strictly locations I've found but are still highly useful. One example is different scripts: https://cdn.jsdelivr.net/gh/sampatt/media@main/posts/2025-04-26-geoguessr/image/2025-04-26-12-02.png https://cdn.jsdelivr.net/gh/sampatt/media@main/posts/2025-04... Both Cambodia and Thailand have Google Street View coverage, and given their geographical proximity it can be easy to confuse them. One trick to telling them apart is their language. They're quite different. Of course I can't read the languages but I only need to identify which is which. This is a great starting point at the easier levels. The reason the pros seem magical is because they're tapping into much less obvious information, such as the camera quality, camera blur, height of camera, copyright year, the Google Street View car itself, and many other 'metas.' It gets to the point where a small smudge on the camera is enough information to pinpoint a specific road in Siberia (not an exaggeration). They memorize all of that. When possible I make the images for the cards myself, but there are also excellent sources that I pull from (especially for the non-location specific cards), such as Plonkit: https://www.plonkit.net/ https://www.plonkit.net/
- bobro 1y agoDid you include location metadata with the photos by chance? I’m pretty surprised by these results.
- SamPatt 1y agoNo, I took screenshots to ensure it. Your skepticism is warranted though - I was a part of an AI safety fellowship last year and our project was creating a benchmark for how good AI models are at geolocation from images. [This is where my Geoguessr obsession started!] Our first run showed results that seemed way too good; even the bad open source models were nailing some difficult locations, and at small resolutions too. It turned out that the pipeline we were using to get images was including location data in the filename, and the models were using that information. Oops. The models have improved very quickly since then. I assume the added reasoning is a major factor.
- vessenes 1y agoA) o3 is remarkably good, better than benchmarks seem to indicate in many circumstances B) it definitely cheats when it can — see this chat where it cheated by extracting EXIF data and wasn’t ashamed when I complained about it cheating: https://chatgpt.com/share/6802e229-c6a0-800f-898a-44171a0c7de4 https://chatgpt.com/share/6802e229-c6a0-800f-898a-44171a0c7d...
- SamPatt 1y agoAs a further test, I dropped the street view marker on a random point in the US, near Wichita, Kansas, here's the image: https://cdn.jsdelivr.net/gh/sampatt/media@main/posts/2025-04-26-geoguessr/image/2025-04-26-13-05.png https://cdn.jsdelivr.net/gh/sampatt/media@main/posts/2025-04... I fed it o3, here's the response: https://cdn.jsdelivr.net/gh/sampatt/media@main/posts/2025-04-26-geoguessr/image/2025-04-26-13-04.png https://cdn.jsdelivr.net/gh/sampatt/media@main/posts/2025-04... Nailed it. There's no metadata there, and the reasoning it outputs makes perfect sense. I have no doubt it'll be tricky when it can be, but I can't see a way for it to cheat here.
- tylersmith 1y agoThis is right by where I grew up and the broadcast tower and turnpike sign were the first two things I noticed too, but the ability to realize it was the East side instead of the West side because the tower platforms are lower is impressive.
- SecretDreams 1y ago> These models have more than an individual mind could conceivably memorize. #computers
- joenot443 1y agoSuper cool, man. Watching pro Geoguessr is my latest break-time activity, these geo-gods never cease to impress me. One thing I'm curious about - in high level play, how much of the meta involves knowing characteristics about the photography/equipment/etc. that Google used when they shot it? Frequently I'll watch rainbolt immediately know an African country from nothing but the road, is there something I'm missing?
- olex 1y agoIn the stream commentary for some of competitive Geoguessr I've watched, they definitely often mention the color and shape of the car (visible edges, shadow, reflections), so I assume pro players know which cars were used where very well.
- wongarsu 1y agoAlso things like follow cars (some countries had government officials follow the streetview car), the season in which coverage was created, camera glitches, the quality of the footage, etc. There is a lot of "legitimate" knowledge. With just a street you have the type of road surface, its condition, the type of road markings, the bollards, and the type of soil and vegetation next to the road, as well as the presence and type of power poles next to the road, to name a few. But there is also a lot of information leakage from the way google takes streetview footage.
- SamPatt 1y agoSpot on. Nigeria and Tunisia have follow cars. Senegal, Montenegro and Albania have large rifts in the sky where the panorama stitching software did a poor job. Some parts of Russia had recent forest fires and are very smokey. One road in Turkey is in absurdly thick fog. The list is endless, which is why it's so fun!
- simonw 1y agoDo you have a feel for how often StreetView published fresh imagery? When that happens, is there a wild flurry of activity in the GeoGuessr community as players race to figure out the latest patterns?
- roxolotl 1y agoOne thing I’m curious about is if they are so good, and use a similar technique as humans, because they are trained on people writing out their thought processes. Which isn’t a bad thing or an attempt to say they are cheating or this isn’t impressive. But I do wonder how much of the approach taken is “trained in”.
- otabdeveloper4 1y ago> how much of the approach taken is “trained in”. 100% of it is. There is no other source of data except human-generated text and images.
- neurostimulant 1y ago> when I asked it how, it mentioned that it knows I live nearby. > The process for how it arrives at the conclusion is somewhat similar to humans. It looks at vegetation, terrain, architecture, road infrastructure, signage, and it just knows seemingly everything about all of them. Can we trust what the model says when we ask it about how it comes up with an answer?
- simonw 1y agoNot at all. Models have no invisible internal state that they can access between prompts. If you ask "how did you know that?" you are effectively asking "given the previous transcript of our conversation, come up with a convincing rationale for what you just said".
- kqr 1y agoOn the other hand, since they "think in writing" they also do not keep any reasoning secret from us. Whatever they actually did is based on past transcript plus training.
- throwaway314155 1y agoRight but the reasoning/thinking is _also_ explained as being partially or completely performative. This is made obvious when mistakes that show up in chain of thought _don't_ result in mistakes in the final answer.l (a fairly common phenomenon). It is also explained more simply by the training objective (next token prediction) and loss function encouraging plausible looking answers.
- GeorgeDewar 1y agoThat writing isn't the only "thinking" though. Some thinking can happen in the course of generating a single token, as shown by the ability to answer a question without any intermediate reasoning tokens. But as we've all learnt this is a less powerful and more error-prone mode of thinking. So that is to say I think a small amount of secret reasoning would be possible, e.g. if the location is known or guessed from the beginning by another means and the reasoning steps are made up to justify the conclusion. The more clearly sound the reasoning steps are, the less plausible that scenario is.
- brundolf 1y agoI find this type of problem is what current AI is best at: where the actual logic isn't very hard, but it requires pulling together and assimilating a huge amount of fuzzy, known information from various sources They are, after all, information-digesters
- is-is-odd 1y agoit's just all compression? always has been
- skydhash 1y agoIt takes a lot of energy to compress the data. And a lot to actually extract something sensible. While you could just just optimize the single problem you have quite easily.
- fire_lake 1y agoWhich also fits with how it performs at software engineering (in my experience). Great at boilerplate code, tests, simple tutorials, common puzzles but bad at novel and complex things.
- brundolf 1y agoYep. But wonderful at aggregating details from twelve different man pages to write a shell script I didn't even know was possible to write using the system utils
- fundingshovel 1y agoI use it for this a lot.
- genewitch 1y ago[flagged]
- HenryBemis 1y agoIs it 'only' "aggregating details from twelve different man pages" or has it 'studied' (scraped) all (accessible) code in GitHub/GitLab/Stachexchange/etc. and any other publicly available coding repositories on the web (and for the case of MS the Git it owns)? Together with descriptions of what is right and what is wrong.. I use it for code, and I only do fine tuning. When I want something that is clearly never done before, I 'talk' to it and train it on which method to use, and for a human brain some suggestions/instructions are clearly obvious (use an Integer and not a Double, or use Color not Weight). So I do 'teach' it as well when I use it. Now, I imagine that when 1 million people use LLMs to write code and fine tune it (the code), then we are inherently training the LLMs on how to write even better code. So it's not just "..different man pages.." but "the finest coding brains (excluding mine) to tweak and train it".
- bjourne 1y agoGeoguessr pro zi8gzag tried out one of the AIs in a video: https://www.youtube.com/watch?v=mQKoDSoxRAY https://www.youtube.com/watch?v=mQKoDSoxRAY It was indeed extremely impressive and for sure would have annihilated me, but I believe it would have no chance to beat zi8gzag or any other top player. But give it a year or two and I'm sure it will crush any human player. Geoguessr is, afaict, primarily about rote memorization of various features (such as types of electricity poles, road signage, foilage, etc.) which AIs excel at.
- simonw 1y agoLooks like that video uses Gemini 2.0 (probably Flash) in streaming mode (via AI studio) from a few months ago. Gemini 2.5 might do better, but in my explorations so far o3 is hugely more capable than even Gemini 2.5 right now.
- neves 1y agoTry Alibaba's https://chat.qwen.ai/ https://chat.qwen.ai/ Activating reasoning
- intalentive 1y agoI wonder how it compares with StreetCLIP.
- matthewdgreen 1y agoI was absolutely gobsmacked by the three minute chain of reasoning this thing did, and how it absolutely nailed the location of the photo based on plants, the color of a fence, comparison with nearby photos, and oh yeah, also the EXIF data containing the exact lat/long coordinates that I accidentally left in the file. https://bsky.app/profile/matthewdgreen.bsky.social/post/3lnqblr5zys2t https://bsky.app/profile/matthewdgreen.bsky.social/post/3lnq...
- SamPatt 1y agoLol it's very easy to give the models what they need to cheat. For my test I used screenshots to ensure no metadata. I mentioned this in another comment but I was a part of an AI safety fellowship last year where we created a benchmark for LLMs ability to geolocate. The models were doing unbelievably well, even the bad open source ones, until we realized our image pipeline was including location data in the filename! They're already way better than even last year.
- ghaff 1y agoI was and am pretty impressed by Google Photo/Lens IDs. But I realized fairly early on that of course it knew the locations of my iPhone photos from the geo info stored in the photo.
- SamPatt 1y agoI dropped into Google Street View and tried to recreate your location, how did I do? https://cdn.jsdelivr.net/gh/sampatt/media@main/posts/2025-04-26-geoguessr/image/2025-04-26-17-51.png https://cdn.jsdelivr.net/gh/sampatt/media@main/posts/2025-04... Here's the model's response: https://cdn.jsdelivr.net/gh/sampatt/media@main/posts/2025-04-26-geoguessr/image/2025-04-26-17-52.png https://cdn.jsdelivr.net/gh/sampatt/media@main/posts/2025-04... I don't think it needed the EXIF data. I'd be curious if you tried it again yourself.
- dudeinhawaii 1y agoThis is super easy to test though (whether EXIF is being used). Open up Geoguessr app, take a screenshot, paste into O3. Doing this, O3 took too long (for the guessing period) but nailed 3 of 3 locations to within a kilometer. Edit: An interesting nuance of modern OpenAI chat interface is the "access to all previous chats" element. When I attempted to test O4-mini using the same image -- I inspected the reasoning and spotted: "At first glance, the image looks like Ghana. Given the previous successful guess of Accra Ghana, let's start in that region".
- pinkorchid 1y agoNote that they can claim to guess a location based on reasonable clues, but actually use EXIF data. See https://news.ycombinator.com/item?id=43732866 https://news.ycombinator.com/item?id=43732866 Maybe it's not what happened in your examples, but definitely something to keep an eye on.
- OhNoNotAgain_99 1y ago[dead]
- SamPatt 1y agoYes, I'm aware. I've been using screenshots only to avoid that. Check my last few comments for examples without EXIF data if you're interested to see o3's capabilities.
- panarky 1y agoAnd yet many will continue to confidently declare that this is nothing more than fancy autocomplete stochastic parrots.
- otabdeveloper4 1y agoI will. I am willing to bet that most of this geolocation success is based on overfitting for Google streetview peculiarities. I.e., feed it images from a military drone and success rate will plummet. You're anthropomorphizing a likelihood maximization calculator. And no, human brains are not likelihood maximization calculators.
- FergusArgyll 1y agoThis overconfidence without making one attempt kills me. Gemini 2.5 pro is free, try it! I gave it 2 pictures I took on a film camera. I then screenshotted the pictures (to remove exif, even though there isn't any). It nailed them. You lost your bet, now change your mind
- hammock 1y ago> It looks at vegetation, terrain, architecture, road infrastructure, signage, and it just knows seemingly everything about all of them. Someone explain to me how this is dystopian. Are Jeopardy champions dystopian too? It’s not crazy to be able to ID trees and know their geographic range, likewise for architecture, likewise for highway signs. Finding someone who knows all of these together is more rare , but imo not exactly dystopian Edit: why am I being downvoted for saying this? If anyone wants to go on a walk for me I can help them ID trees, it’s a fun skill to have and something anyone can learn
- cyanbane 1y agoHave you gleaned anything watching o3 make decisions on a photo? ( i.e. have you noticed if it has thought of anything you.. and other higher level players similar to you... have not? )
- SamPatt 1y agoThis is an interesting question. I watch the output with fascination, mostly because of the sheer breadth of knowledge. But thus far I can't think of anything that is categorically different from what humans do, it's just got an insane amount of knowledge available to it. For example, I gave it an image from a town on a small Chilean island. I was shocked when it nailed it, and in the output it said, "I can see a green wooden street sign, common to Chilean coastal towns on [the specific island]." I have an entire flashcard section for street signage, but just for practicality I'm limited to memorizing scores, possibly hundreds of signs if I'm insanely dedicated. I would still probably never have this one remote Chilean island. It does that for everything in every category.
- maayank 1y ago> when I asked it how, it mentioned that it knows I live nearby Did it mention it in its chain of thought? Otherwise, it could definitely output something because of X and then when asked why “rationalize” that it did it because Y
- larodi 1y agoIs it meaningful to conclude that this is an algorithm that pro GGsrs all follow, and one of them perhaps explained somewhere and the model took it? Is geo-guessing something that can be presented as algorithm or steps? Perhaps it is not as challenging as it seems, given one knows what to look for? not as challenging... as say complex differential geometry.
- RataNova 1y agoMakes me wonder what the ceiling even is for human players if AI can now casually flex knowledge that would take us years to grind out.
- zaik 1y agohttps://www.youtube.com/watch?v=QRqKPDJYyLE https://www.youtube.com/watch?v=QRqKPDJYyLE
- redbell 1y ago> It will use information it knows about you to arrive at the answer.. and when I asked it how, it mentioned that it knows I live nearby. Oh! RIP privacy :( I’ve pretty much given up on the idea that we can fully protect our privacy while still getting the most out of these services. In the end, it’s a tradeoff—and I’ve accepted that.
- zzzeek 1y agoGeoGuessr, well I guess that must have been a great training source for the models
- Suppafly 1y ago>I have a flashcard deck with hundreds of entries to help me remember road lines, power poles, bollards, architecture, license plates, etc. You're basically training yourself the same way an AI is trained at that point.