3 ms·
Okay, I decided to benchmark a bunch of AI models with geoguessr. One round each on diverse world, here's how they did out of 25,000: Claude 3.7 Sonnet: 22,75
by CSMastermind 1y ago
Okay, I decided to benchmark a bunch of AI models with geoguessr. One round each on diverse world, here's how they did out of 25,000:
Claude 3.7 Sonnet: 22,759
Qwen2.5-Max: 22,666
o3-mini-high: 22,159
Gemini 2.5 Pro: 18,479
Llama 4 Maverick: 14,316
mistral-large-latest: 10,405
Grok 3: 5,218
Deepseek R1: 0
command-a-03-2025: 0
Nova Pro: 0
- nemo1618 1y agoNeat, thanks for doing this!
- msephton 1y agoHow does Google Lens compare?
- CSMastermind 1y agoI tried it but as far as I can tell Google Lens doesn't give you a location - it just describes generally what you're looking at.
- msephton 1y agoI had cause to try Google Lens today and found the location to exact address thanks to a veterinary clinic which was in the background of an image. ChatGPT got the country but wrong city.
- arresin 1y agoWhat about 04-mini-high ?
- CSMastermind 1y agoOpenAI's naming confuses me but I ran o4-mini-2025-04-16 through a game and it got 23,885
- arresin 1y agoInteresting. It supports what they said (this is the model with good visual reasoning)