4 ms·
Author here - it's a fair criticism, and I point it out in the article. However, I kept it in for a few reasons. I'm trying to show the model's full capabiliti
by SamPatt 1y ago
Author here - it's a fair criticism, and I point it out in the article. However, I kept it in for a few reasons.
I'm trying to show the model's full capabilities for image location generally, not just playing geoguessr specifically. The ability to combine web search with image recognition, iteratively, is powerful.
Also, the web search was only meaningful in the Austria round. It did use it in the Ireland round too, but as you can see by the search terms it used, it already knew the road solely from image recognition.
It beat me in the Colombia round without search at all.
It's worthwhile to do a proper apples and apples comparison - I'll run it again and update the post. But the point was to show how incredibly capable the model is generally, and the lack of search won't change that. Just read the chain of thought, it's incredible!
- k4rli 1y agoIt's still as much cheating as googling. Completely irrelevant. Even if it were to beat Blinky, it's not different from googlers/scripters.
- SamPatt 1y agoI disagree. I ran those rounds again, without search this time, and the results were nearly identical: https://news.ycombinator.com/item?id=43837832 https://news.ycombinator.com/item?id=43837832
- IanCal 1y agoI tried the image without search and it talked about Dornbirn anyway but ended up choosing Bezau which is really quite close. edit - the models are also at a disadvantage in a way too, they don't have a map to look at while the pick the location.
- SamPatt 1y agoYes, I re-ran those rounds and it made the same guesses without search, within 1km I believe. You're right about not having a map - I cannot imagine trying to line up the Ireland coast round without referencing the map.
- LeifCarrotson 1y agoThere's some level at which an AI 'player' goes from being competitive with a human player, matching better-trained human strategy against a more impressive memory, to just a cheaty computer with too much memorization. Finding that limit is the interesting thing about this analysis, IMO! It's not interesting playing chess against Stockfish 17, even for high-level GMs. It's alien and just crushes every human. Writing down an analysis to 20 move depth, following some lines to 30 or more, would be cheating for humans. It would take way too long (exceeding any time controls and more importantly exceeding the lifetime of the human), a powerful computer can just crunch it in seconds. Referencing a tablebase of endgames for 7 pieces would also be cheating, memorizing 7 terabytes of bitwise layouts is absurd but the computer just stores that on its hard drive. Human geoguessr players have impressive memories way above baseline with respect to regional infrastructure, geography, trees, road signs, written language, and other details. Likewise, human Jeopardy players know an awful lot of trivia. Once you get to something like Scrabble or chess, it's less and less about knowing words or knowing moves, but more about synthesizing that knowledge intelligently. One would expect a human to recognize some domain names like, I don't know, osu.edu: lots of people know that's Ohio State University, one of the biggest schools in the US, located in Columbus, Ohio. They don't have to cheat and go to an external resource. One would expect a human (a top human player, at least) to know that taxilinder.at is based in Austria. One would never expect any human to have every business or domain name memorized. With modern AI models trained on internet data, searching the internet is not that different from querying its own training data.
- mrlongroots 1y agoTo reframe your takeaway: you want to benchmark the "system" and see how capable it is. The boundaries of the system are somewhat arbitrary: is it "AI + web" or "only AI", and it is not about fairness as much as about "what do you, the evaluator, want to know".
- tshaddox 1y ago> There's some level at which an AI 'player' goes from being competitive with a human player, matching better-trained human strategy against a more impressive memory, to just a cheaty computer with too much memorization. Finding that limit is the interesting thing about this analysis, IMO! And a lot of human competitions aren't designed in such a way that the competition even makes sense with "AI." A lot of video games make this pretty obvious. It's relatively simple to build an aimbot in a first-person shooter that can outperform the most skilled humans. Even in ostensibly strategic games like Starcraft, bots can micro in ways that are blatantly impossible for humans and which don't really feel like an impressive display of Starcraft skill. Another great example was IBM Watson playing Jeopardy! back in 2011. We were supposed to be impressed with Watson's natural language capabilities, but if you know anything about high-level Jeopardy! then you know that all you were really seeing is that robots have better reflexes than humans, which is hardly impressive.