5 ms·
Yeah, the author does note that in the article. He also points it out in the conclusion: > If it’s using other information to arrive at the guess, then it’s no
by silveraxe93 1y ago
Yeah, the author does note that in the article. He also points it out in the conclusion:
> If it’s using other information to arrive at the guess, then it’s not metadata from the files, but instead web search. It seems likely that in the Austria round, the web search was meaningful, since it mentioned the website named the town itself. It appeared less meaningful in the Ireland round. It was still very capable in the rounds without search.
- rafram 1y agoSeems like they should've just repeated the test. But without the huge point lead from the rounds where it cheated, it wouldn't have looked very impressive at all.
- silveraxe93 1y agoPeople found the original post so impressive they were saying that it had to be coming from cheating by looking at EXIF data. The point of this article was to show it doesn't. It got an unfair advantage in 1 (and say 0.5) out of 5. With the non-search rounds still doing great. If you think this is unimpressive, that's subjective so you're entitled to believe that. I think that's awesome.
- godelski 1y agoSorry, I think I misread you. I think you said People accused it of cheating by reading EXIF data. They were wrong, it cheated by using web search. That makes the people that accused it of cheating wrong and this post proves that. And is everyone forgetting that what OpenAI shows you during the CoT is not the full CoT? I don't think you can fully rely on that to make claims about when it did and didn't search
- SamPatt 1y agoThat's inaccurate. It beat me by 1,100 points, and given the chain of thought demonstrated that it knew the general region of both guesses before it employed search, it would likely have still beaten me in those rounds. Though probably by fewer points. I will try it again without web search and update the post though. Still, if you read the chain of thought, it demonstrates remarkable capabilities in all the rounds. It only used search in 2/5 rounds.
- godelski 1y agoI'd be interested at capabilities without web search. The displayed CoT isn't the full CoT so it's hard to know if it really is searching or not. I mean it isn't always obvious when it does. Plus, the things are known to lie ¯\_(ツ)_/¯
- SamPatt 1y agoI do understand the skepticism, and I'll run it again without search to see what happens. But a serious question for you: what would you need to see in order to be properly impressed? I ask because I made this post largely to push back on the idea that EXIF data matters and the models aren't that capable. Now the criticism moves to web search, even though it only mattered in one out of five rounds. What would impress you?
- mattmanser 1y agoYou're kinda being your own worse enemy though. "Technically cheating"? Why even add the "technically". It just gives the impression that you're not really objectively looking for any smoke and mirrors by the AI.
- SamPatt 1y agoI hear you - but I had already read through the chain of thought which identified the right region before search, and had already seen the capabilities in many other rounds. It was self-evident to me that the search wasn't an essential part of the model's capabilities by that point. Which turned out to be true - I re-ran both of those rounds, without search this time, and the model's guesses were nearly identical. I updated the post with those details. I feel like I did enough to prove that o3's geolocation abilities aren't smoke and mirrors, and I tried to be very transparent about it all too. Do you disagree? What more could I do to show this objectively?
- godelski 1y ago> What would impress you? I want to be clear that you tainted the capacity to impress me by the clickbait title. I don't think it was through malice, but I hope you realize the title is deceptive.[0] (Even though I use strong language, I do want to clarify I don't think it is malice) To paraphrase from my comment: if you oversell and under deliver, people feel cheated, even if the deliverable is revolutionary. So I think you might have the wrong framing to achieve this goal. I am actually a bit impressed by O3's capabilities. But at the same time you set the bar high and didn't go over or meet it. So that's going to really hinder the ability to impress. On the other hand, you set the bar low, it usually becomes easy to. It i slike when you have low expectations for a movie and it's mediocre you still feel good, right? [0] https://news.ycombinator.com/item?id=43836791 https://news.ycombinator.com/item?id=43836791
- SamPatt 1y agoI did repeat the test without search, and updated the post. It made no difference. Details here: https://news.ycombinator.com/item?id=43837832 https://news.ycombinator.com/item?id=43837832
- clhodapp 1y agoThe question is not only how much it helped the AI model but rather how much it would have helped the human. This is because the AI model could have chosen to run a search whenever it wanted (e.g. perhaps if it knew how to leverage search better, it could have used it more). In order for the results to be meaningful, the competitors have to play by the same rules.