4 ms·
To those argue that LLMs might cheat by using EXIF, I saw a post recently on twitter (https://x.com/tszzl/status/1915212958755676350 https://x.com/tszzl/status/
by hashemian 1y ago
To those argue that LLMs might cheat by using EXIF, I saw a post recently on twitter (https://x.com/tszzl/status/1915212958755676350 https://x.com/tszzl/status/1915212958755676350) and out of curiosity, screen-captured the photo and passed it to O3. So no EXIF.
You can read the chat here: https://chatgpt.com/share/680a449f-d8dc-8001-88f4-60023323c70f https://chatgpt.com/share/680a449f-d8dc-8001-88f4-60023323c7...
It took 4.5m to guess the location. The guess was accurate (checked using Google Street View).
What was amazing about it:
1. The photo did not have ANY text
2. It picked elements of the image and inferred based on those, like a fountain in a courtyard, or shape of the buildings.
All in all, it's just mind-blowing how this works!
- thegeomaster 1y agoSee my other comment: https://news.ycombinator.com/item?id=43804041 https://news.ycombinator.com/item?id=43804041 4o can do it almost as well in a few seconds and probably 10-50x fewer tokens: https://chatgpt.com/share/680ceeff-011c-8002-ab31-d6b4cb622e52 https://chatgpt.com/share/680ceeff-011c-8002-ab31-d6b4cb622e... o3 burns through what I assume is single-digit dollars just to do some performative tool use to justify and slightly narrow down its initial intuition from the base model.
- HarHarVeryFunny 1y agoI don't see how this is mind blowing, or even mildly surprising! It's essentially going to use the set of features detected in the photo as a filter to find matching photos in the training set, and report the most frequent matches. Sometimes it'll get it right, sometimes not. It'd be interesting to see the photo in the linked story at same resolution as provided to o3, since the licence plate in the photo in the story is at way lower resolution than the zoomed in version shown that o3 had access to. It's not a great piece of primary evidence to focus on though since a CA plate doesn't have to mean the car is in CA. The clues that o3 doesn't seem to be paying attention to seems just as notable as the ones it does. Why is it not talking about car models, felt roof tiles, sash windows, mini blinds, fire pit (with warning on glass, in english), etc? Being location-doxxed by a computer trained on a massive set of photos is unsurprising, but the example given doesn't seem a great example of why this could/will be a game changer in terms of privacy. There's not much detective work going on here - just narrowing the possibilities based on some of the available information, and happening to get it right in this case.
- deleted 1y ago[deleted]
- simonw 1y agoIf you want to be impressed I suggest trying this yourself on your own photos. I don't consider it my job to impress or mind-blow people: I try to present as realistic as possible a representation of what this stuff can do. That's why I picked an example where its first guess was 200 miles off!
- HarHarVeryFunny 1y agoI'm not a computer. I expect a computer to also do better than me at memorizing the phone book, but I'm not impressed by it.
- simonw 1y agoIn that case, are you at all surprised that this technology did not exist two years ago?
- skydhash 1y agoDid it not, or no one was interested enough to build one? I’m pretty certain there’s a database of portraits somewhere where they search id details from photograph. Automatic tagging exists for photo software. I don’t see why that can be extrapolated to landmarks with enough data.
- hyperlink014 1y agoIt absolutely tried to use EXIF data when I asked it to guess the location. Here is proof - https://imgur.com/a/CHde2Cx https://imgur.com/a/CHde2Cx I couldn't attach the chat directly since it's a temporary chat.