3 ms·
http://en.wikipedia.org/wiki/Scale-invariant_feature_transform http://en.wikipedia.org/wiki/Scale-invariant_feature_transfo... Facebook almost certainly has mo
by drats 15y ago
http://en.wikipedia.org/wiki/Scale-invariant_feature_transform http://en.wikipedia.org/wiki/Scale-invariant_feature_transfo...
Facebook almost certainly has more photo information than TinEye or Flickr, and of indoor environments probably more than Google (which has reverse image search too). Across any given bar or hospital Facebook would have maybe 5-10 other people with albums tagged with the name/gps/check-in. They'd only need one other album though.
SIFT more or less turns every image into a bag-of-words. Your single photo, even at different angles, is going to have a heavy match with photos they have. If you upload a whole album they are going to have tons of matches and they can be more or less certain of the location. To say nothing of adding even the most basic geoip-to-city lookups that would narrow you down to at least five cities that you and your social network inhabit. But the extra information they have is besides the point, SIFT is enough; hospital rooms look alike to us, to SIFT they don't.
- yogrish 15y agoBut why that whole technology for just making a Suggestion? There is something else they are working on...something really BIG and a complete game changer I guess.
- drats 15y agoWell if the accuracy is 95% why not turn all Facebook users into one huge mechanical turk batch process to get to 99.9%? And it's just convenient for people not to have to tag. Then you have the data sitting around for mapping the insides of all buildings like Google has it for street view. Sure your Asimo-style robot butler/"something really BIG" will be that more efficient with internal mappings of most public spaces when you release it in 2030, but the convenience is sufficient. Remember that Facebook beat Myspace more or less on interface (combined with a few other factors like exclusivity), they aren't going to let Google or some other competitor get ahead of them by having an interface convenience edge of 5-10% which might cause a Myspace-to-Facebook style exodus. Kings who have committed regicide on their predecessor are all too aware of how they got into power.
- drats 15y agoAlso consider this, use SIFT on video stills and collect all the handheld video of a particular concert. Use machine learning to combine all the audio tracks and clean them up into high quality audio. Stitch the video together, using textures from the higher resolution stills people take, to allow people to relive the concert with a massive panoramic video (or 3D) with high quality sound. Use it to launch a music competitor to Google or destroy Ticketmaster (well overdue..) as bands and venues won't have to hire video production companies to record concerts if they sell tickets through FB. All that technology is in current research papers and prototypes at the moment, it would probably only take two years to put it together at worst if they aren't working on it already.
- dbarlett 15y agoPreviously: http://news.ycombinator.com/item?id=3293324 http://news.ycombinator.com/item?id=3293324
- rmc 15y agoOr it could just a be a way to get users to add more meta data to their data, which allows them to increase engagement
- obtu 15y agoThey might be combining some low-accuracy image similarity to match on images from his friends list, and infer the location from those other images.
- jsaxton86 15y agoIf I were to take a picture of a baby at a hospital, wouldn't the majority of the features be of the baby, not the hospital? I suppose if there's at least one picture in the album that's of the actual hospital, that's probably all you would need to infer the rest of the set was taken at the same hospital.
- T-hawk 15y agoThere's plenty of ways to infer location information from a single picture focusing on a baby. The hospital could be identified from something as small as one piece of paper in the background with letterhead or another identifier. Or the face of a nurse in the background that was previously known to be at this hospital. Quite possibly the room layout or particular pieces of equipment in certain arrangements. Landmarks outside a window. Pictures leak a ton of side channel information outside of their subject matter.
- iradik 15y agoIt's not that sophisticated. It's a picture of a baby being taken at a hospital. There's probably hundreds of photos of babies being taken with GPS from that hosptial. FB image recognition probably thinks all these babies are the same "person". Then a separate system comes along and takes photos that do not have location info and tries to match them up with ones that do. It finds a match with the "hospital baby" and then asks for verification. Person says YES and it adds that baby to the "hospital baby" pool as well.
- tibbon 15y agoOne other even simpler possibility, Facebook is looking at where you checked in or made status updates from and comparing the timestamps? "I just had a baby" - updated from hospital (Upload photo with timestamp near that status time) Match the two
- ArbitraryLimits 15y ago> hospital rooms look alike to us, to SIFT they don't Well, to SIFT they'll all look different. Except for the hospital room in Portland, OR and the hospital room in Portland, ME that each happen to have a sign with "Portland Hospital" visibile in the picture. While I'm alarmed about the potential of computer vision to compromise my privacy, I have yet to see anything in actual use even be competent, let alone alarming.
- apu 15y agoSIFT is not enough for this task. There are many computer vision researchers working on this application (automatically inferring location of an image), and accuracies are pretty low right now, except for popular landmarks. The problem is that single SIFT features are not very distinctive, and so you need many of them in common between two images to get reliable matches. So you essentially need images taken from very similar locations. You can ameliorate this requirement slightly by using fancy tricks in how you aggregate SIFT features together, but the fundamental constraint remains. This precondition is satisfied at popular landmarks, since there is a good chance that someone's taken a photo from the same location you're standing at, but in general, this condition is not that easy to meet, and hence the poor accuracy of current approaches. Finally, SIFT is great only for roughly-2d, distinctively textured areas. However, a lot (maybe most?) of the world does NOT fall into this category -- many things are either not textured, not distinctive, or are 3d. Something other than local descriptors (of which SIFT is the best known example) will be needed to understand these kind of scenes. If you're curious about this line of research, a good place to start is the IM2GPS paper, which was the first major work (that I'm aware of) to look at this problem: http://graphics.cs.cmu.edu/projects/im2gps/ http://graphics.cs.cmu.edu/projects/im2gps/