4 ms·
People are really good at face recognition in real-life, because you have not just a static 2-d view of a person completely out-of-context, but a fully dynamic
by apu 12y ago
People are really good at face recognition in real-life, because you have not just a static 2-d view of a person completely out-of-context, but a fully dynamic 3-d view of someone you've probably seen before and you know the lighting environment you're in, meaning you can easily factor out those effects.
A computer algorithm operating on single 2d images (like LFW) has none of that. If you were to ask people to do the same task with the same data, they'd probably still do pretty well (much better than computers), but perhaps not perfectly.
The differences between image-to-image (same person) and different people is exactly what makes this problem so tough in the general case. It was shown about 20 years ago now that faces span a fairly low-dimensional manifold, and across different parts of this manifold, faces of different people do look much more similar than faces of the same person.
I probably shouldn't have explicitly written that formula for recognition, since it's not actually the formula, but more a general scaling rule-of-thumb. But even if we take it as given (hypothetically), there would be several issues. First, remember that 97.53% is on this particular dataset, which is very special in many ways (e.g., it was all collected over the course of a year from photos on Yahoo News, of public figures, with relatively lower-resolution images, and often very distinctive backgrounds). On a more realistic dataset, these numbers would be much lower.
Second, having more images of a person does help, but not nearly as much as you'd hope, because now there's an even bigger chance that you might accidentally match against someone else who happened to have their photo taken in the same pose, lighting conditions, and facial expression as your test photo, and it's very tough for algorithms to discount those confounding factors.
As for other applications of face recognition, let me describe one broad area that I think is pretty exciting, rather than a bunch of specific instances. One of the big shifts in user interface design is going to be "personalization" (to varying degrees). If a program or device or robot can recognize who you are, it can pro-actively change its settings/behavior/performance/etc. to better suit you. A very simple example is the new Kinect, which does some sort of recognition to load your saved profile/controller preferences. But that's just the tip of the iceberg.
And then of course there's personal photo collection management. If your photo organizer knew who everyone was in all your photos, it would be immensely useful in a number of ways. For starters, you could easily search for "photos of me & and my sister" or "my family", etc. You could also do more advanced analysis to start to get at more subtle things. For example, "my high school debate trip" might not immediately seem like it's related to face recognition, but in fact recognition might get you 90% of the way there.
Finally, if you've made it this far, I might as well plug my recent paper, "Photo Recall" [1], which looks at how to do advanced searches on your photo collection. Our system doesn't currently handle faces, but if you look at the kinds of queries we can do, it should become clear how they might extend with faces.
[1] http://homes.cs.washington.edu/~neeraj/projects/photo-recall/ http://homes.cs.washington.edu/~neeraj/projects/photo-recall...
- pessimizer 12y agoThanks, that's a lot to think about. >Second, having more images of a person does help, but not nearly as much as you'd hope, because now there's an even bigger chance that you might accidentally match against someone else who happened to have their photo taken in the same pose, lighting conditions, and facial expression as your test photo, and it's very tough for algorithms to discount those confounding factors. I was thinking that having two sets (for example a set of five and a set of three) of known matching images would mitigate that - in that you'd have 15 pairings to check rather than just one. >If a program or device or robot can recognize who you are, it can pro-actively change its settings/behavior/performance/etc. to better suit you. A very simple example is the new Kinect, which does some sort of recognition to load your saved profile/controller preferences. But that's just the tip of the iceberg. Personal devices and systems would rarely if ever be expected to distinguish between 1000s of people, though. Are you thinking more in terms of public devices, such as locks on house or car doors? That would seem as if it would be impossible just because out of thousands, someone could look more like you did when you set it up than you currently do.
- apu 12y agoThe problem with more pairings is that you have a greater chance of identifying the right person because of their additional photos, but that gets offset by the huge number of photos of everyone else in the database that you will now also get lots of "hits" on. For example, you might get lucky and 2 of the 3 test photos you have match well with 2 of the 5 database photos of the right person. But if you have enough people in the database, you're almost guaranteed to find some random person where all 3 of the test photos match quite well. So it often comes down to how you combine the results from the different test images. Do you take the max? Do you average them? etc. Each has different tradeoffs, and it's a topic under research at the moment ("image set recognition"), but I'd say there are no compelling results yet. Some personal devices might indeed be tuned for relatively fewer people, but these are precisely the cases that are the toughest: distinguishing between people in the same family, who're all going to look much more similar to each other than random people in a large database. But there are many scenarios where it might actually be distinguishing between 1000s of people. To take but one example, advertising. On the web, advertising is popular not just because it's "easy" and one of the few ways of successfully monetizing a site, but also because knowing more about a person allows for more targeted advertising which has a much bigger ROI. How would this translate to the "real world?" Personalization is clearly one of the ways to do this, and face recognition is ideal in many respects, given that it works from a distance, is non-invasive, and cheap to implement. (Obviously, there's a separate debate to be had about privacy and other issues, but I think it's hard to deny that many companies are interested in this.) I'm not sure that face recognition will be used for security purposes like house/car locks, just because it seems like other devices might get you better security sooner and cheaper (like RFID tags used in some car keys now, or some sort of ID from your phone).