4 ms·
Looks like this is using training data from the PASCAL VOC object detection challenge [1], which is the standard benchmark for evaluating object detection perfo
by apu 14y ago
Looks like this is using training data from the PASCAL VOC object detection
challenge [1], which is the standard benchmark for evaluating object detection
performance in computer vision.
Object detection is an extremely tough problem (some would say it is the computer vision
problem ;-)), and while we've made a lot of progress in the past decade, the best
methods are still terrible [2] -- average detection precision between 30-50%.
For reference, most consumer applications require an AP of 90+% to be considered
usable.
So if this is a completely automated solution, it's not going to be able to do
much better, unless the creators can make massive (I mean orders-of-magnitude)
improvements on the state-of-the-art.
But that being said, there are some applications where lower performance is
acceptable. And if you add some manual verification, you could conceivably
make this much better (with an increase in latency, though). Another possibility
is to specialize on a certain type of input image (e.g., if you're a company
taking photos in your warehouse, where all your photos look very similar and/or
you can control the lighting and environment).
Still, I'm excited to see companies attempting to take object detection out to
the real world. All the best to these guys!
[1] http://pascallin.ecs.soton.ac.uk/challenges/VOC/ http://pascallin.ecs.soton.ac.uk/challenges/VOC/
[2] http://pascallin.ecs.soton.ac.uk/challenges/VOC/voc2011/results/index.html http://pascallin.ecs.soton.ac.uk/challenges/VOC/voc2011/resu...
- eloisius 14y agoIsn't the 30-50% only applicable to doing object recognition? I.e. multi-classification. In this case, you have to tell it which object you're looking for.
- apu 14y agoThe relevant table on the results page is Table 3, which is detection performance. Classification is actually an easier problem (see Table 1), in part because the types of scenes in which different classes appear are often quite different, making it easy to avoid some "easy" mistakes.
- EwanG 14y agoOne of my main hobbies is photography. I do mainly outdoor shots, and really enjoy macros of flowers. The problem being that "oh last weekend I took an amazing shot of a purple flower" isn't all that helpful for someone who is trying to find a picture of an iris. When someone comes up with an algorithm that can take my shot, compare it to a library, and tell me what wildflower it is, I will be a happy camper. I suspect Flickr and 500px will also become more valuable places since it would be possible to correlate geotagged shots with flora to document what seems to be there.
- apu 14y agoIt's not quite what you want, but I worked on Leafsnap [1], which automatically identifies trees by their leaves, using computer vision techniques. We focused on leaves since they are present throughout much more of the year than flowers. Our free apps also include high-resolution, high-quality photos of all aspects of the species we cover -- leaves, flowers, fruits, bark, etc. So you can at least browse through and compare the flowers you're looking at with those in the app. Our current coverage is of the trees of the northeast US (about 200 species), but we are working on expanding that. [1] http://leafsnap.com http://leafsnap.com