3 ms·
I commend you. Very few people actually try feeding images into these things to show how bad they are. For a real shock at how poor performance is, try feeding
by pakl 10y ago
I commend you. Very few people actually try feeding images into these things to show how bad they are.
For a real shock at how poor performance is, try feeding frames of video in. (Video is different because it generally doesn't have carefully stereotypically framed, exposed, in-focus content).
Edit: examples of ResNet applied to video by a colleague http://blog.piekniewski.info/2016/08/12/how-close-are-we-to-vision/ http://blog.piekniewski.info/2016/08/12/how-close-are-we-to-...
- Capt-RogerOver 10y agoVery informative quote: So where do the reports of superhuman abilities come from? Well since there are many breeds of dogs in the ImageNet, an average human (like me) will not be able to distinguish half of them (say Staffordshire bullterrier from Irish terrier or English foxhound - yes there are real categories in ImageNet, believe it or not). The network which was "trained to death" on this dataset will obviously be better at that aspect. In all practical aspects an average human (even a child) is orders of magnitude better at understanding/describing scenes than the best deep nets (as of late 2015) trained on ImageNet.