4 ms·
I have been working on this problem for the last year. As it turns out this problem is especially prominent when doing object detection in the real world on a d
by Datenstrom 7y ago
I have been working on this problem for the last year. As it turns out this problem is especially prominent when doing object detection in the real world on a drone platform. Besides a large number of angles/contexts being able to move in 3D adds a large number of scales also, and there is hardly any open training data besides the VisDrone dataset[0] which doesn't even address the scale issue.
It is certainly an interesting problem though. I can't talk much about my work but if anyone wants to collaborate on something open source addressing the core problem check my profile.
[0]: http://www.aiskyeye.com/upfile/Vision_Meets_Drones_A_Challenge.pdf http://www.aiskyeye.com/upfile/Vision_Meets_Drones_A_Challen...
- cs702 7y agoI recently saw a paper (which I reposted here on HN a while back) that perhaps could help with your problem. The paper proposes an approach that sort of induces models to learn representations of objects that are good at predicting "the most agreed-upon" rotations of the inputs. I'm not explaining it well. Anyway, a model in the paper achieved SOTA on a change-of-viewpoint dataset with a really tiny number of parameters. Might be worth a look: https://arxiv.org/abs/1911.00792 https://arxiv.org/abs/1911.00792
- Datenstrom 7y agoI have been looking for an excuse to use capsule networks. I can't investigate it as a possible solution at work because even if good results are achieved the performance problems[0] would prevent real world use. Definitely might look into that as a side project though. I have also been looking for an excuse to do something beyond like SENet in Julia. Maybe a new benchmark like this is what we need to get out of the rut. [0]: delivery.acm.org/10.1145/3330000/3321441/p177-Barham.pdf
- fheinsen 7y agoI'm the author of that paper. Happy to answer questions about it here.