4 ms·
I'm not 100% sure about this but I think part of the reason why its called you only look once is that it is single frame object detection. I do know that there
by aeleos 9y ago
I'm not 100% sure about this but I think part of the reason why its called you only look once is that it is single frame object detection. I do know that there are other types of networks that can previous frames into account, but the models themselves are much larger (in terms of VRAM requirements for loading into memory) and more computationally intensive. This network is special because it can run with very little power. From what I remember it can run on a 10W board at 6fps which compared to networks only a few years ago is 10x decrease
- Animats 9y agoIt's clearly single-frame; you can watch recognition succeed and fail from frame to frame. The "person" recognizer misses some clear faces, so it's not heavily face-oriented. Since it's single-frame, it's not recognizing articulated motion. Recognition seems to be limited to "person", "motorbike", "tie", "cell phone" (a gun, actually), "umbrella", "truck" (misidentified part of a train) "bench" (a railing) and "horse" (motorbike with duffel seen from rear). "Person", "umbrella", "tie", and "motorbike" seem to work; the others are kind of random. The trouble with running recognizers on Hollywood movies is that they have many conventions of what appears on screen and how big it is on screen. Are they training on such data? Good data sets would be side views from a moving vehicle, like Google StreetView data or just a GoPro pointed sideways while driving around.