7 ms·
PP-YOLO Surpasses YOLOv4 – State-of-the-art object detection techniques
- sillysaurusx 6y agoSuppose someone wanted to train a model to identify which decade a photo was taken in. What would be the current SOTA architecture for that type of task? (Suppose also that you had a few million labeled examples.) I like yolo because it’s a production grade object defector. It seems harder to find a production grade classifier. One amusing but dumb idea would be to use yolo for this: train the model on “photo from 1930,” “photo from 1940,” etc, where the bounding boxes cover the entire photo. But I’m curious what the professional solution might be.
- sangfroid_bio 6y agoDepends on whether you are looking for effectiveness vs efficiency. YOLO et al. are optimised for general feature extraction. They are ridiculously good if you need to build something quickly. If you are looking for the most efficient way e.g. lowest processing power per image then use a network that is optimised for the particular feature you are trying to extract.
- sillysaurusx 6y agoHighest accuracy. I was mainly wondering which model ML engineers reach for circa 2020 for classification tasks. Searching for classifiers brings up article after article presenting toy classifiers, but very little with a production focus.
- dontreact 6y agoI've seen people get good results with Efficientnet, and it scales to be quite large. In general, I look at whatever the best state of the art is on imagenet, but then look for implementations that look easy to use among the top few.
- rocauc 6y agoI am guilty of using your dumb idea with satisfactory performance. I'd recommend EfficientNet (or one of its may variations) for an off-the-shelf SOTA classifier. https://paperswithcode.com/sota/image-classification-on-imagenet https://paperswithcode.com/sota/image-classification-on-imag...
- modeless 6y agoEasy, just use an ImageNet classifier architecture with one category for each decade. No need to bother with object detection at all. You could fine-tune or train from scratch. Image classification is probably the single most-researched task with the widest variety of models available. You could select an architecture based on ease of implementation, efficiency, absolute accuracy, or any combination. Papers With Code has great lists of state-of-the-art models. Take your pick: https://paperswithcode.com/sota/image-classification-on-imagenet https://paperswithcode.com/sota/image-classification-on-imag...
- ericjang 6y agoOut of curiosity, what would such a model be used for? I worry that as ML is becoming more popular and powerful, people are jumping to use it on problems from which the answer cannot possibly be identified accurately from the inputs alone. This model's predictions would entirely based on historical trends of what images looked like, rather than something objective like carbon dating. I am not discouraging someone from building such a model, but it would be really helpful to know the context for which such a model is being developed. If it's just a hobby investigation, it would be cool to see how "predictable" dates are from images. I could even see it being used in forensics to provide a "first guess" as to when an image occurred, and helping with triaging of evidence. However, things become deeply problematic if the result from the image is fed to people as "ground truth" simply because the model was found to be accurate on a validation dataset. I certainly wouldn't want this model to be used to determine whether a suspect is innocent / guilty, or to be used naively by museums to date photographs.
- theplague42 6y agoFor example, creating a scrapbook of somebody's life from scanned photographs that lack any metadata.
- taneq 6y agoCould be useful for identifying faked historical photos?
- site-packages1 6y agoI would probably first start with examples from different eras and try to build a statistical classifier on it without resorting to detectors. Those seem like big overkill for such a task.
- nl 6y agoYolo isn't really suitable for classification - it's an object detector. For maximum accuracy there are high-quality well tested ResNet128 or ResNet152 implementations for PyTorch[1] and TensorFlow[2] that most people would use as a basis for classification tasks where the highest accuracy is needed. Lower quality ResNets (eg ResNet50) run a lot faster. Note that in this post, PP-YOLO replaces the older YOLO backbone with a ResNet50 (ResNet50-vd-dcn to be precise) backbone. EfficientNet is another good option[3], as is SE-ResNet. For this specific task it's not entirely clear what/how the classification is supposed to work though. The source of the photos is really, really important: if they are physical photos then the scanner used is important. And things like different film have different tone, and storage of physical photos matters a lot. [1] https://pytorch.org/hub/pytorch_vision_resnet/ https://pytorch.org/hub/pytorch_vision_resnet/ [2] https://tfhub.dev/google/imagenet/resnet_v2_152/feature_vector/4 https://tfhub.dev/google/imagenet/resnet_v2_152/feature_vect... [3] https://tfhub.dev/google/collections/efficientnet/1 https://tfhub.dev/google/collections/efficientnet/1
- p1esk 6y agoIt depends on if you're OK with the classifier overfitting on potentially meaningless details (like a shade of color of the paper most commonly used in that decade, etc) or if you actually want to classify the images based on content (style of clothes, etc).
- atty 6y agoSlightly tangential, but has anyone had a chance to use PaddlePaddle? I played around with it for a little bit a few months ago, and found it to be generally a regression in use-ability when compared to Pytorch or Tensorflow V2. I’d be interested to know what someone more experienced with it thinks.
- epberry 6y agoPretty interesting that the conclusion supports a slightly worse detector with a better framework.
- Imnimo 6y agoI think at a certain point, FLOP count will be more important than FPS. Like once you're running at real time, there aren't a lot of applications that care about 120 FPS vs 110 FPS. But there are a lot of situations where you care about the total number of operations (regardless of GPU parallelism) because you want to run on an edge device or have power constraints.
- CompleteSkeptic 6y agoThere actually is some work (https://arxiv.org/abs/2003.13630 https://arxiv.org/abs/2003.13630) claiming that FLOPS are a poor measure of real-world performance - with some of the more recent FLOP-efficient models actually running slower than older models.
- flafla2 6y agoForgive me if I’m being dense, but shouldn’t we expect performance to degrade if flop count per unit time is decreased, assuming performance is defined as overall runtime (FPS in this case)? It’s a trade off scenario where runtime performance is being balanced by other concerns such as power consumption.
- ganstyles 6y agoWe have started to hit this in the last two years. Models get deeper, more FLOPS but not practically better metrics except FPS. But for our use cases and on our hardware constraints, 30fps on a model that's a couple years old outperforms even the Yolov99 or whatever is SOTA.
- jimmySixDOF 6y ago>because you want to run on an edge device or have power constraints. Plug for TinyML [1] https://www.tinyml.org https://www.tinyml.org
- CompleteSkeptic 6y agoThis isn't directly relevant to PP-YOLO, but I'm surprised roboflow is still promoting "YOLOv5" - despite that model not having an associated paper and it not being made by the authors of the previous YOLO's.[1] The ML community has been asking the authors of that model to rename their project[2] because they are basically stealing publicity by making it seem like the next version of YOLO, despite its performance being worse than that of YOLOv4.[3] Roboflow has deflected this in the past by claiming they don't know if "YOLOv5" is the correct name[4], but by continuing to promote it, they are directly supporting it. In fact, I wouldn't be surprised that their claim of not being affiliated with Ultralytics to be either false or a half truth, given that all the top pages about "YOLOv5" were done by roboflow, including the first official announcement.[5] [1] https://github.com/AlexeyAB/darknet/issues/5920 https://github.com/AlexeyAB/darknet/issues/5920 [2] https://github.com/ultralytics/yolov5/issues/2 https://github.com/ultralytics/yolov5/issues/2 [3] https://github.com/AlexeyAB/darknet/issues/5920#issuecomment-642812152 https://github.com/AlexeyAB/darknet/issues/5920#issuecomment... [4] https://blog.roboflow.ai/yolov4-versus-yolov5/ https://blog.roboflow.ai/yolov4-versus-yolov5/ [5] https://blog.roboflow.ai/yolov5-is-here/ https://blog.roboflow.ai/yolov5-is-here/
- renewiltord 6y agoManufactured controversy.
- quietbritishjim 6y agoLegitimate confusion.
- airstrike 6y agoCan one trademark algorithm names? Not sure if possible but would be the obvious solution
- dheera 6y agoI don't think the original authors mind others using the word "YOLO". It's just ass-holish to call it YOLOv5 if you're not the original author, if at the very least because the original author is probably already working on something they plan to release eventually as "YOLOv5". If they had called it FOO-YOLO, YOLO-BLAH, YOLO++, or literally anything else, it would probably be perfectly fine.
- jcims 6y agoDo the image collections that these models are trained on have EXIF data? Is that included in the training?
- yeldarb 6y agoUsually, it is not. There have been some attempts to combine image and text data into a hybrid model but I’m not sure how widespread it is. Ex: http://cbonnett.github.io/Insight.html http://cbonnett.github.io/Insight.html
- jcims 6y agoI’ve been thinking of enduring some exit data (lens focal length, aperture, gravity vector, focus distance, etc) directly into the image as a band of scale bars across the bottom to inject a reference frame into the image directly. Not sure if it will do anything, just curious if it would help.
- mxuribe 6y agoIs something like Steganography[0] what you had in mind? (I'll assume your intent is less about hiding data, and more about inserting metadata "inline" for convenience, right?) If so, that sounds pretty handy! [0] = https://en.wikipedia.org/wiki/Steganography https://en.wikipedia.org/wiki/Steganography
- nl 6y agoYou can inject any data (in numeric form if you want) as a second data input into your neural network (assuming you are doing something custom). For example we do this to represent the position of specific sub-image parts we extract in an original image. > just curious if it would help Depends what you are trying to do. For example normal CNNs aren't rotation invariant[1] so if you know a gravity vector it can be useful to make your image upright. (Whilst CNN's aren't rotation invariant, it's common practice to augment training data by applying some rotation to the same image, so depending on how the CNN was training it may be fine) [1] https://stats.stackexchange.com/questions/239076/about-cnn-kernels-and-scale-rotation-invariance https://stats.stackexchange.com/questions/239076/about-cnn-k..., https://stackoverflow.com/questions/41069903/why-rotation-invariant-neural-networks-are-not-used-in-winners-of-the-popular-co https://stackoverflow.com/questions/41069903/why-rotation-in...
- tmabraham 6y agoAs a sidenote, can they get Redmon's name correct? In [1] and [2] they call him Redmond, and in [3] they call him PJ Reddie, which is his username and not his real name. It's not even that hard to be correct here... [1] https://blog.roboflow.ai/pp-yolo-beats-yolov4-object-detection/ https://blog.roboflow.ai/pp-yolo-beats-yolov4-object-detecti... [2] https://blog.roboflow.ai/a-thorough-breakdown-of-yolov4/ https://blog.roboflow.ai/a-thorough-breakdown-of-yolov4/ [3] https://blog.roboflow.ai/yolov5-is-here/ https://blog.roboflow.ai/yolov5-is-here/
- rocauc 6y agoThanks for the copyedits - I've updated to "Joseph Redmon."
- KingOfCoders 6y agoWas using "YOLOv5" (hope they merge all the efforts or relabel) yesterday and was amazed on how easy it was (no input image scaling or manipulation) and how fast it was with my model (<1h on RTX2080). Also on how easy it was to use in general (runs, ...) and how easy it was to install (Ubuntu 20.04). To me PyTorch is much more convenient than Darknet.
- mpfundstein 6y agonot only to you :-)
- 29athrowaway 6y agoBaidu should be sanctioned. It is one of the companies responsible for what's happening to Uighurs and other minorities. https://www.youtube.com/watch?v=OQ5LnY21Hgc https://www.youtube.com/watch?v=OQ5LnY21Hgc Computer vision technology, face recognition, object detection, image segmentation... it's all being weaponized. AI/ML frameworks should have more restrictive licenses that forbid mass surveillance.
- CommieDetector 6y agoStill just a clever way to write a lot of if else clauses