9 ms·
Cleaning algorithm finds 20% of errors in major image recognition datasets
- groar 6y agoUsing simple techniques, they found out that popular open source datasets like VOC or COCO contain up to 20% annotation errors in. By manually correcting those errors, they got an average error reduction of 5% for state-of-the-art computer vision models.
- jessermeyer 6y agoGarbage in garbage out.
- rathel 6y agoNothing is however said about how the errors are detected. Can an ML expert chime in?
- ArnoVW 6y agomy guess would be using some sort of active learning. In other words: 1) building a model using the data set 2) making predictions using the training data 3) finding the cases where the model is the most confused (difference in probability between classes is low) 4) raising those cases to humans https://en.wikipedia.org/wiki/Active_learning_(machine_learning) https://en.wikipedia.org/wiki/Active_learning_(machine_learn...
- thibaut-duguet 6y agoI'm a Product Manager at Deepomatic and I have been leading the study in question here. To detect the errors, we trained a model (with a different neural network architecture than the 6 listed in the post), and we then have a matching algorithm that highlights all bounding boxes that were either annotated but not predicted (False Negative), or predicted but not annotated (False Positive). Those potential errors are also sorted based on an error score to get first the most obvious errors. Happy to answer any other question you may have!
- Zenst 6y agoWas the corrected datasets larger or smaller than the originals? Would also be interesting to see these improved datasets run thru simulation of crashes with existing datasets and see how they handle? Though not sure how you would go about that beyond approaching current providers of such cars for data to work thru and suspect they may be less open to admitting flaws and with that, may be a stumbling block. Certainly makes you wonder how far we can optimise such datasets to get better results. I know some ML datasets are a case of humans fine tuning and going thru examples and classifying them, and wonder how much that skews or effects error rates as we all know humans error.
- thibaut-duguet 6y agoTo answer your first question, we had both bounding boxes added and removed, and depending on the dataset, the main type of error was different (I'd say it was overall more objectifs that were forgotten, especially small objects). It would indeed be very interesting to see the impact of those improved datasets on driving, which is ultimately the task that is automated for cars. We've been working on many projects at Deepomatic not only related to autonomous cars, and we did see some concrete impact of cleaning the datasets beyond performance metrics.
- rathel 6y agoThank you for the explanation.
- alexchamberlain 6y agoie you get some to check where the model and the annotations disagree.
- deleted 6y ago[deleted]
- liquidify 6y agoCurious if you could find errors by comparing the results from the different models. Places where models disagree with each other more often would be areas that I would want to target for error checking.
- captain_price7 6y agoplus we'll have to register simply to see a few examples of mislabeling...that was disappointing
- thibaut-duguet 6y agoI've added screenshots of errors in the blogpost so that you have an idea of the errors we spotted. Let me know what you think of them.
- thaumasiotes 6y agoA couple notes on those screenshots: - In the cars-on-the-bridge image, the red bounding box for the semitruck in the oncoming lanes is too small, with its upper bound just above the top of the semi's windshield, ignoring the much taller roof and towed container. - In the same image, there are red bounding boxes around cars that exist, and also red bounding boxes around non-cars that don't exist. If false positives and false negatives are going to be represented in the same picture, it'd be nice to use different colors for them, so the viewer can tell whether the error was identified correctly or spuriously. - I have trouble understanding the "bus" screenshot. The caption says "(green pictures are valid errors) – The pink dotted boxes are objects that have not been labelled but that our error spotting algorithm highlighted." In other words, the green-highlighted pictures are false negatives considered from the perspective of the original data set, and the red-highlighted pictures are true negatives. Or alternatively, the green-highlighted pictures are true positives from the perspective of the error-spotting algorithm, and the red-highlighted pictures are false positives. What confuses me is that all 9 pictures are labeled "false positive" by the tabbing at the top of the screenshot.
- magicalhippo 6y ago> Create an account on the Deepomatic platform with the voucher code “SPOT ERRORS” to visualize the detected errors. Nice ad.
- thibaut-duguet 6y agoOur platform is actually designed for enterprise companies, so we don't provide open access unfortunately.
- magicalhippo 6y agoStill, couldn't you have included an example or two in the article no to illustrate the kind of errors we're talking about?
- scribu 6y agoI signed up and still couldn't see the errors. I just see 3 datasets with generic annotations.
- thibaut-duguet 6y agoThe process is actually a bit complicated but let me explain it to you. Once you are on a dataset, click on the label that you want and use the slider at the top right corner of the page to switch modes (we call it smart detection). You should then be able to access three tabs and the errors are listed in the False Positive and False Negative tabs (I've added a screenshot in the blogpost so that you can make sure to be at the right place). Let me know if you have any problem, thanks!
- scribu 6y agoThanks, I can see them now.
- CydeWeys 6y agoWhy aren't these data sets editable instead of static? Treat them like a collaborative wiki or something (OpenStreetMap being the closest fit) and allow everyone to submit improvements so that all may benefit. I hope the people in this article had a way to contribute back their improvements, and did so.
- seveibar 6y agoI'm working on this[1], my theory is the lack of a good IDE (rather than simple crowdsourcing interface) is the reason why it hasn't been done. Imagine if github had an integrated ide for editing large datasets. Also see dolt which is doing good work here. [1] https://github.com/UniversalDataTool/universal-data-tool https://github.com/UniversalDataTool/universal-data-tool
- lmkg 6y agoOne major use of the public datasets in the academic community is to serve as a common reference when comparing new techniques against the existing standard. A static baseline is desirable for this task. You could maybe split the difference by having an "original" or "reference" version, and a separate moving target that incorporates crowdsourced improvements.
- CydeWeys 6y agoThis sounds like a revisioning system would help a lot. Have a quarterly or annual release cycle or something, so that when you want to compare performance across techniques, you just train both of them to the same target (and ideally all the papers coming out at roughly the same time would already be using the same revision anyway). You'd always work with a versioned release when training models, and you'd only typically work with HEAD when you were specifically looking to correct flaws in the data (as the authors in the linked article are).
- 6gvONxR4sf7o 6y agoThe datasets serve as benchmarks. You get an idea for a new model that solves a problem current models have. These ideas don't pan out, so you need empirical evidence that it works. To show that your model does better than previous models, you need some task that your model and previous models can share for both training and evaluation. It's more complicated than that, but that's the gist. It would be so wasteful to have to retrain a dozen models that require a month of GPU time each on to serve as baselines for your new model...
- kent17 6y ago> We then used the error spotting tool on the Deepomatic platform to detect errors and to correct them. I'm wondering if those errors are selected on how much they impact the performance? Anyway, this is probably a much better way of gaining accuracy on the cheap than launching 100+ models for hyperparameter tuning.
- kent17 6y ago20% annotation error is huge, especially since those datasets (COCO, VOC) are used for basically every benchmark and state of the art research.
- rndgermandude 6y agoAnd people wonder why I am still a bit skeptical of self-driving cars....
- s1t5 6y agoIn one of his fastai videos Jeremy Howard makes the point that wrong labels can act as regularization and you shouldn't worry too much about them. I'm a bit skeptical as to how far you can push this but you certainly don't need perfect labelling.
- groar 6y agoThat is true up to a certain point (for instance, in my experience, having bounding boxes that are not pixel-perfect acts as a regularizer), but there is also a good chance that you are mislabelling edge cases, situations that happen rarely, and that definitely hurts the performance of the neural network to make a correct prediction on these difficult / uncommon scenarios.
- kingvash 6y agoWe did some interesting experiments with Go where we inverted the label of who won and measured what impact that had on the final model. This is a binary label so it's probably more impactful (it's the only signal we are measuring) From memory it had only a small impact (2% strength) with ~7% of results flipped, at 4% it was hard to measure the impact (<1%)
- strbean 6y agoAlso, this is applies to mislabeled data in your training set, right? Not a good thing if it is in your test set.
- 6y ago
- frenchie4111 6y agoBest I can tell, they are using the ML model to detect the errors. Isn't this a bit of an ouroboros? The model will naturally get better, because you are only correcting problems where it was right but the label was wrong. It's not necessarily a representation of a better model, but just of a better testing set.
- groar 6y agoIf I understand correctly they actually did not change the test set.
- frenchie4111 6y agoAh, I guess I missed that
- fwip 6y agoThe title here seems wrong. Suggested change: "Cleaning algorithm finds 20% of errors in major image recognition datasets" -> "Cleaning algorithm finds errors in 20% of annotations in major image recognitions." We don't know if the found errors represent 20%, 90% or 2% of the total errors in the dataset.
- deleted 6y ago[deleted]
- groar 6y agoYes agreed with that ! I can't change the title unfortunately
- gringomarketing 6y agoGringo Marketing Article Spinner is constructed to offer the very best spinning tools with the greatest worth for all users and all languages. We know that premium article is crusial for every single people and company to meet their target marketing needs and requirements. We understand what is needed to supply advanced, yet user friendly software application to provide users the capability to make content quick create the short articles they need with no high costs or complicated settings. Visit https://gringomarketing.com/article-rewriter https://gringomarketing.com/article-rewriter
- m0zg 6y agoAn idea on how this could work: repeatedly re-split the dataset (to cover all of it), and re-train a detector on the splits, then at the end of each training cycle surface validation frames with the highest computed loss (or some other metric more directly derived from bounding boxes, such as the number of high confidence "false" positives which could be instances of under-labeling) at the end of training. That's what I do on noisy, non-academic datasets, anyway.
- jontro 6y agoWeird behaviour on pinch to zoom (macbook). It scrolls instead of zooming and when swiping back nothing happens. Another example of why you should never mess with the defaults unless strictly necessary.
- benibela 6y agoThese things are why I stopped doing computer vision after my master thesis