4 ms·
"The simple idea is to do stuff like train your model on randomized subsets of your data and then compare its performance to using all the data you have." How
by SolaceQuantum 7y ago
"The simple idea is to do stuff like train your model on randomized subsets of your data and then compare its performance to using all the data you have."
How do you do this when you cannot verify that your data, in subset or in whole, is accurate? And furthermore you don't know how inaccurate it is?
- TomVDB 7y agoMy intuitive take on it: You train on a subset of the initial data. Even if the data has a certain number of incorrect frames, it should still do a decent job getting a lot of things right. Then you manually loop through all the images of the data set for which the network has detected something that isn't present in the annotations (and vice versa). If the network correctly identified a missing item that wasn't in the original set, all you need to do is press "correct" (and, again, vice versa). You now have an improved data set. Retrain, rinse, repeat. Eventually, you'll converge to a case where you have consistency between training and annotations. And then, you manually go through all images again to weed out the final mistakes. The benefit of this method is that it's much faster to click "correct" that it is to draw rectangles on the screen to label something.
- floatingatoll 7y agoThe drawback of this method is that it's easy to miss a pedestrian amidst the sea of rectangles already drawn. You would be far better off paying people $x/hour to look at an image for 30 seconds and answer the question "Does this image have any people in it?" Y/N and watching for the human who says "Yes" when the AI says "No". Their accuracy rating will help distinguish who is best able to detect pedestrians that AI and other people missed (and who is just random-clicking Y/N for pay), and their group effort will ensure that someone eventually sees the pedestrian, even if no one else has. Asking them to draw boxes distracts them from their job, which is "verify that we are able to detect human beings with perfect accuracy vs. a hundred people trying to detect human beings". (At worst, ask them to click on the person. No need for a box. Either it's a person or it isn't. If it is, and your AI missed it, then what they think is the right kind of box to draw is the least of your concerns.)
- DrStalker 7y agoWhy pay people to do this when you can make a CAPTCHA that requires them to do it for free? Google must have a really good data set from all those "click the boxes containing X" tests they make people do.
- floatingatoll 7y agoMost people don't have the funds on hand to use predatory pricing^ tactics to train machine learning algorithms. ^ "the pricing of goods or services at such a low level that other suppliers cannot compete and are forced to leave the market", to quote Google's definition (that was itself taken from some other company's dataset).
- TomVDB 7y agoBoth methods aren't mutually exclusive.
- dTal 7y agoI've been thinking a lot about this sort of thing lately, and isn't it the case that ideally you shouldn't need to manually confirm or reject mismatches? If the learning program maintains a probability density for "training labels are wrong", definitive ground truth should be unnecessary - eventually it will figure out the mismatches by itself. As I understand it, this is the core of recursive Bayesian estimation. At the end of the day we don't really have ground truth for anything - it's all filtered through senses with error bars. So any learning process needs to be robust to that.
- MauranKilom 7y agoSure, you can modify your model to better account for bad training data. But you could also fix the training data. The previous comment pointed out that fixing the training data (to a high degree at least, using the predictions of an intermediate version) is significantly faster than the initial labeling.
- joshvm 7y agoYou need to have a known-good test dataset that is as representative as possible. Something that you are absolutely sure is golden - i.e. someone has gone through it manually and verified all the labels are correct and complete. Then it doesn't matter quite so much if your input labels are noisy, because if you perform well at test time, your model is working. If you have no idea of the accuracy of any of your data then you're probably asking the wrong question of it. You can do things like test for consistency using cross-validation, e.g. does half of your dataset predict the other half with the same kind of performance? But that can't detect the same errors repeated throughout your data.