4 ms·
I agree with the conclusion. This is totally unsurprising to me as a ML engineer. If you put garbage data into the model, you get garbage predictions. That does
by fnbr 5y ago
I agree with the conclusion. This is totally unsurprising to me as a ML engineer. If you put garbage data into the model, you get garbage predictions. That doesn’t strike me as particularly novel. The same is true for cooking, after all.
However- this has been truly shocking to all of the non-technical stakeholders I’ve worked with. They take the stance that any large amount of data can be used to do ML on, presumably because they don’t know too much about what doing ML is like.
So I’m convinced the author is right, and I’m also convinced that there will be many attempts to use ML on EMRs.
- Forgeties79 5y agoGarbage in -> Garbage out is basically a Newtonian law at this point haha
- momenti 5y agoThat's not entirely true. Neural networks are fairly robust to noisy training data (a.k.a. garbage).[0] Well, stochastic gradient descent has the noise in its name. More training data can compensate for noisy data to some extent.[1] I'm not sure know if model size can also compensate for noisy data though, but would not be surprised if it did. [0] https://arxiv.org/abs/1705.10694 https://arxiv.org/abs/1705.10694 [1] https://arxiv.org/abs/2202.01994 https://arxiv.org/abs/2202.01994
- midjji 5y agoThere are very specific conditions for this to hold, mostly that the incorrect sample is surrounded by correct ones, and that the model is small enough or the error vanishingly rare. Notably the reference you gave also shows horrendous generalization performance, so its really just showing how easy it is to overparametrize. Input errors can be accounted for to some extent, but also under specific circumstances, eg. https://openaccess.thecvf.com/content_CVPR_2020/html/Eldesokey_Uncertainty-Aware_CNNs_for_Depth_Completion_Uncertainty_from_Beginning_to_End_CVPR_2020_paper.html https://openaccess.thecvf.com/content_CVPR_2020/html/Eldesok...
- Forgeties79 5y agoI mean the argument there is basically if there is enough good quality data, the bad data is somewhat (or mostly) compensated for. To which I would argue that is no longer a “garbage in garbage out” situation as most people use it.
- jerf 5y agoThere is the old quote that we've all seen: "On two occasions I have been asked, 'Pray, Mr. Babbage, if you put into the machine wrong figures, will the right answers come out?' I am not able rightly to apprehend the kind of confusion of ideas that could provoke such a question." - Charles Babbage I will say, I have some good news for the late great Charles Babbage. For the most part, people now do indeed understand that if you put small amounts of wrong figures into a machine, the wrong answers will come out. If nothing else, pocket calculators and math class have given them the direct experience of this. However, it seems that people still expect that if you put gigabytes or petabytes worth of wrong figures into a machine that somehow the right answer will pop out. Ah well. The road never ends, you know.
- d1sxeyes 5y ago> However, it seems that people still expect that if you put gigabytes or petabytes worth of wrong figures into a machine that somehow the right answer will pop out. The interesting fact is that if you put in lots of correct figures, and only one, slightly wrong figure, then the answer may be correct to an acceptable degree. For example, take 100 values, all exactly one, and find the mean. You'll find 1. If you take 99 values of exactly one and one value of two, the mean of your sample will be 1.01, which is close enough to still be useful. In some interpretations, it may even be rounded to 1, meaning that in some circumstances, incorrect figures can indeed sometimes lead to correct answers. Or if you're trying to find out what adding 1 and 4 gives you, but accidentally you add 2 and 3, you get the correct answer despite incorrect inputs. I think people are assuming that if you put gigabytes or petabytes worth of data into a machine, the number of 'wrong figures' will be lost as noise.
- paulmd 5y agothe problem is that for medical coding, this translates to "a small number of procedures will be coded wrong", and that's not a meaningfully better situation than "a small number of procedures can't be coded", and in fact is probably worse. So you need a reasonably high confidence threshold, and really in most cases you probably want to have a human manually review the problem (or questionable) cases.