2 ms·
I worked in cancer genomics (and applications of ml to that) and have a couple of papers to my name. (Software methods so repeatable in the sense you can rerun
by micro_cam 5y ago
I worked in cancer genomics (and applications of ml to that) and have a couple of papers to my name. (Software methods so repeatable in the sense you can rerun them but that doesn't mean the data will mean anything).
Every ML practitioner should spend some time working with genetic data just to realize how weird things can get when you have millions of sparse features for a few thousand cases and all sorts of batch effects. Like to do it well you need to control for the lightbulbs in the microarray machine you were using or the sequencing center used, minor version of the sequencing technology and reagents order the subjects were sequenced in, who ran the machine, geographic origin of subjects etc.
And then when you go out and try to tie your data into the literature you need to be aware of all sorts of self reinforcing bias in which genes get published on.
Amazing and artful applications of mathematics have been developed to address all of these and the care bioinformatic papers address things like cross validation with makes ML look like a joke. Beyond that its even common to go out and get more data or validate with different technology before publishing but it is still easy to get things wrong.
If we truly want reproducible research we need to address the batch effects with less noisy sequencing machines and an assembly line like approach to generating orders of magnitude more data as cheaply as possible including new data on new subjects to verify studies after the fact.