4 ms·
I'd be really interested in knowing where you disagree and why, if you have time to do a writeup.
by native_samples 5y ago
I'd be really interested in knowing where you disagree and why, if you have time to do a writeup.
- dekhn 5y agoThere's not much to say. DM made a significant improvement to homology-based modelling using some very recent technology and fair amount of CPU power. They collaborate with a long-time CASP player (https://en.wikipedia.org/wiki/David_Tudor_Jones https://en.wikipedia.org/wiki/David_Tudor_Jones) who has been doing leading edge protein folding for decades. And CASP is a competition, which DM excels at. Finally, they picked the "easy" category of CASP, for which there is a huge amount of side data which ultimately, when properly processed, provides structural constraint information good enough to produce very high quality structures. By not addressing ab initio protein folding, they didn't solve the protein folding problem, they solved the "homologous protein structure prediction problem", which is far less grand. I did CASP back in 2001 or 2002 using alignment methods and protein threading, nothing special (we did terribly). But I told everybody at the conference that ML would eventually surpass the best human predictions, and people said (correctly) that at the time: we didn't have enough data (3D structures) to train on, we didn't have algorithms that were any good (practical backprop for deep NNs, embeddings, transformers), and there wasn't enough CPU time. The data problem was addressed (by having huge numbers of sequence alignments for most protein structures), the algorithms have obviously gotten better, especially transformers, and DeepMind has a lot of CPU (well, TPU). So basically my prediction came true when the underlying restrictions were removed (which to me seems "obvious"). The academic community would have done this two years after DM if DM hadn't done it (although their wins probably would be much smaller). DM just got there faster, by exploiting a number of advantages. None of this really moves anything forward. Being able to predict protein structures of homologous proteins is not the grand challenge problem, and what DM did tells us nothing about the physics of folding. And it doesn't produce results good enough to do structure based drug design. Really, they should have just put out a PR that said they did well in CASP and would be open sourcing the model, the training data, and a trained model, and then done that six months later.
- dekhn 5y agoOne update: they have now released a github project and dataset to train the model. So the second half of my last sentence has been resolved.
- native_samples 5y agoThanks.