16 ms·
No, DeepMind has not solved protein folding
- physicsguy 6y agoHighly recommend the recent book Science Fictions by Stuart Richie which discusses in one chapter how scientists often write the press releases themselves and then the media just copy them almost verbatim. This is obviously a little different because the scientists are working for a company rather than a University, but I think the same thing applies here.
- gspr 6y agoReally? I'm only familiar with university PR people writing them on the scientists' behalf, inducing a lot of cringe on the part of the latter.
- throwaway74453 6y agoThis is very different. DeepMind submitted their solutions to a hidden test set at CASP14, with a known metric and a value that the subfield considered to mean solving protein folding. They cleared the bar with AlphaFold2. This is not a PR stunt by a company trying to spin their meager results.
- saadalem 6y agoThe bottom line? Understanding disease biology and biological networks is the rate-limiting step in revolutionizing drug discovery. No matter how well the biophysics and structural biology progresses, medicinal chemists could be targeting the wrong protein the whole time. That being said, this achievement almost seems miraculous given where we were decades ago, and it's waiting for us as we work hard to figure out the disease biology piece of the puzzle.
- dekhn 6y agoYes, this 100%. Back when I started my PhD program in biophysics, the exact opposite was stated: "if you can determine the 3D structure of the protein that causes a disease, you can target it with a drug.". I wasted decades believing that paradigm because really smart people kept repeating it. unfortunately there is no clear next step for the disease biology part of the study, as far as I can tell, except to collect enormous amounts of high quality data about diseases, typically one or a few at a time, and hope you get lucky finding something (IE, serendipity is just as important as intelligence).
- m00x 6y agoYou don't get determinism when working within the body. Having a high enough accuracy will give you a "good idea" of the interactions it might have with other proteins and substances, but can't account for the millions of other interactions they might have with other particles. Most of the bioinformatics aren't deterministic, but still rely on stochastic measures. DNA sequencing is done by sampling, then predicting the rest and as far as my own biology teachings go, it's categorized as "solved". Sequencing might get better, but we've accepted it as a solution for the moment. Arguing the term "solved" is just pedantic. We know it's not 100%, but the actual usefulness of improving the prediction of a few more Angstroms isn't going to make a huge practical difference. What matter is that we can start actually building tooling and lab tests on this method.
- sdenton4 6y agoShouldn't the next step just be predicting a distribution of possible structures, instead of just the most likely? That doesn't seem crazy given where we're at on the machine learning side; predicting distributions is almost the default at this point, in many areas.
- Scandiravian 6y agoThat is actually one of the "classic" ways of doing it. There's a bunch of different Bayesian models for protein structure - as an example, I implemented a model for structural changes in evolution for my master's thesis The big challenge for these kind of models is the curse of dimensionality. Since every atom in the structure can potentially interact with every other atom, it's tricky to make a joint distribution for the entire sequence and it's rare to have a model that's both accurate and parallelizable, so the field hasn't benefitted much from the advances in for instance GPU computing
- dekhn 6y agoThe protein folding problem isn't part of the body- it's a reduction in which the protein is folded in regular water solvent with no other components. I'm not sure what you mean by "DNA sequencing is done by sampling then predicting the rest". DNA sequencing works by oversampling and then making a "call" about the specific base in a position given the evidence. Regions without data are described as N with estimated length M, rather than a "prediction".
- xpe 6y agoIn the spirit of clarification, I want to share snippets from three people mentioned in the Nature article [1]: John Moult: ''' “This is a big deal,” says John Moult, a computational biologist at the University of Maryland in College Park, who co-founded CASP in 1994 to improve computational methods for accurately predicting protein structures. “In some sense the problem is solved.” ''' Andrei Lupas: ''' AlphaFold is unlikely to shutter labs, such as Brohawn’s, that use experimental methods to solve protein structures. But it could mean that lower-quality and easier-to-collect experimental data would be all that’s needed to get a good structure. Some applications, such as the evolutionary analysis of proteins, are set to flourish because the tsunami of available genomic data might now be reliably translated into structures. “This is going to empower a new generation of molecular biologists to ask more advanced questions,” says Lupas. “It’s going to require more thinking and less pipetting.” ''' Mohammed AlQuraishi: ''' “I think it’s fair to say this will be very disruptive to the protein-structure-prediction field. I suspect many will leave the field as the core problem has arguably been solved,” he says. “It’s a breakthrough of the first order, certainly one of the most significant scientific results of my lifetime.” ''' [1]: https://www.nature.com/articles/d41586-020-03348-4 https://www.nature.com/articles/d41586-020-03348-4
- xpe 6y agoI share these quotes because people don't have the same ideas about what "solving" protein folding means. I'd suggest finding clearer concepts. For example: * What degree of accuracy is attainable via each technique? * How much (a) wall clock time / (b) overall compute resources are required for each technique? * What use case(s) fit best with each technique? I offer these because I see a lot of energy expended in people digging in and defending their definitions, rather than understanding what other people mean.
- m00x 6y agoI think we can just go from prior works. DNA sequencing uses AI to sample the DNA at certain spots, then predict what the amino acids/genes will be from there. There aren't more significant research towards finding new ways of solving DNA sequencing since this method is good enough and can improve from more data and better models. We consider it "Solved" in this case. Tons of tooling was built on top of it and until we can get the true sequence of amino acids quickly and cheaply, it's not going away.
- Pensacola 6y agoThe aspect of this story that fascinates me is that one could argue that DeepMind has not, in fact, solved anything. An inscrutable, black-box ML model has. While such an algorithm could ideally predict protein folding in any case, it can't explain anything about why, and therefore can't really advance the science. An analogy might be that if you trained an AI model on billiard balls, it could become really good at telling you where a ball will end up when you hit it, but it could never tell you that the reason is that f=m*a, meaning it will do nothing to advance the science.
- roywiggins 6y agoPainstaking experimental ascertainment of protein structure doesn't really tell anyone "why" it ended up that way either, but it's still been worth doing and advances science a great deal. https://www.alpco.com/dorothy-hodgkins-discovery-insulins-3d-structure https://www.alpco.com/dorothy-hodgkins-discovery-insulins-3d...
- dekhn 6y agoRight. The challenging thing about protein folding is that we've been able to probe protein folding in numerous ways but none of them actually give a direct picture of what the process of folding looks like. That is, the specific physical configuration trajectories/envelope followed by an ensemble of folding proteins.
- deleted 6y ago[deleted]
- hajimemash 6y agoJust because you don't understand exactly how something works, don't mean it's not useful. Someone who doesn't understand exactly how everything works, like how do bikes stay upright and how does gravity work, can still use the bike to acquire food or advance science.
- Barrin92 6y ago>Just because you don't understand exactly how something works, don't mean it's not useful The OP didn't say that it is not useful, what they implied was that it is not actually science, which is correct. Science is a system that produces and organises knowledge. Chomsky made this point many years ago in a similar debate in linguistics. Statistical learning might produce results, but it tells us virtually nothing about the underlying laws or structures that govern language use. ML in its current from is effectively the modern version of behaviourism and will, or already does, suffer from the same issues.
- ampdepolymerase 6y agoIn practice it doesn't matter. What Deep Mind has done far outstrips previous results. Now that we know neural networks for protein predictions is not a dead end, the accuracy will easily improve with time.
- Yajirobe 6y agoYes, we will just map out/prepare more training data, train better networks. Prepare more and more training data... until we have mapped out all proteins known and there is nothing left to predict for the networks...
- chorsestudios 6y agoThe 2020 AlphaFold team mapped OFR8, a protein associated with COVID-19. It seems unlikely that mapping out everything we currently know of would be the end of their research. COVID-19 is probably not the last deadly pathogen we will encounter.
- whoisburbansky 6y agoThis is exactly how I feel; protein folding is almost chaotic in the sense that tiny differences in specific atomic locations end up having huge functional impacts, which is completely unlike, say, neural machine translation, where a slightly garbled translation is still intelligible for the most part. I don't quite see how this approach to protein folding helps if you can't actually be sure about the predicted structure's functionality without doing the expensive experimental verification.
- foota 6y agoI think part of the answer here is that you can more easily verify than you can determine. Are x y and z where we expect them? Yes? Looks good.
- whoisburbansky 6y agoExcept you can’t, right? Figuring out where we expect them to be involves finding out where they are in the first place.
- anthony_doan 6y agoYeah a bit too much hype around ML. I was in my org conference and the guy that head AI team was stating how Deep Learning going to change the world. That it'll write software in a few years (he's from the applied math field). He also glossed over many on going problems with Deep Learning. We're in healthcare industry and this guy is pushing Theranos like level of snake oil. The people that don't know a lick about ML or Statistic or both rely on these people. The doctors are relying on that dude and I for these inputs. And the dude is selling theranos like stuff. I think the problem is the medical field, cs, and stat are all high level field, requiring years of training to acquire mastery and knowledge. So it's rare that someone would have all three to be able to have a impartial view or a good overview of pros and cons.
- king_magic 6y agoHuh? Who exactly is pushing "Theranos-like level of snake oil" here? What "dude" are you referring to?
- deleted 6y ago[deleted]
- kevinskii 6y agoMohammed AlQuraishi elaborated on this in his 2018 "What Just Happened?" blog post that was discussed on HN when the news was announced. [1] It's a fairly long read, but he goes into a lot more detail as to why the problem isn't yet solved. Interestingly, he also notes that pharma and academia should feel some embarrassment from DeepMind's achievement: "What is worse than academic groups getting scooped by DeepMind? The fact that the collective powers of Novartis, Pfizer, etc, with their hundreds of thousands (~million?) of employees, let an industrial lab that is a complete outsider to the field, with virtually no prior molecular sciences experience, come in and thoroughly beat them on a problem that is, quite frankly, of far greater importance to pharmaceuticals than it is to Alphabet. It is an indictment of the laughable “basic research” groups of these companies, which pay lip service to fundamental science but focus myopically on target-driven research that they managed to so badly embarrass themselves in this episode." [1] https://moalquraishi.wordpress.com/2018/12/09/alphafold-casp13-what-just-happened/ https://moalquraishi.wordpress.com/2018/12/09/alphafold-casp... (Edited to clarify that the blog post was not recent.)
- aazaa 6y agoTo summarize: 1. "...only two-thirds of DeepMind’s solutions were comparable to the experimentally determined structure of the protein. ..." 2. "... the average or root-mean-squared difference (RMSD) in atomic positions between the prediction and the actual structure is 1.6 Å (0.16 nm). That’s about the size of a bond-length." 3. It may be "... more difficult to predict the structures of proteins with folds that are not well represented in the database of solved structures." 4. "... the method cannot yet reliably tackle predictions of proteins that are components of multi-protein complexes."
- screye 6y agoA couple of other valuable perspectives: https://explainthispaper.com/ai-solving-protein-folding/ https://explainthispaper.com/ai-solving-protein-folding/ Another supporting article from Derek Lowe's (think Medical Science's Stratchery. Highly acclaimed and usually cynical) blog : https://blogs.sciencemag.org/pipeline/archives/2020/11/30/protein-folding-2020 https://blogs.sciencemag.org/pipeline/archives/2020/11/30/pr...
- screye 6y agoI wonder if these select industry labs with full creative freedom (DeepMind, Brain, OpenAI) will end up being a Xerox Parc/ Bell labs moment for our generation. A group of incredibly intelligent people allowed to do whatever fundamental research they want at the expense of parent company's business, whose value will not be appreciated for a couple decades until everything new in technology points back to a fundamental innovation coming out of that lab.
- fock 6y agonow that we have this inevitable discussion venue after everyone has digested the PR (and some of the people her might have even had a chance to ask some questions to deepmind) some remarks (questions - if anyone from the team is here): - I thought that open code was becoming the standard and also being pushed by Google? But apparently this does not apply to deepmind, because of $$$s? For the original Alphafold (which actually was 3 models) no code is available except, where it had to be (their nature publication on one of the three models used in the competition). - why did they not participate with all the available proteins? I guess it's some loophole in the rules, to allow for greater "improvements" even when models are not super general, but from a naive scientific view that is absolutely stupid. - maybe they went fully generative this time, but in CASP13, 2/3 of their models were just blackbox-predictors, which they optimized with simulated annealing. Given that the configurational sample space for the protein is huge, doing that seems still rather costly. I wonder how the actual spectra fitting works and compares to that and why experimentalists could not go this route as well? (just do simulated annealing, until the spectrum fits). - they already trained on all known proteins, yet with some they are still far off. Seems like it's not solved really, though results are certainly impressive and it could be a great tool for any person interested in frozen protein structures.
- sandGorgon 6y agodoes anyone know where the latest code for this is ? i can see CASP 13 is here - https://github.com/deepmind/deepmind-research/tree/master/alphafold_casp13 https://github.com/deepmind/deepmind-research/tree/master/al...
- fock 6y agothis is not CASP13. It's one of 3 deepmind models from CASP 13 (and doesn't contain any feature-related code). And not even the most interesting imho (Variational LSTM/CNN-Autoencoder which generates structures from sequences). Just be aware that they are doing this for $$$s and PR - not for the good of the people like their website suggests - although they really try to put a smokescreen around this.
- dnautics 6y agoJeez. Clickbait title, tons of fallacy, like goalpost moving. Some context on myself: I have 15+ years of postgraduate chemistry/molecular bio/biophysics/biochemistry experience, then quit to go into tech, worked for AI hardware and now AI-driven software startups, mostly as backend (not implementing ML models, but I know how to do that and have done some small models for myself). I'm a pessimist about this AI innovation cycle's prospects in general AI and in particular drug discovery. Has protein folding been solved? Yes. As a practitioner, what I want to be able to do is pull a sequence that I've retrieved from DNA, drop it into a computer, and get the structure out. Intuitive insight can FOLLOW from those results. For example: In one of my projects, I was able to look at the structure (a homolog had been solved, so I just did a dumb alignment and threading), identify that it was acting as an NPN transistor (unpublished details), fix the electron flow through the enzyme and improve the yield (https://link.springer.com/article/10.1186/1754-1611-7-17 https://link.springer.com/article/10.1186/1754-1611-7-17). Later, I looked at the structure of the enzyme, modified the surface charge in one particular part of it, and improved electron throughput again (https://www.mdpi.com/1422-0067/16/1/2020 https://www.mdpi.com/1422-0067/16/1/2020). This was with primitive "protein folding through homology" tools, now there's an good chance I could do these sorts of things with proteins I don't have homologous structure for. These are the sorts of things that protein folding enables. One more thought -- I bet that DeepMind can do things like make it obvious where there are certain posttranslation sites (like FeS clusters) or make it obvious where there is a cryptic phosphorylation or glycosylation site (sequence holdover from a previous mutant that no longer has its expected functionality) because it's been buried. Will there be corner cases? Yes. Probably deepmind will have difficulty solving the fold of amyloid fibrils. Probably deepmind will have some difficulty with super-strange post-translational modifications (think things like GFP's core fluorophore), or if you design a drug where you splice in an unnatural or D-amino acid, or oddities like that. > we are not at the point where this AI tool can be used for drug discovery I disagree. Sure, it probably won't be able to find a small molecule binding site. It almost certainly won't be able to design a drug that has a long-range allosteric effect (think Gleevec's super strange mechanism of action). But, deepmind WILL be able to help design biologics that, for example, can interact with bump-hole mutations. There was never an expectation that "protein folding" solves every problem in the drug discovery pipeline. That's out of scope for the basic problem. As for this: > AI methods rely on learning the rules of protein folding from existing protein structures. Come on. There is no method that doesn't rely on learning the rules of protein folding from existing structures. Even de novo MD-modeling has tweaking fudge factors (we could call it "dark biochemical fields" -- think: what is the expected dielectric constant around a tyrosine residue?? No way we're calculating that from the schroedinger equation) that are empirically derived to get your results.
- ramraj07 6y agoAnother academic professor trying to undermine something that just makes their entire way of doing things look stupid, that's all there is. When the human genome was sequenced another entrepreneur came in (Venter), said "You guys are morons to spend billions over decades, let me show how an actual smart team would do it" and bet them to it in a fraction of time and cost. Yet the consortium that spent billions on the human genome project still got congratulated. What's ironic is that when they said they sequenced the human genome they still had as many asterisks to that statement as this broad DeepMind statement has. In fact, the sequencing of the human genome claim was probably more disingenuous than DeepMind because, and this is important, they didn't make up this definition - the CASP organizers did. This team just met a pre defined standard as what's accepted as "solving of protein folding" and this brilliant team met this challenge. Academics calling this hyperbole should fire every university's PR team because the amount of hyperbole they add to every press release about a paper where they "cured" cancer in mice is 10x larger than this. Academia is fundamentally broken; the cracks started appearing in the sixties (read Hammings lecture notes), and have all but metastatasized throughout, especially in biology. We are all dying faster because of this (Google the Alzheimer's cabal). Academia is now just bunch of overperforming hacks who are honestly not that good at anything except sitting in circles in NIH grant review panels giving millions to each other while giving "constructive criticism" like this to what is clearly a monumental achievement, because if they don't then they reinforced how useless they are together as a group. Going back to this hacky article, of course this is the first step, and there is much to be ironed out. But in the words of Sydney Brenner, a great scientist from a time when actually smart, humble people became professors, "The entry of large numbers of American ... into the field will ensure that all the chemical details ... will be elucidated." [1] We are definitely now in a fairly deterministic path towards figuring out protein folding for all practical purposes. The few academics who still have humility and foresight see this. DeepMind is obligated to release enough results and there are still some sane minds left in academia that they will do what they are okay at, which is filling in the details. [1] http://nemaplex.ucdavis.edu/General/Biographies/SBrenner.htm http://nemaplex.ucdavis.edu/General/Biographies/SBrenner.htm
- aurizon 6y agoAs dnautics implies, this is like you make a model of a mousetrap - along comes a mouse - snap, what is the shape of the closed mousetrap? By this I mean proteins fold into a shape, and DM gives us some idea what that looks like, then it enters into a reaction - what is the shape now? Many are enzymes and snap and release, snap and release. Some poisons have the same initial shape - snap - then no release - reaction site has been poisoned - the body thinks it has lots, but the body dies of the poison. Some toxins that kill by ribosomal inactivation are like this, mushroom poisons, ricin are like this. Mankind would love to be able to reverse ricin and mushroom poisonings, which usually happen days after the toxic event. Not being an expert, this is just an analogy.
- the__alchemist 6y agoWe can't consider this solved, or that we have an understanding of how the process works until we can do it ab-initio. Improvements in machine-learning and other experiment-based algorithms are great, since they're the best we have. We need a breakthrough in chemistry and the ability to solve the Schrodinger equation in 3D (Or something similar) to truly solve this. Ie generate the electronic structure statically without using arbitrary, tuned constants, then evolve it over time. We know the rules and constraints, but unfortunately, can't solve it without using approximations, and fitting to experimental data. Machine-learning approaches will always suffer from over-fitting; they can produce practical results for known cases, but their predictive power is limited. (But still impressive!)
- MiddleWinger 6y agoYes, the did.
- openasocket 6y agoThere's one thing this critique didn't entirely make clear to me, hoping someone here could answer. Is there a relatively efficient algorithm to tell that a particular folding solution is correct (or within a certain distance of correct)? If there was, at least then this AlphaFold would either tell you "here, this is more or less what this protein folds into" or "I can't figure out the folding". Even if it told you it can't figure out the folding a large portion of the time, at least it isn't giving you garbage data.
- m3at 6y agoI see a lot of negative reactions in the comments, but at least this part seems fair to me: > That advance will be much clearer once their peer-reviewed paper is published (we should not judge science by press releases), and once the tool is openly available to the academic community DeepMind has a tendency for hyper inflated PR (not the only ones mind you), wanting for the scientific process to run its course before claiming victory sounds good to me.
- nl 6y ago> DeepMind has a tendency for hyper inflated PR Can you point at some? AlphaZero (emphasis: AlphaZero, not AlphaGo) is probably the most significant AI breakthrough of the last 25 years (maybe more - I think it's more significant than AlexNet on ImageNet) and there was very little hype about it. AlphaGo got quite a lot of PR, but a lot of that came from the Go community, especially in Korea.
- marco_craveiro 6y agoComplete lay person here, so please bear with me if the question is very silly. Can this help the work on Perovskite solar cells [1] by any chance? I ask because it appears AlphaFold has applications in crystallography and (to the lay person) it seems that finding the right crystal structure is key in making a Perovskite solar cell that can last for decades. [1] https://en.wikipedia.org/wiki/Perovskite_solar_cell https://en.wikipedia.org/wiki/Perovskite_solar_cell