52 ms·
AlphaFold: a solution to a 50-year-old grand challenge in biology
- nmca 6y agohttps://deepmind.com/blog/article/alphafold-a-solution-to-a-50-year-old-grand-challenge-in-biology https://deepmind.com/blog/article/alphafold-a-solution-to-a-...
- partingshots 6y agoI continue to be impressed by how quickly DeepMind has managed to progress in such a short time. CASP13 was a shocker to all of us I think, but many were skeptical as to the longevity of the performance DeepMind was able to achieve. I believe with CASP14 rankings now released, it's safe to say that they've proven themselves. Congratulations to the team! This work will have far reaching impacts, and I hope that you continue to invest heavily in this area of research.
- whimsicalism 6y agoProgress like this was, in my view, inevitable after the invention of unsupervised transformers. It'll be genetics next. e: although AlphaFold appears to be convolutionally based! I suspect that'll change soon.
- klmr 6y ago> It'll be genetics next. Which part of genetics are you thinking of? Much of genetics isn’t amenable to this kind of ML, because it isn’t some kind of optimisation problem. And many other parts don’t require ML because they can be modelled very closely using exact methods. ML does get used here, and sometimes to great effect (e.g. DeepVariant, which often outperforms other methods, but not by much — not because DeepVariant isn’t good, but rather because we have very efficient approximations to the exact solution).
- whimsicalism 6y agoWhat do you mean? Genetics is amenable because the genome is a sequence that can be language modeled/auto-regressed for depth of understanding by the network. There are plenty of inferences that you would want to do on genetic sequences that we can't model exactly and there is some past work on doing stuff like this, although biology is usually a few years behind. https://www.nature.com/articles/s41592-018-0138-4 https://www.nature.com/articles/s41592-018-0138-4 e: for clarity
- garmaine 6y agoThis is word salad.
- whimsicalism 6y agoRude. I would appreciate substantive criticism, especially when I'm linking papers in Nature starting to do exactly what I'm talking about.
- garmaine 6y agoI cannot give constructive feedback to something which is incomprehensible. "the genome is a sequence that can be language modeled/auto-regressed for depth of understanding by the network" The genome is not a sequence so much as a discrete set of genes which are themselves sequences which specify construction plans for proteins. That distinction is important. Language modeling in the context of machine learning typically means NLP methods. Genetics is nothing like natural language. Auto-regression is using (typically time series) information to predict the next codon. This makes very little sense in the context of genetics since, again, the genetic code is not an information carrying medium in the same sense as human language. Being able to predict the next codon tells you zilch in terms of useable information. "Depth of understanding by the network" ... what does that even mean??? The above sentence is a bunch of popular technical jargon from an unrelated field thrown together in a nonsensical way. AKA word salad.
- alquemist 6y agoFWIW, transformers is to sequences what convnets is to grids, modulo important considerations like kernel size and normalization. Think of transformers as really wide (N) and really short (1) convolutions. Both are instances of graphnets with a suitable neighbor function. Once normalization was cracked by transformers, all sort of interesting graphnets became possible, though it's possible that stacked k-dimensional convolutions are sufficient in practice.
- whimsicalism 6y agoI work in the field, I don't need the difference explained to me. > Think of transformers as really wide (N) and really short (1) convolutions Modern transformer networks are not "really short" and you're also conflating the difference between intra- and inter- attention. There is still a pitched battle being waged between convnets and transformers for sequences, although it looks like transformers have the upper hand accuracy wise right now, convnets are competitive speed-wise.
- Veedrac 6y ago> e: although AlphaFold appears to be convolutionally based! I suspect that'll change soon. “For the latest version of AlphaFold, used at CASP14, we created an attention-based neural network system” ?
- the8472 6y ago> but many were skeptical as to the longevity of the performance DeepMind was able to achieve For a non-biologist, on what is this skepticism based? Just purely based on following ML news it looks like the trend for ML solutions has been that they've overtaken expert-systems once they've gained a solid foodhold in a field. Maybe this is some perception bias. Are there any cases where ML performed decently but then hit a ceiling while expert systems kept improving?
- whimsicalism 6y agoML is a super overloaded term. There are definitely cases where machine learned statistical solutions do not perform as well as the systems tuned by the experts, but if you can define the task well and get the data for a deep solution, usually those will overtake.
- penagwin 6y agoThis. I believe technically just linear regression could be considered "machine learning".
- misnome 6y agoI've seen people at bio conferences actively calling linear regression machine learning.
- diab0lic 6y agoThis is likely because linear regression meets most widely accepted definitions of machine learning. [0][1] It is simple and very effective when learning in linear space. [0] https://en.wikipedia.org/wiki/Machine_learning https://en.wikipedia.org/wiki/Machine_learning [1] https://www.cs.cmu.edu/~tom/mlbook.html https://www.cs.cmu.edu/~tom/mlbook.html
- b3kart 6y agoSorry, I don’t get it. Are you saying fitting a linear regression model to data and making predictions somehow isn’t machine learning? I am confused.
- deleted 6y ago[deleted]
- lucidrains 6y agoAmazing day for structural biology! If it weren't for the pandemic, I would be out at the bars celebrating tonight!
- postingpals 6y agoHeh, soon you'll be able to do that too when the vaccine comes out. What a great end to the year.
- xyzal 6y agoDoes it mean there is no point in playing fold.it anymore?
- breck 6y agoYes, no point, as far as I understand it.
- IgniteTheSun 6y agoConsidering the resource requirements for this AI approach mentioned in the article, its unlikely that its been tested on more than a few tens to hundreds of proteins. This may only work on a subset of the proteome so I would think it worth it to continue playing if you find it to be a fun past-time.
- nynx 6y agoThose were the requirements for training it.
- hobofan 6y agofold.it was always more geared towards being edutainment than actually contributing solutions. Of the ~20 publications made related to fold.it over a decade, ~5 of them seem to have contributed to solving structures, while the rest of them are about the game itself.
- flobosg 6y agoBesides structure prediction, Foldit is used for the inverse problem: protein design.
- m-p-3 6y agoI'm wondering what this means for folding@home.
- breck 6y agoI'm pretty sure this means they can pack it up? Or point their infra to a different problem?
- touisteur 6y agoOr just do billions of inferences per second. Next step?
- great_tankard 6y agoFolding@home mostly tries to calculate protein dynamics using already solved structures, so their work is still critical.
- matsemann 6y agoI came here wondering the same. Is this based on work done by folding@home for instance? (As in, it used their precomputed stuff as training data)
- chetan_v 6y agoFirst Nobel prize for AI from this?
- xgulfie 6y agoHopefully they give one out for this, if only so I can say I'm a Nobel Prize contributor
- TRcontrarian 6y agoNo way, you were on the team? Congrats.
- aardvarkr 6y agoWe’ll know ten years from now
- harperlee 6y agoNot knowing a lot about biotechnology, I read the article and it sounds great, but how big is this as a gamechanger? Can someone comment on how big are the implications of this in, let’s say, 5 years from now, on day to day life? Does this mean that biotech is going to explode? Or just that drugs will come to market faster, perhaps cheaper for rare diseases, but from the same industry structure as always?
- candiodari 6y agoThis will allow us to discover much more about the structure of the cell (of "life") at a before this unprecedented speed. We should find many, many more mechanisms and targets for medicine, but it takes 10-20 years to bring a new medicine to market. So in 5 years you'll see exactly zero new medicines pop up.
- piyh 6y agoNo new medicines, but way more biotech tools. Higher yield GMO plants, foundational research into disease, science backed recommendations for lifestyle changes to avoid disease that previously eluded us, some crazy stuff happening in animal models. The progress in biotech the past 20 years makes moore's law look slow.
- pmastela 6y agoI agree. The main inhibitor of speed that products of this advancement will be deployed at will likely be determined by local policies. Though, given just how profound some of the impacts on medicine might be, the speed at which they can be deployed might become a matter of national security (a healthier population bodes well for a healthier economy which in turn strengthens national security). Hopefully this competition shortens the time-to-market for all these new medicines.
- fabian2k 6y agoProtein folding is a big and important problem, so this is certainly big news if it works as well as it seems. But I wouldn't assume that this changes everything, we can already determine how proteins fold by experimental work. The disadvantage is that this is a lot of work, though the methods there also improved a lot. One question is how robust the predictions are that DeepMind produces. I would also assume that right now it can't e.g. determine protein structures in the present of other small molecules, or protein complexes. A lot of the interesting stuff lies in the interactions between molecules. And in general in life sciences any new development will take at least a decade until it hits day to day life, likely even more. We're living with a exception to this rule right now due to the pandemic, but in general things take quite a bit of time in that space.
- seek3r 6y agoKudos to DeepMind. I’m eager to read their paper.
- breck 6y agov2 looks amazing. that jump is even more incredible than the first. More context from v1 in 2018: https://moalquraishi.wordpress.com/2018/12/09/alphafold-casp13-what-just-happened/ https://moalquraishi.wordpress.com/2018/12/09/alphafold-casp...
- TrackerFF 6y agoWhat are the immediate real-world applications of this? Just asking, because I have very little knowledge in this area.
- candiodari 6y agoGiven the DNA code for one of the "machines" that run cells, we can generate an atomic model of that machine. This means we can "compile" (one part of) the DNA code. It was already possible, but so slow that entire datacenters would spend months calculating this for a single protein and even then we can't use them on the really complex ones at all, necessitating things like neutron spectroscopy which are totally insane, and only work on like 1% of proteins. This is useful because for example chemical simulation tools don't run on DNA code, but on atomic models. And also to produce "images" of the molecules (images between quotes because most proteins are too small to interact with reasonable photons, and no interaction with photons means you can't see them in any way) DNA has other parts that are really important but we don't understand at all yet, where this doesn't help at all. This applies to sections of DNA sent to ribosomes, to produce actual molecules. Besides that, there are pieces of DNA that "index" the DNA, pointers (from one gene to another), triggers (that for instance start production of an enzyme based on some external influence, like detection of a marker molecule) and export markers (that tell you what to do once the protein is produced, for example, mark a protein to be removed from the cell, incorporated into the cell membrane, or for instance used inside the cell nucleus, and there's also one that essentially says "at this point stop producing a protein and instead couple the rest of the DNA code to the end of the protein you just made").
- empiricus 6y agoI am actually scared. This plus CRISPR means real nanotechnology is within reach.
- jcims 6y agoI think this is the interesting part because there aren't going to be the same regulatory hurdles for using ribosomes to manufacture technology as there are for medicines. Synthetic organelles that weave fibers, build metamaterials, etc could lead to pretty magical advances in our capability.
- entropicdrifter 6y agoPerhaps we'll live to see The Diamond Age
- wrinkl3 6y agoCan't wait to join a distributed computing bacchanalia.
- dynamite-ready 6y agoFar from an expert here, but your comment makes me think of Michael Crichton's 'Prey', if you've not already read it. Not that I wish to add to your apprehension.
- enchiridion 6y agoMy thought as well. I wonder what the world will look like in 20 years because of this. I'm willing to bet it will be staggeringly different than what most people are expecting.
- marcosdumay 6y agoThere is still at least one NP-hard problem on the way, that is creating a protein with a desired format.
- deleted 6y ago[deleted]
- echelon 6y agoThis sounds wonderful and frightening. On the one hand, now we can engineer drugs at light speed. But wasn't protein folding supposed to be NP-hard? Can deep learning find the cracks in P vs NP? Perhaps making clever guesses at prime factors because it learned some weird structural fact that has eluded mathematicians. If we break crypto, there goes the modern world. Banks, bitcoin, privacy, Internet, the whole shebang. (I obviously am not an expert in computational complexity and hope that some domain experts can chime in and assuage my fears.)
- karl-j 6y agoI think I'm almost as uninformed as you, but I believe it comes down to the difference between perfect solutions and close enough solutions. Consider the classic NP problem of the traveling salesman problem. "[Modern heuristic and approximation algorithms] can find solutions for extremely large problems (millions of cities) within a reasonable time which are with a high probability just 2–3% away from the optimal solution." [0] When close enough is enough, NP problems can often be solved in P time, and I suspect this is one of those cases. For crypto however, close enough is not enough. [0] https://en.wikipedia.org/wiki/Travelling_salesman_problem#Heuristic_and_approximation_algorithms https://en.wikipedia.org/wiki/Travelling_salesman_problem#He...
- aparsons 6y agoThere is probably a team at DeepMind working on cracking simple crypto. Problem is, it can be difficult to cast the problem properly/“correcty”. How does a one way function get represented?
- ichbinwiederda 6y agoFar from an expert on complexity theory, but NP-hard problems can be approximated in polynomial time. With Deep Learning you are doing approximation. So this is nothing ground breaking in that respect.
- Vervious 6y agothere are also a variety of problems that are hard to approximate.
- unchocked 6y agoBeen out of the field for a while, could someone currently in it qualify these results? Hyperbolic title notwithstanding, they approach 90% median free modeling accuracy. The "other 90%" still remains to be solved...
- asdfasgasdgasdg 6y agoI don't think anyone on HN is going to have more authority to qualify the results than the independent experts quoted in the linked article. Among whom are numbered a Nobel laureate, the president of the group that designs the tests of protein folding systems, and the former CEO of Genentech+current CEO of Calico.
- dekhn 6y agoArt's a smart guy and I have a lot of respect for his biological intuition, but his understanding of computational biology is very limited.
- asdfasgasdgasdg 6y agoI would imagine that he is not assessing this advancement merely using his own personal expertise, but rather the combined expertise of the resources he represents. CEOs don't just look at problems and potential solutions. They have people who look at those things, and then tell them their opinion. In any case, you've picked a nit with one of the three people quoted. Any objections to the other two?
- dekhn 6y agoMy main objection to Vivek (the Nobel Prize winner) is the prize in that case should have gone to my advisor, Harry Noller. John Moult... he's a nice guy but I think he's being a bit breathless here.
- asdfasgasdgasdg 6y ago
- _greim_ 6y agoAt Sun back in the day our workstations tended to have fairly promiscuous login settings, so one of my coworkers took the liberty to launch folding@home on every machine in the org. Listing running processes one day, I saw this thing pegging my CPU; asked around and others had it too. A virus!?! Then he fessed up. Kinda miffed at first but ultimately really cool, so we let the thing keep running. That was my introduction to the whole protein folding problem, and it's really great to see this milestone!
- dekhn 6y agoI ran Folding@Home at Google on hundreds of thousands of fast Xeon cores for over a year. I concluded at the end that unbiased MD simulations are not an effective use of computer time.
- xbmcuser 6y agoThis is a lot bigger than people are assuming if protein folding can be done quickly and cheaply it will trickle down to a lot more than medicine. It is going to advance bio fuels, food production and a lot more.
- CJefferson 6y agoHas anyone got any good other references for this? After some of the dodgy experiments related to alpha zero (comparing to purposefully degraded chess systems), I'd love to see some independent analysis.
- sanxiyn 6y agoCASP is that independent analysis...
- CJefferson 6y agoTrue, but I haven't seen an independent discussion of the CASP results. There is a good chance this is great, but I don't trust deepmind press releases.
- andi999 6y agoI am also wondering. I generally find these kind of approaches hard to believe, but this might be my prejudices.
- syncsynchalt 6y agoThe article in Science implies that we have independent confirmation of predictions yielding useful results, beyond the challenge itself: > The organizers even worried DeepMind may have been cheating somehow. So Lupas set a special challenge: a membrane protein from a species of archaea, an ancient group of microbes. For 10 years, his research team tried every trick in the book to get an x-ray crystal structure of the protein. “We couldn’t solve it.” > But AlphaFold had no trouble. It returned a detailed image of a three-part protein with two long helical arms in the middle. The model enabled Lupas and his colleagues to make sense of their x-ray data; within half an hour, they had fit their experimental results to AlphaFold’s predicted structure. “It’s almost perfect,” Lupas says. “They could not possibly have cheated on this. I don’t know how they do it.”
- CJefferson 6y agoThanks, that really is convincing.
- piva00 6y agoThis sounds big, like really really big. At least from my old times providing my idle computing resources to Folding@Home and following that project, this seems like the major golden milestone for protein folding.
- FrojoS 6y agoExactly what I was thinking. In a very small way many of us tried to help with this problem back in the day. Makes it feel even more important. Now I'm waiting for the equivalent news about SETI@Home ;-)
- iandanforth 6y agoTitle as submitted is hyperbole, please fix?
- breck 6y agoI don't think it is. Look at the graph.
- sanxiyn 6y agoIt is not a hyperbole.
- EgoIncarnate 6y agoI agree. "AlphaFold achieves a median score of 87.0 GDT". While this is a major advance, to me 100 GDT would be 'solved', not 87.
- ashtonbaker 6y ago> To me Are you a domain expert? Because: > According to Professor Moult, a score of around 90 GDT is informally considered to be competitive with results obtained from experimental methods.
- The_rationalist 6y agobut experimental methods have not solved protein folding either. AlphaFold has'nt solved protein folding but I can't wait to see their progress for ALphaFold 3. What would be informatively useful would be to know how much accuracy is needed on average for drug engineers, I'd say that 99% is more likely to be the minimum to make solid inferences
- ashtonbaker 6y ago> but experimental methods have not solved protein folding either. I might be missing something here, but isn't "experimental methods" just shorthand for "our best knowledge of a protein's structure, obtained via NMR or X-ray crystallography"? In that case, I'm not sure what "solving" protein folding even means - literally zero mean error? We can't know/solve anything beyond our best knowledge, that's tautological. > What would be informatively useful would be to know how much accuracy is needed on average for drug engineers. Yeah that would be interesting, but: > I'd say that 99% is more likely to be the minimum to make solid inferences ...what are you basing this on?
- Rochus 6y agoGreat. So then farewell CARA (http://cara.nmr.ch/doku.php http://cara.nmr.ch/doku.php), we had a good time.
- Rochus 6y agoAfter I had some time to think about it, I come to a different conclusion. Contrary to my first assumption, Bio NMR (in contrast to crystallography) will become more and more important, since the method allows to study the dynamic properties of proteins. With the structure predicted by DNNs, the chemical shifts to be expected in the NMR spectra can be calculated; the assignment problem is thus largely eliminated. Bio NMR can then be used specifically to study the "parts that move".
- mabbo 6y agoSometimes announcements like this are a bit over-the-top. But what really, to me, cements the 'big-deal' of this is the "Median Free-Modelling Accuracy" graph half way down the page. Scores of 30-45 for 15 years. Now scores of 87-92. This isn't a minor improvement, it's a leap forward.
- entropicdrifter 6y agoNot to mention the fact that two years ago they took it from 45% to >60%. If they can continue improving, even with an exponential decay in rate of improvement, this is certainly a stunning example of technological disruption.
- Zenst 6y agoEven without any improvement, the amount of grunt-work the AI can pre-do and get down to a short-list - that in itself will see changes in progress speeding research up.
- kordlessagain 6y ago> and get down to a short-list There's no reason to believe the list will contain all solutions, however.
- patagurbon 6y agoNo but it will hopefully contain some. Which for many if not most problems is all that matters
- treis 6y agoThat is an impressive improvement, but I think you've missed the most important point: >a score of around 90 GDT is informally considered to be competitive with results obtained from experimental methods So DeepMind is to the point where it's a question of whether their generated model or the experimentally determined structure is closest to the actual physical structure.
- phonebucket 6y agoThis is a huge jump forward. Last year's performance already was a big step up over the previous, and this seems to go much further. So big kudos to the research team. Nonetheless, I'd like to hear more from specialists outside the context of a marketing blog post before I fully buy into a claim of a solution. There's also a rabbit hole about what 'solution' actually means. Is the performance sufficient for any protein folding prediction application that might arise in the future?
- yarabarla 6y agoMan, I remember running folding@home years ago on my terrible laptop. Now this was done with what they say is equivalent to only 100-200 GPUs. Crazy to see how far we've come in just a short amount of time.
- sumtechguy 6y agome too... should have done bitcoins :)
- jjk166 6y agoNow onto the much harder problem of doing the reverse: taking an arbitrary structure and determining an amino-acid sequence that will fold into it.
- Rochus 6y agoWhat for?
- jjk166 6y agoThe forward folding problem lets you determine structures from a known genetic sequence. So for example you could very quickly sequence the genome of a virus and figure out how it worked much faster than current methods allow. The reverse folding problem lets you specify a structure and then make a genetic sequence to produce it. For example you could look at this virus to see how it infects its host, then design a custom protein to act as an anti-body stopping it, which is a capability we don't currently have. Forward folding is certainly useful, but reverse folding would be revolutionary.
- Rochus 6y agoThe set of all proteins which can potentially be expressed in an organism is known. Now maybe we also get decent (static) structure information for these. But the interaction of a virus with the host cell is much more complex. There is much more than just an amino acid sequence involved. And these parts are all moving, so a static picture as we now can create faster than before does not contain all the information necessary to fully understand the functions.
- woeirua 6y agoThis is a big step forward, but the outstanding question as far as to whether or not this is useful for evaluating novel proteins, is going to be how good is the confidence metric at telling the user to trust or not trust the results. You can see from their examples, that AlphaFold is very good but not perfect. I imagine for some proteins it will still give misleading or erroneous results and if you can’t tell when that happens without verifying the structure experimentally then this will likely not be that useful for new science.
- asdfasgasdgasdg 6y ago> the outstanding question as far as to whether or not this is useful for evaluating novel proteins That is not an outstanding question. The test on which DeepMind scored high marks is a test of how well the algorithm folds novel proteins -- proteins whose ground-truth structure has not yet been published.
- woeirua 6y agoWe’d have to see the distribution of GDT scores evaluated on unknown proteins to say anything about how confident we can be. If the distribution is tightly distributed around the median then great, this works really well. If the variance is large though then you’re going to have a hard time using this for meaningful predictions.
- foota 6y agoAccording to the article there's a confidence score as well. As long as this is sufficiently predictive of errors either a tight or wide distribution is likely acceptable.
- woeirua 6y agoWe need to see the relationship between confidence and GDT score. If you have a nice relationship then again everything is great. But... most confidence metrics from neural networks do not have a nice relationship to the primary metric.
- aparsons 6y agoFascinating work. I wonder if this approach works to model interactions (no reason it shouldn’t). The interactions of proteins with other proteins and well as as molecules like lipids, water and electrolytes form the basis for cellular processes. If that can be inferred correctly, you are looking at the building blocks of a “human simulator”.
- elevenoh 6y agoSo the median accuracy went from ~58% (2018) to 84% (2020) in 2 years? Does 84% == solved? Also, any low hanging frut implications for longevity tech?
- dekhn 6y ago100% accuracy is "solved".
- randcraw 6y agoSolving the inverse problem would be even more valuable -- given a specific shape (and other biochemical desiderata), what sequence of amino acids would create that protein? As hard as the protein folding problem is, the inverse problem is harder still. THAT is the one true grail.
- dekhn 6y agoWe "solved" this at Google years ago using Exacycle. We ran Rosetta (the premier protein design tool) at scale. The visiting scientist (who later joined GOogle and created DeepDream) said it worked really well "I could just watch a folder and good designs would show up as PDB files in a directory".
- gfodor 6y agoYou can't get 100% accuracy on something for which you don't or can't know the ground truth.
- dekhn 6y agoThe protein folding problem is predicated on the idea that there is a ground truth (a single static set of atomic coordinates with positional variances). If your point is that even experimental methods can't truly reach 100% (due either to underlying motion in the protein, or can't determine the structure), that's more or less what Moult is saying (they more or less arbitrarily define ~1A resoution and GDT of 90 as the "threshold at which the problem is solved").
- ampdepolymerase 6y ago@dang, please combine the thread with https://news.ycombinator.com/item?id=25253488 https://news.ycombinator.com/item?id=25253488
- lawrenceyan 6y agoEarlier post on this with direct results: https://news.ycombinator.com/item?id=25253488 https://news.ycombinator.com/item?id=25253488
- jeffbee 6y agoPretty interesting that they only used about $15k worth of resources (retail price) to achieve this. It's not a technique that would have been out of reach for other organizations based only on not being able to afford the compute.
- mdjt 6y agoBased on the going rate of a 32-core TPUv3 slice ($32/hr USD) running "for a few weeks", isn't this closer to $65k USD?
- entropicdrifter 6y agoOne could buy 200 GPUs for cheaper, I think that's where the other comment's price estimate came from.
- jeffbee 6y agoIt says $1,752/mo for v3-8, so I just multiplied it 8x.
- mdjt 6y agoFair enough, that calculation is still a bit off if they used 128 cores (16x instead of 8x). Not that it really matters...
- allenz 6y agoThat’s only for the final model. To find it, they’d need to run 1,000 experiments, trying many high-level approaches, many architectures for each component, hyperparameter search, and multiple seeds. Large machine learning projects need $10M in capital.
- jeffbee 6y agoI bet it's still a lot less than they spent training AlphaStar.
- moritonal 6y ago
- mensetmanusman 6y agoThis is amazing, if we can simulate multi-protein interactions, you could imagine in our lifetimes being able to see a fully computation driven simulation of a human blood cell. That would be a huge breakthrough.
- visarga 6y agoWhat amazed me most was that they used hundreds of millions of unlabelled protein scans. This means we can collect massive data in a new modality, besides the usual suspects: images, video, audio, text, lidar and sensors. Soon I expect neural implant data to be massive as well. They surely did unsupervised training on raw data and then fine-tuning on the 170K labelled sequences. I expect the data volume could be increased by orders of magnitude in the next couple of years and we'll see a GPT-3 like jump.
- comicjk 6y agoCASP (Critical Assessment of protein Structure Prediction) is calling it a solution. To quote from the article: "We have been stuck on this one problem – how do proteins fold up – for nearly 50 years. To see DeepMind produce a solution for this, having worked personally on this problem for so long and after so many stops and starts, wondering if we’d ever get there, is a very special moment." --Professor John Moult Co-founder and chair of CASP
- dekhn 6y agoIt's an improvement- and a big one- but not a solution to the problem. It mainly shows just how stuck the community had gotten with their techniques and how recently improvements in DNNs and information theory methods can be exploited if you have lots of TPU time.
- aardvarkr 6y agoIt’s officially recognized as a solution.
- cambalache 6y agoWell, it's not. Nature does not have a committee sorry. Proteins are delicate "machines" where even a a small change in the sequence (and thus the 3D structure) as small as a few amino-acids would change effectively the structure and the function of it. On top of that, proteins are dynamic beasts. In any case, it's a great advance, but DM, as many companies likes a little bit too much to tout its own horn.
- deleted 6y ago[deleted]
- ClumsyPilot 6y agoI am not sure we are talking about the same thing -i.e. there is a solution for hunger, but it's not a solved problem.
- mncharity 6y agoAdditional commentary in Science: https://www.sciencemag.org/news/2020/11/game-has-changed-ai-triumphs-solving-protein-structures https://www.sciencemag.org/news/2020/11/game-has-changed-ai-... (submitted by furcyd : https://news.ycombinator.com/item?id=25254888 https://news.ycombinator.com/item?id=25254888 ).
- ramraj07 6y agoThe most amazing part: > The organizers even worried DeepMind may have been cheating somehow. So Lupas set a special challenge: a membrane protein from a species of archaea, an ancient group of microbes. For 10 years, his research team tried every trick in the book to get an x-ray crystal structure of the protein. “We couldn’t solve it.” > But AlphaFold had no trouble. It returned a detailed image of a three-part protein with two long helical arms in the middle. The model enabled Lupas and his colleagues to make sense of their x-ray data; within half an hour, they had fit their experimental results to AlphaFold’s predicted structure. “It’s almost perfect,” Lupas says. “They could not possibly have cheated on this. I don’t know how they do it.”
- pmastela 6y agoLike the old Arthur C. Clark quote goes: “Any sufficiently advanced technology is indistinguishable from magic” -- unless it might be cheating in which case throw them a curve ball. Kudos to the DeepMind team for making magic happen.
- rsiqueira 6y ago"A sufficiently advanced Artificial Intelligence would be indistinguishable from God." (Way Of The Future - AI Church)
- mensetmanusman 6y agoThat would require the AI to exist outside of time and space.
- AlexCoventry 6y agoWhat's the actual news, here? AlphaFold is amazing, but it's been around for a while.
- typon 6y agoAlphaFold 2. The article specifically mentions it.
- AlexCoventry 6y agoThanks.
- verroq 6y agoSo who will have access to this? DeepMind never publishes their models.
- randcraw 6y agoI suspect DM will sell this as a service, especially to corporations like pharmas who create small molecule drugs. If their method works as advertised, it may rejuvenate the flagging prospects of Rational Drug Design, the guiding R&D drug development methodology behind most new molecular entities (drugs) for the past ~25 years, which has not proven to be the clear economic win that had been hoped.
- TomJansen 6y agoAccording to [1], they must release enough information for others to replicate the AI model: "As a condition of entering CASP, DeepMind—like all groups—agreed to reveal sufficient details about its method for other groups to re-create it. That will be a boon for experimentalists, who will be able to use accurate structure predictions to make sense of opaque x-ray and cryo-EM data." [1]: https://www.sciencemag.org/news/2020/11/game-has-changed-ai-triumphs-solving-protein-structures https://www.sciencemag.org/news/2020/11/game-has-changed-ai-...
- deleted 6y ago[deleted]
- EgoIncarnate 6y ago"AlphaFold achieves a median score of 87.0 GDT". Game changing, and a huge improvement, but not 100% solved. Also this is for static folding. Dynamic folding and interaction is a much harder problem. Those need to be tackled too before I would consider protein folding 'solved'.
- nabla9 6y agoThey solved the latest folding competition benchmark set. Shorter problems are easy to solve. Median score is mix of easier hand harder problems. Next year competition will have new set of much bigger and harder problems to solve. This seems like a leap, not solved as in having solution that just works and scales.
- hans1729 6y ago>Those need to be tackled too before I would consider protein folding 'solved' Semantics. From a systemtheoretical point of view, dynamic folding is an abstraction of static folding; solve (i.e. understand the underlying mechanisms) static folding and you can start progressing on dynamic folding, building up on your previously achieved solution. Wether it's solved or not depends on wether you mean `general folding` or the `entire spectrum of folding` when considering the problem.
- 6gvONxR4sf7o 6y agoSolve could mean understanding the underlying mechanism, but in this case, I don’t think that’s how they did it.
- hans1729 6y agoMy intuition for deeplearning was exactly that, statistical inference of underlying mechanisms. But I haven't read the paper yet, so you might be right
- ramraj07 6y agoIt's probably never going to be solved though right. To truly solve protein folding we'd have to have a program that can stimulate a small but still significant system at the QM level; looks like deep learning can get us 60% (conservatively estimating the whole problem domain ) but not all the edge cases, just like it did in other problem domains as well.
- hsnewman 6y agoThat's kinda a big deal.
- cs702 6y agoTwo years ago, after DeepMind submitted its first set of predictions to CASP (Critical Assessment of protein Structure Prediction), Mohammed AlQuraishi, an expert in the field, asked, "What just happened?" https://moalquraishi.wordpress.com/2018/12/09/alphafold-casp13-what-just-happened/ https://moalquraishi.wordpress.com/2018/12/09/alphafold-casp... Now that the problem of static protein structure prediction has been solved (prediction errors are below the threshold that is considered acceptable in experimental measurements), we can confidently answer AlQuraishi's question: Protein Folding just had its "ImageNet moment." In hindsight, AlphaFold v1 represented for protein structure prediction in 2018 what AlexNet represented for visual recognition in 2012.
- flobosg 6y agoAlQuraishi described the progress made in CASP13 (2018) as “two CASPs in one”. This one is an even bigger breakthrough.
- Seanambers 6y agoI particularly like the rant on pharmaceuticals companies lack of basic research. My impression has been that medical progression have been slow for quite some time, nice to see that there are some truth to that. In the end software and tech companies might just eat up the pharmaceutical industry as well. - It's all just code at some level. The Deepmind team did this with ; "We trained this system on publicly available data consisting of ~170,000 protein structures from the protein data bank together with large databases containing protein sequences of unknown structure. It uses approximately 128 TPUv3 cores (roughly equivalent to ~100-200 GPUs) run over a few weeks, which is a relatively modest amount of compute in the context of most large state-of-the-art models used in machine learning today." So it wasn't out of reach for academia, pharmaceuticals, or others with a bit of resources.
- flobosg 6y agoYeah, it was a big slap in the face. But, to be fair, most of the scientific and technological advances (sequencing efforts, structural genomics projects, etc.) that generated the data used by DeepMind came from academia and, to a lesser extent, the pharma industry.
- lgeorget 6y agoSee also the piece in Nature about the topic: https://www.nature.com/articles/d41586-020-03348-4 https://www.nature.com/articles/d41586-020-03348-4
- vadansky 6y agoJust to add to this whole "It's not solved! Yes it is!" discussion. Note that >According to Professor Moult, a score of around 90 GDT is informally considered to be competitive with results obtained from experimental methods. So if we go by >= 90 as solved: >In the results from the 14th CASP assessment, released today, our latest AlphaFold system achieves a median score of 92.4 GDT overall across all targets. they solved for their targets, but >Even for the very hardest protein targets, those in the most challenging free-modelling category, AlphaFold achieves a median score of 87.0 GDT (data available here). They basically admit they still haven't "solved" it for "most challenging free-modelling category" Take that as you will, not sure how useful the ">= 90 is solved" criteria is since they call it "informal" themselves.
- pretendscholar 6y ago87 GDT sounds pretty much solved to me if 90 is the benchmark
- garmaine 6y agoThat’s shifting goal posts. The hardest structures are also going to be harder experimentally. What makes them hard to predict is the very close energies involved in different folding pathways. Those close energies mean there will be more variant structures which change by use the experimental approach too.
- deleted 6y ago[deleted]
- fastball 6y agoWhat do you mean you're not sure how useful ">= 90" is as a criteria? You literally said why it is useful in your comment: > 90 GDT is informally considered to be competitive with results obtained from experimental methods. It's informal because we don't have a true "gold-standard" for determining a protein's folded structure – the best we have is experimental methods of trying to determine the structure which still have a great deal of error (compared to other things we can measure). So all we can do is say "the GDT between two experimental measurements (of the same protein) is often around 90, so if we get there with predictive models that's pretty much just as good". As soon as we have better experimental methods for determining protein tertiary structure, you can be sure we will require predictive models to deliver better results too. Until then, the point is that the delta between any two experimental determinations of folded structure is approximately the same as the delta between an experimental determination and an AlphaFold guess. So the AlphaFold guess may as well be an experimental measurement. Except the AlphaFold guess happens fairly trivially (once you give it the DNA sequence[1]), where as the experimental method is involved and expensive. [1] Or the primary structure, I'm unsure what inputs are given to AlphaFold.
- Quarrel 6y agoAs someone who wrote a thesis many moons ago about protein folding, this is pretty astonishing to see. Yay science.
- haolez 6y agoDoes this make Folding@Home obsolete?
- hoppla 6y agoI am puzzled me about “AI-knowledge”. Have we really learnt anything? Is distilling the knowledge from AlphaFold just as a hard problem as solving protein folding?
- fairity 6y agoIf you forgot how to do long division, but still had a calculator, wouldn't the calculator still be useful?
- 6gvONxR4sf7o 6y agoI hate headlines like “X has solved Y.” How often have we see computer vision and natural language solved at this point, whenever a model does well enough in a benchmark? Their own article doesn’t even have that headline. This is a massively cool thing that’s happened. Why ruin it with a massively hyperbolic headline?
- TheRealPomax 6y agoBecause only the experts in this field get to tell us, the laymen, what "solving the protein folding problem means", and they defined it not as "perfect" but as "more than good enough to be acceptable as correct result". Which this did. X has actually solved Y. That's not so much "massively cool", that's historical.
- 6gvONxR4sf7o 6y agoI think the “they” you’re referring to is only whatever PR person wrote the headline. Nowhere in the substance of this (PR!) post does it refer to it as anything but a great leap. When an expert in the field outside of deepmind says protein folding has been solved, I’ll believe it.
- nharada 6y agoIt does appear other experts in the field are claiming this: https://twitter.com/MoAlQuraishi/status/1333383769861054464 https://twitter.com/MoAlQuraishi/status/1333383769861054464
- falcor84 6y agoI don't think I ever saw a headline saying natural language is solved; who's claiming that?
- danaris 6y agoThe "solved protein folding" part isn't even in the article. It appears to be clickbait editorialization by whoever submitted the link.
- breatheoften 6y agoAnyone care to muse about appropriate investment strategies based on the not previously feasible research approaches that might now be possible? Should we expect to see faster progress in large well capitalized bioscience companies -- or a sudden increase in the viability of smaller biotech and/or biotech startups ...? Are we gonna see top talent fleeing the old biotech companies to start their own ventures with a new belief that the potential for huge reward might suddenly seem achievable? What kind of companies do we think will be the first that are able to translate this new knowledge into profits?
- enchiridion 6y agoI agree. I think a company that is working on large scale automated bio experiments would be well positioned to take advantage of something like this. What companies are doing that work?
- 0xBA5ED 6y agoDoes this give the ability to engineer cures for currently incurable diseases?
- jfarlow 6y agoIn short - certain ones, yes. This should be one step (that was a bottleneck) in helping a company with a fixed budget do an order or magnitude more 'experiments' with the same amount of resources. Lab resources are expensive and fixed, so if you can pre-compute what you need, you can get right to the more powerful results. We design proteins for immunotherapies - this kind of thing would help us more rapidly design our proteins (and more efficiently use our wet-lab resources to speed existing projects). For others, some drugs are hard build without knowing how they will interact - this could both provide new 'targets' to go after, but also might help prevent projects that would otherwise accidentally target an important protein.
- fogleman 6y agoHow will this get into the hands of those who could use it?
- sanxiyn 6y agoRealistically speaking, if you are a scientist who could use this and you mailed DeepMind, they will probably run it for free and send you the result. It would be a good PR.
- mylons 6y ago12-13 years ago in a classroom the professor for my intro to bioinformatics class said if you were to solve this problem, you would win a Nobel prize. Congrats to the team! What an achievement.
- JauntTrooper 6y agoThey will definitely win the Nobel for this.
- gravy 6y agoMaybe combine with https://news.ycombinator.com/item?id=25254772 https://news.ycombinator.com/item?id=25254772 ?
- tgbugs 6y agoMy conclusion reading this is that a gradient is a gradient is a gradient. If you can minimize one, you can minimize them all. The hard work would seem to be figuring out how to transform into a gradient that your hardware can solve. It will also be interesting to see the kinds of systematic errors that will come as a result of the biases in the training set, and whether it can be used to predict what the structures would look like under slightly different conditions (e.g. pH).
- danaris 6y agoThe title here is not merely breathless clickbait, it also has very little to do with the headline of the actual article, which is "AlphaFold: a solution to a 50-year-old grand challenge in biology". I thought the #1 criterion for titles was that they should match the original if at all reasonable...?
- dalbasal 6y agoQuestion for the wise: Assuming optimistic further progress, what are the implications of accurately predicting protein folding? What are we hoping to discover, or succeed in doing?
- justinzollars 6y agoI have a Masters in Biology. This was once described as an impossible problem to solve. A huge achievement.
- m3kw9 6y agoDoes this obsolete Folding@home?
- michaelcampbell 6y agoMy question exactly; or Rosetta @ home, or any of the other protein folding "@home"s. I participate in a few, but would gladly donate my compute resources elsewhere if this is no longer necessary.
- sleepysysadmin 6y agoI made new years prediction about exactly this. I predicted folding@home would die despite huge interest again because of covid.
- uoaei 6y agoNo, they didn't. They approximated a solution to protein folding. The two are different concepts -- this isn't the typical HN pedantry. "Solving" the problem would entail developing an interpretable algorithm for taking a string of amino acids and determining the 3D structure once folded. Approximating a solution would entail simulating that algorithm, which is what their neural network is doing. It is of course usually accurate, but you would expect this with any suitable universal function approximator. Props to DeepMind and congrats to CASP but is it not obvious that this is more hype-rhetoric for public consumption?
- astroalex 6y agoThe distinction you're making between "solved" and "closely approximated" makes logical sense to me. However, if I'm interpreting the AlphaFold results correctly, this distinction isn't practically significant, right? If you can approximate an algorithm with error that is "below the threshold that is considered acceptable in experimental measurements" (to quote another HN comment), then you have something as good as the algorithm itself for all intents and purposes. Therefore the use of the word "solve" doesn't qualify as hype-rhetoric, and the distinction you're making does seem somewhat pedantic (even if technically true). (I'm speaking as someone with only the tiniest amount of stats/ML experience, so I could be totally wrong!)
- colonelcanoe 6y agoIt might be the case that the relevant, practical threshold now tightens. For example, perhaps it is easier to experimentally verify a protein shape predicted by an algorithm than it is to experimentally determine the protein shape?
- intpx 6y agoexactly. Even an incomplete map with somewhat limited resolution makes navigation a hell of a lot easier than flying blind. This effectively is a data reduction solution-- if you have a fuzzy shape of the thing you are trying to model, and you learn the mechanics better with each thing you model, your ability to quickly and accurately reach a goal improves
- SubiculumCode 6y agoRIP folding at home? EDIT: Just throwing this out there: Are there national security issues to think about with this? Can it be used to weaponize computational biology?
- flobosg 6y agoFolding@home tackles a related but different problem. They simulate folding dynamics, i.e. how does a protein reach its folded structure. If AlphaFold gives you a picture of a protein structure, Folding@home shoots a video of that protein undergoing folding.
- amelius 6y agoThis was also one of the main selling points of quantum computers. Makes you wonder what Deep Learning will tackle next. Factorization of large integers?
- optimalsolver 6y agoCan just anyone enter this challenge, or do you have to be part of a major institution?
- dmd 6y agoAnyone.
- optimalsolver 6y agoCan just anyone enter this challenge, or do you have to be part of a major institution?
- flobosg 6y agoI think anyone can take part. There are a few unaffiliated participants.
- amelius 6y agoCurious, what are the sizes of the training and validation/test datasets (number of structures)?
- papaf 6y agoCurious, what are the sizes of the training and validation/test datasets (number of structures)? The proteins are shown on the CASP website [1]. Both the number of residues and number of proteins are bigger than I expected. [1] https://predictioncenter.org/casp14/targetlist.cgi https://predictioncenter.org/casp14/targetlist.cgi
- asbund 6y agoExiting time
- wespiser_2018 6y agoThis will undoubtably change our understanding of human health and biology in many impactful ways in the years to come! The same information we get through x-ray diffraction will now be available 100x or even 1000x cheaper, and using this model can even aid the interpretation of xray diffraction data! What excites me most isn't doing what we can do now, for cheaper (which will surely lead to more effective research methods), but the potential to gain a systematic view of protein structures, either across the genome, species, or through time which will give us a deeper and more fundamental understanding of biology.
- sjg007 6y agoLooks like a transformer model. Anyone have any insights?
- schemescape 6y agoDid AlphaFold2 also have the biggest budget? :) Edit: from the other HN article on this topic: > We trained this system on publicly available data consisting of ~170,000 protein structures from the protein data bank together with large databases containing protein sequences of unknown structure. It uses approximately 128 TPUv3 cores (roughly equivalent to ~100-200 GPUs) run over a few weeks https://deepmind.com/blog/article/alphafold-a-solution-to-a-50-year-old-grand-challenge-in-biology https://deepmind.com/blog/article/alphafold-a-solution-to-a-...
- curiousllama 6y agoActually, no! Or at least, the budget they used (<$100k at retail prices to train the model) is well within the feasible range for other research institutions. In other words, it's less like GPT3 and more like ImageNet.
- whimsicalism 6y agoI don't know - tens of thousands per train is not accessible for most academic institutions when you consider the necessity of ablation studies, experimentation, etc.
- anchpop 6y agoFor a topic like protein folding, it should be
- whimsicalism 6y ago> For a topic like protein folding, it should be Well, I've worked in some academic deep research labs and they did not have the money to do the experiments they wanted to do.
- lacksconfidence 6y agoIs the cost to train really the relevant metric for developing this? It seems like the salary's involved are probably at least 10x whatever they spent on hardware.
- WanderPanda 6y agoWe indeed stand on the shoulders of a small number of giants! I'm infinitely thankful for the work DeepMind is doing. Lets maybe celebrate this accomplishment for one day and start being worried about big tech again tomorrow. Many of the comments here usually suggest that we should live in worries and fear but to my knowledge there is not too much historical evidence for these kind of companies turning evil.
- tpoacher 6y agoWhenever deepmind comes up with something like this, my first instinct is to say "yay for humanity" ... then I remember who they work for, and the second instinct is to say "Ah. Crap."
- purpleidea 6y agoDoes this produce the various different foldings that each protein can often "sit" in? Can it take temperature and other environmental conditions into account? Can you specify that a particular ligand or electrical current is present so that you can see the resultant shape change? Is all the source code for this available so that other scientists can build on top of this, or will we have to go through a paid or SaaS google API to use it?
- spenczar5 6y agoAlphaFold was used to analyze proteins in SARS-CoV-2 (https://www.crick.ac.uk/news/2020-03-05_crick-scientists-support-deepminds-effort-to-make-covid-19-data-available https://www.crick.ac.uk/news/2020-03-05_crick-scientists-sup...). Does anyone know what impact that has had? This is really an amazing moment.
- xphos 6y agoLike this is awesome and a huge advancement but one thing that worries me with an AI solution is that it doesn't really draw us any closer to the why. Why do proteins fold the way they do? We can predict the resulting structure which is extremely significant, we have no clue why. While we get the insight of being able to predict some structures we don't get the insight of why things are happening the way they are. In some cases like this it might not matter but in other cases that insight might actually be way more significant than answer the problem to begin with. Of course we can review over the problem with the additional predictions that AI gives us but this can be haphazardous because what if there is specific sequence spins in some certain way that we and thus the AI has never seen and it goes missed. I'm not a biologist to say this is possible but I known this kind of edge case can come up and what rabbit holes will we go down because we only have the AI implied insight. disclaimer I think the contributions are super useful for science but they do come with worries as does every path of discovery
- Odenwaelder 6y agoI have no idea what you are talking about.
- xphos 6y agoAI solve the process but doesn't give a whole lot of insight into the formulas and the description what's going on. Where we as humans have reasonably found that e = mc^2. However AI would gives us e or m but backboxes us away from seeing that c aka the speed of light was involved(unless we implied that before). There might be interesting relationships that are useful that AI unintentionally masks that could be ground breaking if we could only understand process more holistically. I think a different commenter eluded in this case we think we understand protein folding well we just struggle to synthesis it in a compact mathematical way even though with AI we can simulate the process well for known examples. The issue with AI is we don't know if our current example set includes every case what if there is a strange sequence of amino acid that causes something "weird" to happen that we have haven't seen. AI cannot predict something novel it or us haven't seen which is the issue. The process(if it exists) of how one could solve this problem might also be exportable to other fields if it was formulized with math rather than estimated with AI.
- andy_ppp 6y agoSo, sorry to be a philistine but what specific discoveries will this lead to... will it make it easier to produce antivirals or even molecular machines?
- tim333 6y ago"DeepMind said it had started work with a handful of scientific groups and would focus initially on malaria, sleeping sickness and leishmaniasis, a parasitic disease" https://www.theguardian.com/technology/2020/nov/30/deepmind-ai-cracks-50-year-old-problem-of-biology-research https://www.theguardian.com/technology/2020/nov/30/deepmind-...
- bayeslaw 6y agoas so many time recently the hn crowd proves to be completely clueless and uneducated when it comes to ai.. this is a miracle.. it is THE achievement we'll remember from the past decade when it comes to ai.. if you don't understand why I recommend learning and reading.the level of ignorance and often proud ignorance here is frightening to me.. ppl who downplay this are either stupid in biochemistry or ai or both .. please don't listen to them. this right here is the single biggest news of 2020..
- dr_dshiv 6y ago"It has occurred decades before many people in the field would have predicted. It will be exciting to see the many ways in which it will fundamentally change biological research."
- FrojoS 6y agoBell Labs invented the transistor. Now this. Monopoly money at its best!
- dang 6y agoUrl changed from https://predictioncenter.org/casp14/zscores_final.cgi https://predictioncenter.org/casp14/zscores_final.cgi, which points to this.
- sabujp 6y agoSo it might "be over" for small molecules, but let's see large macromolecules and protein assemblies be predicted
- sabujp 6y agoso are all these protein folding labs and projects e.g. folding at home, etc essentially dead projects now?
- dluan 6y agoI worked in the lab that helped develop folding@home, as well as the game where the crowd was the chaotically trained machine that folded and unfolded one amino acid at a time. This feels like a pretty significant new chapter in the humanity movie. A few times, I get immense pangs of jealousy for younger people a generation or a half before me. And I'm only 30! This is one of those times.
- maxlamb 6y agoIs the team really that young? 20 year olds?
- sgillen 6y agoI think he means that those who are in their teens / even younger now will get to experience immensely cool tech in their lifetime.
- dluan 6y agoWhen I was there, it was a lot of smart grad students and undergrads, and the occasional old professor :) That group held the previous highest scores.
- heycosmo 6y agoFascinating! AlphaFold (and other competitors) seem to use MSA (Multiple Sequence Aligment) and this (brilliant) idea of co-evolving residues to build an initial graph of sections of protein chain that are likely proximal. This seems like a useful trick for predicting existing biological structures (i.e. ones that evolved) from genomic data. I wonder (as very much a non-biologist), do MSA-based approaches also help understand "first-principles" folding physics any better? and to what degree? If I write a random genetic sequence (think drug discovery) that has many aligned sequences, without the strong assumption of co-evolution at my disposal, there does not seem any good reason for the aligned sequences to also be proximal. Please pardon my admittedly deep knowledge gaps.
- flobosg 6y ago> do MSA-based approaches also help understand "first-principles" folding physics any better? Not really. MSA-based approaches, as most structure prediction methods, have as a goal to find the lowest energy conformation of the protein chain, disregarding folding kinetics and basically all dynamic aspects of protein structure. > If I write a random genetic sequence (think drug discovery) that has many aligned sequences, without the strong assumption of co-evolution at my disposal, there does not seem any good reason for the aligned sequences to also be proximal. I don't think I fully understood this, but I'll give it a shot anyway. If your artificial sequence aligns with others, there's a chance that it will fold like them, depending on the quality and accuracy of the multiple sequence alignment. Since multiple sequence alignments are built under the assumption of homology (all sequences have a common ancestor), it's a matter of how far from the "sequence sampling space" your sequence is located compared to the others.
- heycosmo 6y ago> I don't think I fully understood this, but I'll give it a shot anyway. If your artificial sequence aligns with others, there's a chance that it will fold like them, depending on the quality and accuracy of the multiple sequence alignment. Since multiple sequence alignments are built under the assumption of homology (all sequences have a common ancestor), it's a matter of how far from the "sequence sampling space" your sequence is located compared to the others. I understand that similar sequences may fold similarly (although as length increases, I highly doubt it, but IDK). I'm talking about aligned sub-sequences within one chain and their ultimate distance from each other in the final structure. Co-evolution suggests that aligned sub-sequences are also proximal. But manufactured chains did not evolve, therefore the assumption is no longer useful.
- EGreg 6y agoWhat happens when AI is better at everything measurable than humans? Better at conversation. Better at making people laugh, and generate attraction or other emotions, better at motivating them, and organizing movements, etc. Clearly we are not ready for such an efficient system... it would be a big disruption to all human organizations and relations. It would start with Twitter botnets and directing sentiment.
- daxfohl 6y agoThey still suck pretty bad at many physical things. Bipedal robots are a joke. They also dent and rust. It'll be that way for a while. They don't reproduce. But virtual world, say they're better at math. Say they prove all the Clay Millennium Problems. Say they go way beyond those problems and produce some math far beyond human's ability to understand it. I've been thinking about that for a while and have decided it's fine. Math as a profession will still exist. Fact is, there's already a proof for everything mathematicians are investigating (or a proof that there's no proof, recursively), out there somewhere. Mathematicians are just searching for it, so that it can be understood and translated to human language. The fact that AI already knows the answers doesn't mean that human mathematicians are useless: they are still required to uncover the meaning of these results and translate them into human language. AI then is still just a tool that mathematicians use to help them in their search. Similar to how biologists will use AlphaFold. I guess. Or... https://imgur.com/gallery/9KWrH#sv05qpF https://imgur.com/gallery/9KWrH#sv05qpF
- jeffxtreme 6y agoGDT_TS for AlphaFold is now comparable is at experimental levels; but that's based on the class of proteins for which we've been able to determine the 3D structure of the protein, for which there might be selection bias. I wonder if we can determine if this extends to proteins that aren't as keen to determining their 3D structure? For example, certain proteins are more crystallizable than others.. For these non-crystallizable proteins, I wonder if we can say that AlphaFold would generate accurate 3D models? And if possible, might there be a way to map out this uncertainty?
- deeviant 6y ago> I wonder if we can determine if this extends to proteins that aren't as keen to determining their 3D structure? This is already happened. "An AlphaFold prediction helped to determine the structure of a bacterial protein that Lupas’s lab has been trying to crack for years. Lupas’s team had previously collected raw X-ray diffraction data, but transforming these Rorschach-like patterns into a structure requires some information about the shape of the protein. Tricks for getting this information, as well as other prediction tools, had failed. “The model from group 427 gave us our structure in half an hour, after we had spent a decade trying everything,” Lupas says." From: https://www.nature.com/articles/d41586-020-03348-4 https://www.nature.com/articles/d41586-020-03348-4
- jeffxtreme 6y agoAgree this is great to hear, but the fact that they had X-ray diffraction data indicates this protein was indeed crystallizable no? Though the next paragraph in the article shows that DeepMind is indeed working on mapping out reliability: "Demis Hassabis, DeepMind’s co-founder and chief executive, says that the company plans to make AlphaFold useful so other scientists can employ it. (It previously published enough details about the first version of AlphaFold for other scientists to replicate the approach.) It can take AlphaFold days to come up with a predicted structure, which includes estimates on the reliability of different regions of the protein. “We’re just starting to understand what biologists would want,” adds Hassabis, who sees drug discovery and protein design as potential applications."
- The_rationalist 6y agoLet's imagine that as a researcher I make a breaktrhough NN model, but that I need a lot of TPUs/GPUs in order to test it, is there a service for temporarily lending such hardware to me for free/not much ? (e.g google colab ?) Otherwise researchers will plateau with their hardware budget.
- yk 6y agoVery interesting, however now the problem becomes to characterize such machine learning approaches. With traditional simulation methods the authors can usually explain easily in which situation a specific approach is good or bad, with neural networks we don't really have a good approach how to analyze the quality of the prediction.
- dang 6y agoAll: there are multiple pages of comments; if you're curious to read them, click More at the bottom of the page, or like this: https://news.ycombinator.com/item?id=25253488&p=2 https://news.ycombinator.com/item?id=25253488&p=2 We changed the URL from https://predictioncenter.org/casp14/zscores_final.cgi https://predictioncenter.org/casp14/zscores_final.cgi to the blog post, which has more background info.
- kovek 6y agoI've seen you mention this [More] comment a few times now. I like it, though what if you change the design of the More functionality?
- nathancahill 6y agoAlso, what do the traffic stats look like for the second/third pages of big threads like this one? Pretty steep falloff?
- dang 6y agoI assume so, but haven't looked recently. I'll try to do that and report back here later. Feel free to ping me at hn@ycombinator.com if I forget. Edit: ok, for this thread so far, 95% of views are page 1, 4% are page 2, 1% are page 3. For https://news.ycombinator.com/item?id=25065026 https://news.ycombinator.com/item?id=25065026, which had a "more pages" comment at the top: 93% of views were page 1, 5% page 2, 2% viewed page 3. For https://news.ycombinator.com/item?id=23155647 https://news.ycombinator.com/item?id=23155647, which did not have a "more pages" comment at the top: 96% of views were page 1, 3% page 2, 0.5% page 3. Radically overgeneralizing from that, it seems likely that the pinned comment at the top helps a bit in terms of directing people to later pages. How that compares to the mammoth-single-page scenario is hard to say because we don't know how many readers would be scrolling down that far to see those comments. There's likely a power-law dropoff no matter what we do.
- kovek 6y ago
- 29athrowaway 6y agoSo what's going to happen to fold.it and folding@home now?
- flobosg 6y agoSee https://news.ycombinator.com/item?id=25256318 https://news.ycombinator.com/item?id=25256318 and https://news.ycombinator.com/item?id=25256772 https://news.ycombinator.com/item?id=25256772
- flobosg 6y agoSee also the news in Nature: https://www.nature.com/articles/d41586-020-03348-4 https://www.nature.com/articles/d41586-020-03348-4
- uuuuuuuuuuuu 6y agoI feel like DeepMind has a disproportionately large scientific impact relative to its resource pool. How would one (or a group) go about replicating its success?
- ribrars 6y agoI think the key here to replicating the success is the deployment of deep learning effectively. But I would argue that deepmind's resource pool is immense, it's backed by Google. The resources of GPU's (and more advanced TPU's) are in abundance... not to mention the many brilliant PhD scientists who work there.
- piannucci 6y agoThis is so cool. I hope they will also tackle the problem of predicting RNA structures and catalytic activity.
- optimalsolver 6y agoThey should make this into a Kaggle competition. Maybe they might get an even better model.
- troelsSteegin 6y agoCan anyone (yet) provide a sketch of how this works? I saw a mention of "attention", which I vaguely take to be a surrogate for some form of structural information. It's an astonishing result. How does it work?
- Havoc 6y agoDoes something like Folding@Home still have meaning after this?
- notkaiho 6y agoWho'd have thought that the kid who programmed Theme Park would go on to do this kind of work.
- hawkjo 6y agoSo how fast does the prediction model run? Will they soon be publishing predictions for millions of known sequences?
- bredren 6y agoSome very relevant sci-fi art to this news is the show Devs. It specifically looks at how AI can be used for predictions. The show immediately came to mind in reading this news. It aired on Hulu.
- yread 6y agomindblowing how far ahead this model is https://predictioncenter.org/casp14/gdtplot.cgi?group=205&models=first&target=T1024-D1 https://predictioncenter.org/casp14/gdtplot.cgi?group=205&mo...
- yread 6y agoOr even this one: https://twitter.com/TassosPerrakis/status/1333535594002132996/photo/1 https://twitter.com/TassosPerrakis/status/133353559400213299...
- daxfohl 6y agoIs this going to put biologists that study this problem (and there are a lot of them, right?) out of business? This is the tipping point where I think the AI singularity may have teeth. I could see math proofs being the next thing to fall. If AI solves the remaining six millennium problems in the next few years, what does that mean for math researchers?
- russdpale 6y agoIf this is the only thing that comes from AI, or if this is the only lasting application the technology, then all of the research and time and code and frustration will have been worth it.
- antipaul 6y agoWhat size was the test set? From what I gather, training was done on 170,000 Amino Acids (features) and the resultant protein structure (labels). This is out of 200 million possible proteins. How many examples were in the test set? EDIT: Looks like N=100 for test set: “Entrants get amino acid sequences for about 100 proteins whose structures are not known” https://www.sciencemag.org/news/2020/11/game-has-changed-ai-triumphs-solving-protein-structures https://www.sciencemag.org/news/2020/11/game-has-changed-ai-...
- antipaul 6y agoAs usual with ML, I now wonder how “similar” the test set is to the training set, compared to the examples that are neither in the training set, nor the test set: TODO: 200 million - 170,000 training - 100 test ~= 199.8 million proteins
- kxs 6y agoThey trained on 170k sequences/ structures/ proteins, each sequence has 10s to 100s or even 1000s amino acids. Structure is much more conserved than sequence. Out of the 100 targets, roughly 1/4th have no similarity to known structures, so there shouldn't be an overlap for those with the training set. They did very well on those targets.
- gabia 6y agoThe vast majority of structures in the protein data bank are determined by crystallography, which involves putting the protein in a chemical cocktail that causes it to crystallize. The cocktail is very different to the chemical environment in which the protein functions, so an open question is whether the protein structure determined by crystallography (and hence learned by AlphaFold) is representative of the structure in it's natural environment. It would be very interesting if there was a way to use computational techniques to go beyond what crystallography and other experimental techniques (Cryo) can accomplish and determine the protein structure in it's true biological setting. Some research into experimental methods for this include high power X-ray pulses. Nonetheless, impressive work!
- thamer 6y agoThere's something I don't understand about protein shapes. There are tons of software solutions – on the web and offline – to visualize the shape of proteins from their sequence of amino acids. How do these work then, if we don't know how the atoms might be arranged in space? For example, this[1] is the code for SARS-Cov-2's Spike(S) protein. From what I understand of this page it's pretty short, only ~1,757 proteins (corresponding to ~3,821 bases in RNA). And here[2] is a visualization of it in 3D. You'll likely recognize the characteristic mushroom shape that's been portrayed in 3D models of SARS-CoV-2 in the media. How does this software work if there's no real way to tell how the protein is arranged? [1] https://www.ncbi.nlm.nih.gov/nuccore/NC_045512 https://www.ncbi.nlm.nih.gov/nuccore/NC_045512 (search for spike glycoprotein on that page) [2] https://3dmol.csb.pitt.edu/viewer.html?pdb=6X6P&style=stick&select=resn:HOH;invert:1&surface=opacity:0.8;colorscheme:whiteCarbon https://3dmol.csb.pitt.edu/viewer.html?pdb=6X6P&style=stick&...
- jassany 6y agoI work suggest trying to understand how the haemoglobin protein works... The shape of it is not too confusing too look at.
- paraschopra 6y agoAl such proteins have been crystallised and we know their shape experimentally.
- thamer 6y agoThanks for the answer! That explains it. I was looking just at the amino acid sequence and missing a whole lot. I read the protein folding and X-ray crystallography articles on Wikipedia and they had most of the answers I was looking for. I also saw a request being made by this JavaScript 3D viewer to fetch the PDB (Protein Data Bank) file for the model, which is a text file with tens of thousands of lines describing the coordinates of atoms in space as well as their bonds and other structures. It even has some metadata about the way the data was collected. For the spike protein linked above: http://files.rcsb.org/view/6X6P.pdb http://files.rcsb.org/view/6X6P.pdb I find it fascinating that we're even able to scan the 3D structure of molecules with such precision.
- dnautics 6y agoI've long been an AI/ML positivist in the field of protein structure prediction (but not in drug discovery in general), admittedly a bit surprised it was now and not 3-4 years from now... And for a long time I have been saying that a "heuristic" model for folding is going to win (and it looks like it has). However, I would also caution that, there are going to be protein structures that are not in the opus of known structures (being able to solve the structure at all is itself a biasing factor) and AlphaFold's capability to figure those out will be interesting. I would not necessarily be confident it could. (think of issues, like face detection algorithms not being able to correctly identify minorities, e.g.)
- sagebird 6y agoCan it simulate two proteins interacting? IE search for the simplest protein sequence that a.) doesn't affect folding geometry when paired with every known human protein and b) causes the greatest deviation when paired with covid-19 proteins?
- jl6 6y agoDoes this make Folding@Home obsolete?
- Cort3z 6y agoAfter reading this I can't stop thinking about a possible future where we have predicted all possible permutations of different diseases and created vaccines for them. Maybe our kids one day will get an all in one vaccine preventing all viral and bacterial disease.
- hoseja 6y agoHow does it do with intrinsically disordered proteins? Those are some of the more interesting ones and very hard to xray.
- stefanka 6y agoWhere will the development go on from now? We have been working on a geometrical approach that avoids the curse of dimensionality to solving the same problem for the last few months. Now, I wonder whether it makes sense to continue at all (we were and are clearly not ready to participate in the challenge yet). So What remains unresolved? Exact position of side chains? Can their approach be used for protein-protein interaction too?
- Ono-Sendai 6y agoCan some explain why you can't just run a simulation simulating the forces between the amino acids, and let the protein curl up based on those forces?
- isoprophlex 6y agoAccurate modelling on such a detailed level becomes intractable due to long time scales needed for folding, and the presence of forces that are not adequately described at the "ball and spring" level of abstraction that molecular mechanics simulation usually employs. Its better to abstract everything away by a neural net, apparently...
- MarkMc 6y agoJason Crawford answers your question here: https://twitter.com/jasoncrawford/status/1333576261877125121?s=19 https://twitter.com/jasoncrawford/status/1333576261877125121... Short answer: it's extremely computationally expensive
- thetrooper 6y agoGlad to see AI is progressing beyond annoying customer support chatbots and marketing tools. At this rate it will predict the covid pandemic anytime soon now.
- yters 6y agoIs this immune to things like adversarial examples? E.g. will we get a situation where we flip one nucleotide or amino acid, and suddenly AlphaFold is making completely incorrect predictions?
- aazaa 6y agoIt's not clear whether the predictions were for the solution or solid phase. Can anyone speak to that?
- krick 6y agoThis is great and I feel weirdly relieved (considering I don't actually really gain anything from that). That, on the other hand, makes me feel sad and almost depressed every time: > It uses approximately 16 TPUv3s (which is 128 TPUv3 cores or roughly equivalent to ~100-200 GPUs) run over a few weeks, a relatively modest amount of compute in the context of most large state-of-the-art models used in machine learning today
- okarthik42 6y agoIMO, AlphaFold 2 is a great example of industrial research labs making huge breakthroughs. I'm not sure if AlphaFold 2 is over hyped or not (because I don't know anything about protein folding), but given how a lot of computational biologists reacted to the results (the co-founder of CASP seems very impressed :')), I suppose this a big deal. I hope DeepMind becomes the Bell Labs for AI. Bell Labs is the best example of industrial research labs making huge strides. Of course, AI doesn't exist yet, and deep "learning" is nothing but curve-fitting done in fancy ways, but I would not be surprised if DeepMind results in a few Turing and Nobel laureates.