14 ms·
Alphafold
- pjfin123 5y agoI'm assuming you can't run this on any consumer computer?
- pjfin123 5y agoNevermind > The simplest way to run AlphaFold is using the provided Docker script. This was tested on Google Cloud with a machine using the nvidia-gpu-cloud-image with 12 vCPUs, 85 GB of RAM, a 100 GB boot disk, the databases on an additional 3 TB disk, and an A100 GPU.
- sambroner 5y agoThat’s… way closer to consumer than I expected
- lifthrasiir 5y agoExcept for (DGX) A100.
- qeternity 5y agoFor inference... Still accessible, but expensive to run at scale. And training even worse.
- erhk 5y ago2.2TB data
- lasagnaphil 5y agoNah, 4TB disk drives are not that expensive.
- crazysim 5y agoAmazing. That's not a lot of libraries of congresses at all.
- dekhn 5y agowhich is basically nothing. They could put it in a cloud bucket and you could copy it to another bucket in minutes.
- swalsh 5y agoedit I was wrong. Please ignore.
- qeternity 5y agoOk, so biochemists: which bit of the secret sauce are they leaving out?
- dekhn 5y agoFantastic, they released the dataset and code to train the model. Science will be able to proceed. edit: not the code to train the model, just the code to run inference. The underlying sequence datasets include PDB strucrures and sequences, and how those map to large collections of sequences with no known structure (no surprise). Each of those datasets represents decades of thousands of scientists work, along with programmers and admins who kept the databases running for decades with very little grant money (funding long-term databases is something NIH hated to do until recently).
- FredFS456 5y agoThere's a preview paper as well: https://www.nature.com/articles/s41586-021-03819-2 https://www.nature.com/articles/s41586-021-03819-2
- dekhn 5y agoYes, I skimmed the paper already and it wasn't too surprising. There are details that will take some time to parse out to understand how important they are. Personally, I've found over decades that academic papers like that are far less useful to me than a github project and downloadable data that I can inspect, run and modify on my own. Other folks I know could read that paper and write the code in a day, I always wish I could do that.
- gopalv 5y ago> The total download size is around 428 GB and the total size when unzipped is 2.2 TB. Please make sure you have a large enough hard drive space, bandwidth and time to download. > This was tested on Google Cloud with a machine using the nvidia-gpu-cloud-image with 12 vCPUs, 85 GB of RAM, a 100 GB boot disk, the databases on an additional 3 TB disk, and an A100 GPU. This is amazingly detailed for a researcher who wants to follow in the track and also Apache licensed, which is one road-bump out of the way for a commercial enterprise, like an actual drug manufacturer who wants to burn some money trying this out. edit: said the last part too fast, the code has a "the AlphaFold parameters are made available for non-commercial use only under the terms of the CC BY-NC 4.0 license"
- nextos 5y agoAlphafold 2 is very very cool, but we need a little dose of reality. It's still a bit away from really solving protein folding as it was marketed. For example, multi-complex proteins are not well predicted yet and these are really important in many biological processes and drug design: https://occamstypewriter.org/scurry/2020/12/02/no-deepmind-has-not-solved-protein-folding/ https://occamstypewriter.org/scurry/2020/12/02/no-deepmind-h... A disturbing thing is that the architecture is much less novel than I originally thought it would be, so this shows perhaps one of the major difficulties was having the resources to try different things on a massive set of multiple alignments. This is something an industrial lab like DeepMind excels at. Whereas universities tend to suck at anything that requires a directed effort of more than a handful of people.
- timr 5y ago> A disturbing thing is that the architecture is much less novel than I originally thought it would be, so this shows perhaps one of the major difficulties was having the resources to try different things on a massive set of multiple alignments. This is something an industrial lab like DeepMind excels at. Whereas universities tend to suck at anything that requires a directed effort of more than a handful of people. Yeah, the HN commentary on Alphafold has a high heat-to-light ratio. I'm eager to read the paper because the previous description of the method sounded remarkably similar to methods that have been around for ages, plus a few twists. The devil is going to be in the details on this one.
- MrsPeaches 5y ago> high heat-to-light ratio Sorry for the ignorance but what does this mean?
- butMuhCulture 5y agoIt’s trying to say light is more valuable than heat, or some such folksy thing. I cook steak in the dark so I don’t find it to be a very insightful metaphor.
- 5y ago
- fossuser 5y agoDoes anyone on HN work in bio or drug discovery? Could you give an overview of how people can leverage this (or how you might?). From reading around about it, it sounds like there's often a need to find a certain type of molecule to activate/inhibit another based on shape and the ability to programmatically solve for this makes the searching way easier. Is this too oversimplified/wrong? How will this be used in practice. [Edit]: Thanks for the answers!
- timr 5y ago> Could you give an overview of how people can leverage this (or how you might?). Short answer: nobody knows. Traditionally, protein folding is a solution in search of a problem, but that's largely because the predictions were...unusably bad. This was always more of a super-difficult validation problem for the force fields and simulation methods, which could then be used for other problems of greater value (such as rational protein design, or simulation of the motion of proteins with known structures). These predictions are better, but still pretty far from the level of precision that you'd want for any kind of rational drug design, where the exact locations of protein side-chains (for example) matter a lot. You'll note that AlphaFold returns structures that are "relaxed" using one of the oldest simulation systems for proteins: AMBER. So it's not exactly a clean-room solution to the problem, and you can't assume that the details (which matter to drug design) are going to be any better than for the older methods. But that said, if you have a method that can reliably give you a blurry view of the overall shape of a protein, even that could be useful for things like target discovery or inference of biological networks. But this is still a lot closer to pure research than "revolutionizing drug discovery", as is frequently batted around on reddit, HN and the press.
- stupidcar 5y agoThe model parameters are only available for non-commercial use. That's a shame, as I presume there might be a lot of medical startups that would benefit from having this kind protein-folding tech available.
- mikewarot 5y agoUnless I'm mistaken, you could train the model yourself, starting with a random set of values. In time, your error rates would be low enough to have a new set of parameters which you could use however you like.
- stupidcar 5y agoYep, but there's a couple of problems. Firstly, AFAIK Deepmind haven't made all the code and settings they used to train the model available (although the paper does describe the architecture). Secondly, training a machine-learning model of this complexity is generally much more expensive, in terms of time and compute requirements, than using the resulting model. If you're a a medical startup, having an off-the-shelf prediction model you can just start using for all your protein folding needs is a very different proposition from having to train one yourself from scratch. That said, hopefully other researchers and institutions will take Google's research and produce an equivalently powerful model but with a more commercially-friendly open-source license. From some comments in this thread, it sounds like that's already happening, in fact.
- COGlory 5y agoI am a structural biologist. This is one of the handful of topics that overlaps with my field here. I'm very excited to play with this, although it might eventually put me out of a job.
- rllearneratwork 5y agowhy would it put you out of job? Wouldn't it just become one of the tools you use?
- dekhn 5y agoIt would both become a tool he used (to produce initial structures to fit in density maps) and a tool that used his or her output (because alphafold requires known protein structures that are homologous to the one you're predicting).
- nikhilsimha 5y agoThe implicit assumption you are making is that the demand increases in lock step with productivity gains. 100x faster drug discovery, 100x more drugs need to be discovered => same number of people employed. These correlations do hold for technical fields, but logically there should be a point beyond which productivity gains outpace, demand growth / demand could even stop growing. One should either retool to solve a newer problem before this point is reached, or hope that the point is not reached in the span of their career. Oil rig builders for example - manufacturing has been increasingly automated, but the demand for oil rig building has grown consistently. But they should probably look into solving other problems given that demand is shifting.
- sbierwagen 5y ago>but logically there should be a point beyond which productivity gains outpace The limiting factor on drug approval is clinical trials. Once every living person is enrolled in a clinical trial, we will have hit the maximum rate at which humanity can produce new drugs. That might be more than 10x the current rate, but probably less than 1000x.
- thesausageking 5y agoThe PDF is linked in the article: https://www.nature.com/articles/s41586-021-03819-2_reference.pdf https://www.nature.com/articles/s41586-021-03819-2_reference...
- Cas9 5y agoHonest question: since AlphaFold doesn't really _solve_ the protein folding problem (it's NP-complete after all), but only _approximates_ solutions very well, what are the real impacts of this? Isn't a good approximation of a protein enough to cause unexpected problems? How do we know that an approximate structure will perform the same as the correct solution?
- radus 5y agoYes, it is still useful. Even structures obtained through traditional means (eg. x-ray crystallography) are approximations to an extent since there are limits to the resolution that you can obtain and oftentimes regions of proteins are "disordered". Additionally, these structures are only snapshots of a protein in a particular state, which may not completely reflect the dynamics of the protein in its native environment.
- whimsicalism 5y agoYou want to find a protein that has X structure (since structure determines function to a degree). If AlphaFold is substantially more accurate at solving proteins, it can mean that drug discovery is faster, assays are faster, etc. etc. The "unexpected problems" would be caught in the assay stage.
- radus 5y agoKind of disagree with this.. solving protein structures is not the rate limiting step in drug discovery or in biochemical assays -- not by a long shot. See this excellent comment by @dekhn on a related submission: https://news.ycombinator.com/item?id=27849046 https://news.ycombinator.com/item?id=27849046
- dekhn 5y agoThe protein folding problem is not NP complete. The "formal" protein folding problem, as posed (find the set of dihedral angles whose resulting structure has the lowest energy) might be, but that bears only a distant resemblance to how people "solve" the problem today. At the very least, the statement is incorrect because many proteins don't actually fold to their energy minimum, they get stuck in kinetic traps, and the formal PF defintion never accomodated that idea.
- Cas9 5y agoHonest question: since AlphaFold doesn't really _solve_ the protein folding problem (it's NP-complete after all), but only _approximates_ solutions very well, what are the real impacts of this? Isn't a good approximation of a protein enough to cause unexpected problems? How do we know that an approximate structure will perform the same as the correct solution?
- Ultimatt 5y agoThere is a lot of bias in the chat here from a more chemistry and pharma slant. If you ignore this AlphaFold solves in a very meaningful way the problem blocking a lot of science investigation. For comparative and evolutionary analysis structure is far more conserved than sequence. Especially in things like viruses or anything with a high rate of reproduction like bacteria. Just knowing the general fold or overall structure is enough to do structural alignment and tell if two genes are related on that basis, even if their genomic sequence is completely dissimilar. Large groups of researchers rely on sequence homology built from sequences of known structure. But AlphaFold works well in new sequence space to far more accuracy than is needed. If we had an AlphaFold prediction for every known sequence suddenly the evolutionary relationships between all genes and even all species would be far clearer. This on its own unlocks a new foundation to reason about function and molecular interaction with a wholistic systems view without gaps in what we can know with some reasonable assurance. For an analogy think of the difference between having books in different languages describing objects. You know what some of the book in English might say but you dont even know if the book in Spanish is even talking about the same things. AlphaFold is like an AI that transforms all the books into picture books and now we can use image similarity or have one person look at all pictures.
- haihaibye 5y ago> even if their genomic sequence is completely dissimilar I think you mean amino acid homology? (due to synonymous mutations) I looked it up and you're right, protein structure/motifs are much more highly conserved than amino acid sequence https://humgenomics.biomedcentral.com/articles/10.1186/1479-7364-6-10#ref-CR3 https://humgenomics.biomedcentral.com/articles/10.1186/1479-...
- culopatin 5y agoDoes anyone know if this can be made to work with rna fold?
- jfengel 5y agoSo... is it possible to clone this and turn it into a Folding@Home client? How does it do?
- dekhn 5y agono, it wouldn't make sense to do that. Folding@Home is for ab initio where you don't have any prior info for the structure, this is for homology modelling. F@H probes the dynamics of protein folding, this just makes a static prediction.
- kmckiern 5y agoWhere there isn't an available crystal structure, Alphafold can be used to create initial structures for simulation via folding@home, replacing older homology modeling techniques. Source: former folding@home researcher.
- duckerude 5y ago> The AlphaFold parameters are made available for non-commercial use only, under the terms of the Creative Commons Attribution-NonCommercial 4.0 International (CC BY-NC 4.0) license. You can find details at: https://creativecommons.org/licenses/by-nc/4.0/legalcode https://creativecommons.org/licenses/by-nc/4.0/legalcode Does CC BY-NC actually do this? As far as I can tell it only really talks about sharing/reproducing, not using. Or is the only thing prohibiting other commercial use the words "available for non-commercial use only"?
- sillysaurusx 5y agoArtbreeder has some interesting prior art here: nVidia forbid commercial use of StyleGAN, but artbreeder disregarded it and happily sold all the breeding you wanted. No one seemed to care. I suspect that the clause is there to prevent a startup launching on the basis of “see this trained model? Yeah, that’s literally our business model” though, which is a mildly amusing thought, wot wot. So basically, a few tens of thousands, sure. A few million, big G might have a problem. Still, the smart move would be to launch the business anyway, and gamble that you can work out a licensing deal.
- mikewarot 5y agoIf you took their parameters, then trained it for while on a different set of data, it would vary from the original. I wonder how much compute would be required to make the offset far enough to hold up from scrutiny, and in court. Alternatively, you could manually change the network model, add a few hidden layers, etc... modifying the parameters in step, and result in a new model and new parameters. Some training to vary the parameters, and it's now a new work.
- l33tman 5y agoI would hazard a guess here that taking the parameters and continuing training constitutes "using" the parameters. Then when you get the subpoena you would have to explain the thousands of emails and slack messages discussing how you extend their parameters... :)
- dekhn 5y agoI missed an important detail: """an academic team has developed its own protein-prediction tool inspired by AlphaFold 2, which is already gaining popularity with scientists. That system, called RoseTTaFold, performs nearly as well as AlphaFold 2, and is described in a paper in Science paper also published on 15 July""" One of the things I say about CASP has to be updated. It used to be "2 years after Baker wins CASP, the other advanced teams have duplicated his methods and accuracy, and 4 years after, everything Baker did is now open source and trivially reproducible" now, it's baker catching up to DeepMind and it took about a year https://doi.org/10.1126/science.abj8754 https://doi.org/10.1126/science.abj8754
- radus 5y agoVery cool! Great to see this competition between academia and industry yielding improvements on all fronts.
- lukeplato 5y agoI see an interesting geometric similarity between the two models, namely the attention mechanism that learns relationships between structure in embedded reference frames (i.e. 1D,2D embeddings in RoseTTaFold or "local" frames in Alphafold) and the true structure of the intrinsic space reference frames (3D coordinates in RoseTTaFold and "global" frames in Alphafold)
- devindotcom 5y agoAlso announced today was RoseTTAFold from UW's Baker Lab, which claims nearly the same accuracy at much higher efficiencies. There's a public server and paper in Science. More info here and here: https://www.bakerlab.org/index.php/2021/07/15/accurate-protein-structure-prediction-accessible/ https://www.bakerlab.org/index.php/2021/07/15/accurate-prote... https://techcrunch.com/2021/07/15/researchers-match-deepminds-alphafold2-protein-folding-power-with-faster-freely-available-model/ https://techcrunch.com/2021/07/15/researchers-match-deepmind...
- deleted 5y ago[deleted]
- ehsankia 5y agoCould it be that AlphaFold 2 was open sourced in response to this?
- dekhn 5y agoit's very likely the baker submission to science forced DM's hand.
- xvilka 5y agoThat is the way, unlike AlphaFold, to publish everything in open source. Kudos to the research team!
- xor99 5y ago> With RoseTTAFold, a protein structure can be computed in as little as ten minutes on a single gaming computer. I guess its a little less accurate but the quick compute time makes as much difference too. E.g. research students can have multiple less costly mistakes before achieving what they want with the software.
- deleted 5y ago[deleted]
- mensetmanusman 5y agoDistribution of this 2 TB file seems like a good use of torrent…
- jerven 5y agoWorking with one of the team providing the uniref dataset used here. Running any kind of torrent stuff in an university network setting is a central policy fight one just does not want to get into at all. On the other hand outbound network traffic from an university is "free". So the benefit is absolutely minimal from a hosting perspective. It was tried (https://journals.plos.org/plosone/article?id=10.1371/journal.pone.0010071 https://journals.plos.org/plosone/article?id=10.1371/journal...) but it is gone the way of the dodo for the above reasons.
- hermitsings 5y agofodl
- tdfirth 5y agoThis isn't a criticism - I'm just curious to hear people's thoughts on this. When I look at this code, one of my initial reactions is that it does not seem to be very thoroughly tested. Sure, certain modules have been tested (e.g. `model.quat_affine`) but it's not clear how completely. Meanwhile, other modules, for example `model.folding`, have not been tested at all, despite containing large amounts of complex logic. That kind of code that works with arrays is very easy to get wrong and bugs are difficult to spot. My experience working with code written by researchers is that it frequently contains a large number of bugs, which brings the whole project into question. I've also found that encouraging them to write tests greatly improves the situation. Additionally, when they get the hang of testing they often come to enjoy it, because it gives them a way to work on the code without running the entire pipeline (which is a very slow feedback loop). It also gives them confidence that a change hasn't lead to a subtle bug somewhere. Again, I'm not criticising. I am aware that there are many ways to produce high quality software and Google/DeepMind have a good reputation for their standards around code review, testing etc. I am, however, interested to understand how the team that wrote this think about and ensure accuracy. In general, I hope that testing and code review become a central part of the peer review process for this kind of work. Without it, I don't think we can trust results. We wouldn't accept mathematical proofs that contained errors, so why would we accept programs that are full of bugs? edit: grammar
- benschulz 5y agoMy understanding is that it has been manually tested. I.e. it has produced correct results to previously intractable problems. I'm not sure how much automated testing would add at that point.
- dmos62 5y agoUnit testing usually isn't easily replaced by manual testing. If you have, for example, 3 units that can be in 2 different modes each, that's 2^3 different combinations, but only 2*3 unit modes. Testing the end result is more work than testing the units.
- 5y ago