8 ms·
AlphaFold Protein Structure Database
- deleted 5y ago[deleted]
- moyix 5y agoAnyone else getting a 403 Forbidden? If so it might be better to link to the paper instead: https://www.nature.com/articles/s41586-021-03828-1 https://www.nature.com/articles/s41586-021-03828-1
- jkh1 5y agoWorks fine for me. Must have been a temporary glitch.
- jkh1 5y agoDidn't see this post so posted it also. Also relevant: https://www.embl.org/news/science/alphafold-potential-impacts/ https://www.embl.org/news/science/alphafold-potential-impact...
- ricksunny 5y agoI’m sorry but why don’t tbey just release the ability for a user to enter a known real-world sequence’s accession number from Genbank / GISAID, and generate the protein structure from that? Why do they have to abstract the user from the process by only exposing a completed database of the protein structures the Alphafold researchers decided would be worth producing?
- sherjilozair 5y agoDeepMind has already released the open source code and model parameters. The database makes it easier to access the predictions.
- sveme 5y agoI'd guess the ad-hoc simulation of the structure is computationally quite expensive and takes a while, though that's just a guess and I haven't read the original paper yet.
- ricksunny 5y agoIn fact a cost of $1-$4 for the preferred implementation: https://news.ycombinator.com/item?id=27894060 https://news.ycombinator.com/item?id=27894060 The colab provides a slightly-less-accurate version that operates in the cloud. For the real mccoy it seems one must set up one’s own environment and leverage the git repo.
- tazjin 5y agoYou can use the open-source code, and we also have a Colab notebook for that: https://bit.ly/alphafoldcolab https://bit.ly/alphafoldcolab More info: https://deepmind.com/blog/article/putting-the-power-of-alphafold-into-the-worlds-hands https://deepmind.com/blog/article/putting-the-power-of-alpha...
- ricksunny 5y agoThanks for that - I can see why my comment was downvoted now, as the the posted article's FAQ lists these links for those who would like to study their favorite sequenced-but-unmodeled protein. I'm glad Alphafold is as open source as it is, and I recognize that it didn't have to be so. I think I was primed for a knee-jerk reaction because when Alphafold's results were announced back in Dec. 2020, with expressions of what a boon it would be for researchers around the globe, I anticipated there would be a timeline announced for exposing a tool or for the open-sourcing. (The Github repo has only just been released about 6 days ago ...) With all the work on SARS-CoV-2's 'interactome', as well as human proteins & enzymes involved in pharmacology of antiviral drugs under development / repurposing , it's easy to imagine that drug developers would have liked to exercise Alphafold as soon as it was announced. (I myself have wanted a structure for human enzyme OATP1A2 that wasn't available on the PDB for such a drug pharmacology study - quite glad it is available at hand now.. .:) ). Anyway I'm sure good arguments will be made about the need to really 'get it right' before releasing, or internal deliberations on how much to open up vs charging for it. But 7 months lead time during a pandemic is a long time... In all cases thanks again for this innovation's availability now. :)
- ricksunny 5y agoA little addendum (for posterity as this is now an old article post) • RosettaFold came out in academic paper form as well as open-source github repo slightly prior to AlphaFold. Was AlphaFold's decision to open up motivated by RosettaFold's publishing / opening-up activity? So I feel that the extent of AF's altruism in this (while high) deserves some scrutiny, as perhaps it falls short of the extent that several of us commenting interpreted a couple days ago. • While RosettaFold allows online execution to the best of their spec, they currently have a backlog queue of ~3000 jobs going back to mid-July. This is telling about the processing power required for folding. AlphaFold has access to a lot of processing power (even if it could ultimately & reasonably end up being on a charge-per-job basis).
- nharada 5y agoFrom the abstract[1]: > After decades of effort, 17% of the total residues in human protein sequences are covered by an experimentally-determined structure. Here we dramatically expand structural coverage by applying the state-of-the-art machine learning method, AlphaFold2, at scale to almost the entire human proteome (98.5% of human proteins). [1] https://www.nature.com/articles/s41586-021-03828-1 https://www.nature.com/articles/s41586-021-03828-1
- vmception 5y agoBasically they are saying that decades of distributed protein folding was useless and everyone would have had more utility mining cryptocurrency if it existed several years earlier But at least it inspired someone to make and release this
- cing 5y agoJust in case you're not joking, it's worth noting that the majority of distributed molecular simulation (past and present) is spent studying "folded proteins" to discover structures of proteins that are often hidden from methods like AlphaFold (currently). For example, https://www.nature.com/articles/s41557-021-00707-0 https://www.nature.com/articles/s41557-021-00707-0
- dmitryminkovsky 5y ago> experimentally-determined structure refers to structures determined by means of physical examination, with like crystallography, not to attempts at predictive computational analysis prior to AlphaFold, which were not accurate compared to AlphaFold.
- Jabbles 5y agoA third way in which you are wrong is that AlphaFold derives a lot of its power by referring to previously-solved protein structures, or parts of them. It doesn't fold the proteins from scratch in an "alpha-zero" way.
- vmception 5y ago
- sdbrown 5y agoThis is a fabulous convenience! The reach of this ready-to-go data will be much larger (in some directions) than the model and CASP results themselves.
- stephanheijl 5y agoI'm impressed and grateful that DeepMind released this resource, this will save a lot of compute from labs trying to replicate an entire exome for themselves. While some structures look great, there are still some misses here. Important structures like BRCA1 (a well-studied breast cancer associated protein) are just structures for the BRCT and RING domains surrounded by a low-confidence string of amino acids, likely shaped to be globular: https://alphafold.ebi.ac.uk/entry/P38398 https://alphafold.ebi.ac.uk/entry/P38398 Maybe I was wrong for expecting the impossible here, but I was excited to see this specific structure and it appears that there is still work to do. Nevertheless, kudos to Deepmind on their amazing achievement and contributions to the field!
- cing 5y agoEverything between the BRCT and RING domains of BRCA1 is an intrinsically unstructured region which DeepMind correctly predicts, https://pubmed.ncbi.nlm.nih.gov/15571721/ https://pubmed.ncbi.nlm.nih.gov/15571721/ Another famous one would be R-domain of CFTR, which was not resolved in experimental structure determination, and AlphaFold models correctly show disorder there. Nothing to be done in those cases except perform molecular simulation or other experiments to assess dynamic ensembles, https://alphafold.ebi.ac.uk/entry/P13569 https://alphafold.ebi.ac.uk/entry/P13569
- maga 5y agoA curious non-biologist here: how valuable are these low confidence predictions for biologists? In other words, is it hard to predict but easy to check situation as with, say, prime numbers in mathematics?
- toufka 5y agoThe medium-confidence predictions are great for grounding or sourcing intuition. If you're trying to divide up a protein for an experiment and you have to choose where to divy it up - you'd like to use even a bad prediction to help weight an otherwise completely random approach. AND there are great methods to help with this, but they're often custom, time-consuming, and out-of-field for most. So being able to very quickly spot-check using a uniform state-of-the art, for any arbitrary protein, makes it actually pretty useful for certain kinds of pre-experimental guidance.
- dnautics 5y agoyikes, this doesn't even do some basic stuff like trim off pre-protein segments for secreted proteins... Without this, you could get some very incorrect structures.
- deleted 5y ago[deleted]
- visarga 5y agoCitation factory, that's what it is.
- abcc8 5y agoResources as useful as this are bound to be. We do cite our sources after all.
- Ovah 5y agoInteresting that they're porting it to other organisms. Different organisms have variations in ribosomes, post translational modifications and even tRNA repertoire. So it's not a guarantee that two identical DNA sequences will give identical proteins in two different organisms.
- ramraj07 5y ago??? Unless you jump from eukaryotes to archea these are not real concerns. Most PTM markers are very conserved.
- Ovah 5y agoI'd say the jump from eukaryotes to procaryotes is a realistic scenario in recombinant DNA technology. I have some experience with recombinant yeast and PTMs. Degree of glycosylation actually vary a lot depending on strain used and has a huge effect of protein activity. And of course these PTMs affects the crystal structure.
- pelorat 5y agoShouldn't matter? Protein folding is based on the laws of physics after all. If DNA sequences folds differently in different organisms then an external factor is missing.
- Ovah 5y agoWhile the laws of physics remain the same, the folding machinery between species varies to some degree. Protein folding is determined by the unique environment/machinery of a cell. A concrete example is disulphide bonds (S-S, ex cystein-cystein) that require a certain pH to form. The primary pathways of disulphide-bond formation are localized in the endoplasmic reticulum (ER) of eukaryotic cells and the periplasmic space of prokaryotic cells. So two complete different mechanisms to end up with the same bond (protein structure) depending on the organism.
- dnautics 5y ago
- narrator 5y agoGain of function researchers working for the world's militaries will use this research to figure out how to get viruses to attach to receptor sites peculiar to particular races. The people developing the antivirals will have a lot harder time countering these weapons because making antivirals that aren't poisonous in some weird way is a much harder job. If this is not the case, please let me know why, it will really help me sleep better at night. A U.S congressional representative came out of a classified briefing recently and announced that the CCP is hard at work on race specific bioweapons.[1] Unfortunately, I think this is the launch of a new era of weapons we're seeing right now. The biggest development in war since the atom bomb. Like the atom bomb, the big question was will we kill ourselves with this technology. Who knows? [1]https://yournews.com/2021/07/22/2185645/rep-marjorie-taylor-greene-says-she-was-briefed-about-chicoms/ https://yournews.com/2021/07/22/2185645/rep-marjorie-taylor-...
- drcode 5y ago...and many doctors will use it to attach pharmaceuticals to receptor sites of particular cancers.
- narrator 5y agoI'm thinking that the problem is is that it is much harder to develop drugs that only kill cancers very efficiently and don't harm the rest of the body than to tweak viruses that just have to keep the person alive long enough to spread the virus.
- drcode 5y agoI 100% agree your point is valid. The counterargument is "Yes, people can do bad things with protein data, just as they can do bad things with a telephone, like use it to discuss a bank robbery."
- narrator 5y agoThe crazy part is a bioweapons program is really cheap compared to a nuclear weapons program, and now with these new tools it's even cheaper. Before, it was vastly more expensive to do the cycle of creating a new viral protein and testing a bioweapon on human cell culture. Now that process is speeded up millions of times with this technology because that can all take place inside a computer. This is similar to the change with drone weaponry. Before, you had to have large cruise missiles to get pinpoint strikes. Now small countries like Azerbaijan can buy a whole fleet of drone weapons and get the benefits of having a modern air force with pinpoint strikes and even stealth for vastly less money.
- ramraj07 5y agoAs an ex biomedical researcher I was trying to think what protein I should enter and see, and couldn't come up with a protein that I know of, that didn't have a structure already (at least a crude one). That is, we roughly know how most known important proteins look like. This is an amazing tool, and will he indispensable in labs (I'll expect any lab to use this site at least once a year?) But it's not as transformative as some might think.
- amelius 5y agohttps://www.embl.org/news/science/alphafold-potential-impacts/ https://www.embl.org/news/science/alphafold-potential-impact... > A discussion of the applications that AlphaFold DB may enable and the possible impact of the resource on science and society
- pelorat 5y agoDo we really know the structure of every protein that assembles into a human cell?
- seventytwo 5y agoDefinitely not.
- _RPL5_ 5y agoFrom their abstract: --- After decades of effort, 17% of the total residues in human protein sequences are covered by an experimentally-determined structure1. Here we dramatically expand structural coverage by applying the state-of-the-art machine learning method, AlphaFold2, at scale to almost the entire human proteome (98.5% of human proteins). The resulting dataset covers 58% of residues with a confident prediction, of which a subset (36% of all residues) have very high confidence. https://www.nature.com/articles/s41586-021-03828-1 https://www.nature.com/articles/s41586-021-03828-1 --- The metric they use (residues) is a bit unusual (I would have used number of proteins instead), but I assume they wanted to account for ambiguity (such as proteins with partial structures).
- cing 5y ago
- lumost 5y agoI used to do some RNA molecular dynamics simulations in college which were both computationally expensive and difficult to replicate. Having the ability to reasonably predict protein structure is an incredible scientific achievement - however I am curious if anyone here who is better informed has takes on the following. 1. How likely is it that alphafold learned to accurately predict protein structure in the narrow domain of proteins that have been experimentally synthesized and whose structure has been measured? in other words will AlphaFold's results generalize to proteins which cannot yet be synthesized in the laboratory. 2. If Alphafold's accuracy holds, what type of commercial applications does this open up?
- culopatin 5y agoI happen to be working on a database for folds as well. But RNA folds not protein folds. I’m not a bio guy but my gf is and if I understand correctly this is not the same. I hope they are different because it would suck to be me lol. This is my first big boy project and I’m driving solo so it takes me a while to make any progress. But at least now I have this db and genbank to model after
- pelorat 5y agoThere's a lot of news about AlphaFold lately but what about Rossettafold? Wasn't it more accurate and much faster?
- spacecity1971 5y agoQuick question, please excuse my ignorance, but is there a way to extrapolate sequence from structure? In other words, can we design proteins and calculate the sequence required to make it?
- kmckiern 5y agoIt's hard but people do it! This is the field of "protein engineering".
- _RPL5_ 5y agoThis is awesome! When they announced CASP results a few months ago, I was wondering if AlphaFold will be accessible as an API, where you can submit a protein id or a sequence and get back a 3D structure. This database is basically that, except it's free & open to the public. Major props!