19 ms·
AlphaFold reveals the structure of the protein universe
- bifftastic 4y agoHow do they know their structures are correct?
- seydor 4y agothey don't but they are more correct than what others have predicted. Some of their predictions can be compared with structures determined with x-ray crystallography
- cupofpython 4y agodid they come up with their structures independently of the x-ray crystallography, or was that part of a ML dataset for predicting structure
- unlikelymordant 4y agoThe casp competition that they won consists of a bunch of new proteins, the structures of which havnt been published. So the test set is for brand new proteins in that case.
- cupofpython 4y agonice, very cool then
- tomrod 4y agoThis is the right line of questioning. As we solve viewability into the complex coding of proteins, we need to be right. Next, hopefully, comes causal effect identification, then construction ability. If medicine can use broad capacity to create bespoke proteins, our world becomes both weird and wonderful.
- __rito__ 4y agoThey won a decades-long standing challenge predicting the protein structures of a much smaller (yet significantly quite large) set of proteins using a model (AlphaFold). Then they use the model to predict more. Although we don't know if they are correct, these structures are the best (or the least bad) we have for now.
- luma 4y agoSame as any other prediction I'd presume. Run it against a known protein and see how the answer lines up. Predict the structure of an unknown protein, then use traditional methods (x-ray crystallography, maybe STEM, etc) to verify.
- christudor 4y agoThis is exactly right.
- gilleain 4y agoAs a simple example, one measure used to compare a predicted structure against a reference is the RMSD (root mean square deviation). https://en.m.wikipedia.org/wiki/Root-mean-square_deviation_of_atomic_positions https://en.m.wikipedia.org/wiki/Root-mean-square_deviation_o... The lower the RMSD between two structures, the better (up to some limit).
- iandanforth 4y ago"Verify" is almost correct. The crystallography data is taken to be "ground truth" and the predicted protein structure from AlphaFold is taken to be a good guess starting point. Then other software can produce a model that is a best fit to the ground truth data starting from the good guess. So even if the guess is wrong in detail it's still useful to reduce the search space.
- lrem 4y agoDisclaimer: I work in Google, organizationally far away from Deep Mind and my PhD is in something very unrelated. They can't possibly know that. What they know is that their guesses are very significantly better than the previous best and that they could do this for the widest range in history. Now, verifying the guess for a single (of the hundreds of millions in the db) protein is up to two years of expensive project. Inevitably some will show discrepancies. These will be fed to regression learning, giving us a new generation of even better guesses at some point in the future. That's what I believe to be standard operating practice. A more important question is: is today's db good enough to be a breakthrough for something useful, e.g. pharma or agriculture? I have no intuition here, but the reporting claims it will be.
- f38zf5vdt 4y agoThe press release reads like an absurdity. It's not the "protein universe", it's the "list of presumed globular proteins Google found and some inferences about their structure as given by their AI platform". Proteins don't exist as crystals in a vacuum, that's just how humans solved the structure. Many of the non-globular proteins were solved using sequence manipulation or other tricks to get them to crystallize. Virtually all proteins exist to have their structures interact dynamically with the environment. Google is simply supplying a list of what it presumes to be low RMSD models based on their tooling, for some sequences they found, and the tooling is based itself on data mostly from X-ray studies that may or may not have errors. Heck, we've barely even sequenced most of the DNA on this planet, and with methods like alternative splicing the transcriptome and hence proteome has to be many orders of magnitude larger than what we have knowledge of. But sure, Google has solved the structure of the "protein universe", whatever that is.
- lrem 4y agoI recognize your superior knowledge in the topic and assume you're right. But you also ignore where we're at in the standard cycle: https://phdcomics.com/comics/archive_print.php?comicid=1174 https://phdcomics.com/comics/archive_print.php?comicid=1174 ;)
- christudor 4y agoThis video goes some way to explaining how they know the structures are correct: https://www.youtube.com/watch?v=vXZzftX03VY https://www.youtube.com/watch?v=vXZzftX03VY
- ArnoVW 4y agoWe know the structure of some proteins. It's not that it's impossible to measure, it's just very expensive. This is why having a model that can "predict" it is so useful.
- DevX101 4y agoThey compare the predicted structure (computed) to a known structure (physical x-ray crystallography). There's an annual competition CASP (Crtical Assessment of protein Structure Prediction) that does X-Ray crystallography on a protein. The identity of this protein is held secret by the organizers. Then research teams across the world present their models and attempt to predict without advance knowledge, the structure of the protein from their amino acid sequence. Think of CASP as a validation data set used to evaluate a machine learning model. DeepMind crushes everyone else at this competition.
- liuliu 4y agoThe worry is about dataset shifting. Previously, the data were collected for a few hundreds thousands structures, now it is 200m. I think there could be doubts on distributions and how that could play a role in prediction accuracy.
- dalbasal 4y agoCan someone put AlphaFold's problem space into perspective for me? Why is protein folding important? Theoretical importance? Can we do something with protein folding knowledge? If so, what? I've been hearing about AlphaFold from the CS side. There they seem to focus on protein folding primarily as an interesting space to apply their CS efforts.
- fabiospampinato 4y agoYou are basically made of proteins, which are basically folded sequences of amino acids, proteins are molecular machines that are the fundamental building block of animals, plants, bacteria, fungi, viruses etc. So yeah the applications are enormous, from medicine to better industrial chemical processes, from warfare to food manufacturing.
- jebarker 4y ago> proteins are molecular machines Does that imply proteins have some dynamics that need to be predicted too? I remember seeing animations of molecular machines that appeared to be "walking" inside the body - are those proteins or more complex structures?
- gilleain 4y agoYes, very much so. Even for proteins that seems like they are just scaffolding for a catalytic centre can have important dynamics. A classic example is haemoglobin, that 'just' binds to oxygen at the iron in the middle of the haem. Other binding sites remote from the oxygen binding one can bind to other molecules - notably carbon dioxide. The 'Bohr effect' mechanism is outlined here : https://en.wikipedia.org/wiki/Bohr_effect#Allosteric_interactions https://en.wikipedia.org/wiki/Bohr_effect#Allosteric_interac... Even at the lowest level, there is some evidence that ligand binding can affect the structure of the backbone of the protein. For example, peptide plane flipping https://en.wikipedia.org/wiki/Peptide_plane_flipping https://en.wikipedia.org/wiki/Peptide_plane_flipping although I'm not sure where the research is on this nowadays.
- gz5 4y ago>we’re now releasing predicted structures for nearly all catalogued proteins known to science is the result that researchers will now much more quickly 'manually' validate or invalidate the predicted structures for proteins they are working with? i understand it is traditionally a long and complex process, but i imagine it is expedited by having a predicted structure to test as the baseline?
- epups 4y agoAlphaFold is a phenomenal tool that demonstrates how AI can already outclass humans for certain tasks. It is a prime example of a problem space where conventional approaches are simply inferior, and that AI is not just a fancy name but can be extraordinarily powerful.
- hackernewds 4y agoThis reads like something created by an AI. And the mere fact that I can't be sure if it is or not now is testament to the statement this is making
- epups 4y agoIt's a new tool, AlphaFool
- desindol 4y agoIt's the superstition and the amount of parasite words.
- jcims 4y agoJust for shits & giggles. Prompt: Write a short abstract for an article on why AlphaFold isn't just a tremendous technological achievement, but an example of how Artificial Intelligence will allow humanity to continue to achieve technological progress in the future: Response: When DeepMind’s AlphaFold was announced as the winner of the CASP13 protein folding prediction competition, it was hailed as a major breakthrough in both artificial intelligence (AI) and biology. While AlphaFold is undoubtedly a remarkable achievement, it is also a clear example of how AI can be used to solve complex problems that have traditionally been seen as beyond the reach of machines. In particular, AlphaFold represents a significant advance in the ability of AI systems to learn from limited data. Unlike previous AI approaches to protein folding, which relied on large datasets of known protein structures, AlphaFold was able to produce accurate predictions with only a small number of known protein structures to work from. This ability to learn from limited data is likely to be increasingly important as we face problems that are too complex for us to fully understand and collect all the data we would need to solve them using traditional methods. In the future, AI systems like AlphaFold will become increasingly important in helping us to make progress on a wide range of pressing challenges, from developing new medicines to tackling climate change.
- candiddevmike 4y agoIs folding@home obsolete now?
- seventytwo 4y agoGood question… I’d imagine that other methods of folding solutions are still valuable, because AlphaFold needs to be checked.
- dekhn 4y agoIt's not, but the question is (and has long been) whether the energy expended by folding@home is worth the scientific result. IMHO- probably not.
- foxhop 4y agoI would say no, the two approaches may be used to validate each other.
- flobosg 4y agoFolding@home answers a related but different question. While AlphaFold returns the picture of a folded protein in its most energetically stable conformation, Folding@home returns a video of the protein undergoing folding, traversing its energy landscape.
- jarenmf 4y agoThis is probably one of the best applications of AI in science in terms of impact so far. I can't think of any other problem with the same potential impact. EDIT: grammar
- t00 4y agoYou are right and when thinking about it I can see 2 problems which I hope in the future can have even more impact: 1. Using AI to determine the most efficient methods of doing mathematical expressions, transformations and computation algorithms - division, square root, maybe traveling salesman - these which take relatively high amount of CPU cycles to compute and are used everywhere. If inputs and outputs can be assigned to it, AI can eventually build a transformation which can be reproduced using a silicon. 2. Physics phenomena in general, not only organic protein, can be measured and with sufficient ability to quantize them to inputs and experimentally obtained outputs to train the network, we could in theory establish new formulas or constants and progress the understanding of the Universe.
- lrhegeba 4y agothe groundworks, at least partially, happen as you typed this: https://www.nature.com/articles/d41586-021-01627-2 https://www.nature.com/articles/d41586-021-01627-2
- 323 4y agoAI translate has probably a bigger worldwide impact so far.
- jebarker 4y agojarenmf said "in science" - but it is an interesting question how much automated translation has helped scientists translate papers from other languages.
- hijodelsol 4y agoIt even goes both ways - it allows non-native English speakers to publish their work in correct technical/scientific English with far less barriers.
- carbocation 4y agoThe press release is a bit difficult to place into historical context. I believe that the first AlphaFold release was mostly human and mouse proteins, and this press release marks the release of structures for additional species.
- azangru 4y ago> I believe that the first AlphaFold release was mostly human and mouse proteins, More than that. The press release actually contains an infographic comparing the amount of published protein models for different clades of organisms. The infographic shows that the previous release (~1mln proteins) contained proteins of some animal, plant, bacterial, and fungal species.
- deleted 4y ago[deleted]
- kache_ 4y agoThis is an incredible gift to humanity. A huge positive impact. The team should be proud
- klemola 4y agoAs an aside, the protein structure visualizations in the article are pretty. Is there a good source for more?
- alphabetting 4y agohttps://alphafold.ebi.ac.uk/ https://alphafold.ebi.ac.uk/
- flobosg 4y ago* https://pdb101.rcsb.org/motm/ https://pdb101.rcsb.org/motm/ * https://ccsb.scripps.edu/goodsell/ https://ccsb.scripps.edu/goodsell/ * https://pdb101.rcsb.org/sci-art/geis-archive/irving-geis https://pdb101.rcsb.org/sci-art/geis-archive/irving-geis * https://www.digizyme.com/portfolio.html https://www.digizyme.com/portfolio.html * https://www.drewberry.com/ https://www.drewberry.com/ * https://biochem.web.utah.edu/iwasa/projects.html https://biochem.web.utah.edu/iwasa/projects.html * http://onemicron.com/ http://onemicron.com/ * The art of Jane Richardson, of which I couldn’t find a link * This blog has plenty of good links: https://blogs.oregonstate.edu/psquared/ https://blogs.oregonstate.edu/psquared/
- dekhn 4y agoDemis and John will probably win either the Chemistry or Physics Nobel Prize in the next couple of years.
- thomasahle 4y agoSome people are using "AI wins a Nobel price" as the new Turing test. Maybe that is going to happen sooner than they expect. Or maybe the owners of the AI will always claim it on its behalf.
- dekhn 4y agothere's no AI here. This is just ML. All deepmind did here was use multiple excellent resources- large numbers of protein sequences, and small numbers of protein structures, to create an approximation function of protein structure, without any of the deep understanding of "why". Interestingly, the technology they used to do this didn't exist 5 years ago!
- swayvil 4y agoI had a dream about this a few days ago. About complexly wrinkled/crumpled/convolved things. Like a fresh crepe stuffed into the toe of a boot. Bewilderingly complex. But I have a question. Does such contortion work for 3d "membranes" in a 4d space? It's something I'm chewing on. Hard to casually visualize, obviously.
- gspr 4y agoOf course! The term you might wanna start off googling is "curvature of manifolds". What's even neater than "3d thing curving in 4d space" is that these notions can be made precise also without the "in [whatever] space" part (see "intrinsic curvature" and "Riemannian manifold").
- swayvil 4y agoThank you very much.
- alphabetting 4y agoObtaining this dataset prior to alphafold would have cost on the order of $200 trillion. https://twitter.com/wintonARK/status/1552653527670857729 https://twitter.com/wintonARK/status/1552653527670857729 Anyone knowledgeable know if this estimate is accurate? Insane if true
- green-eclipse 4y agoIt's impossible to really put a number on it, because the task itself was impossible. PHDs and the field's top scientists simply couldn't figure out many complicated protein structures after years of attempts, and the fact that there's so many (200M+) mean that the problem space is vast.
- shauryamanu 4y agoEven if that's exaggerated, it might have taken significant time to reach to this stage. Probably on the order of >50 years.
- dekhn 4y agoIt doesn't make any sense on multiple levels. This is a computational prediction and there was no computational alternative- for many of these proteins would never have had a structure solved even if you spent the money. They are just taking $cost_per_structure_solved * number_of_remaining_structures and assuming that things scale linearly like that. Note that crystallographers are now using these predicftions to bootstrap models of proteins they've struggled to work with, which indicates the level of trust in the structural community for these predictions is pretty high.
- knbknb 4y agoOff the top of my head: (200 trillion cost) / (200 million structures predicted) = 1 million per structure. That reflects the personnel cost (5 Yr PHP scholarship, PostDoc/Prof mentorship; inverstment+depreciation for the lab equipment). All this to crystallize 1 structure and characterize its folding behavior. I don't know if this calculation is too simplistic, just coming up with something.
- deleted 4y ago[deleted]
- cm2187 4y agoHow do you know that the predicted structure will be correct? I presume researchers will need to validate the structure empirically. Do we know how good the model has been at predicting so far?
- yuan43 4y ago> Today, I’m incredibly excited to share the next stage of this journey. In partnership with EMBL’s European Bioinformatics Institute (EMBL-EBI), we’re now releasing predicted structures for nearly all catalogued proteins known to science, which will expand the AlphaFold DB by over 200x - from nearly 1 million structures to over 200 million structures - with the potential to dramatically increase our understanding of biology. And later: > Today’s update means that most pages on the main protein database UniProt will come with a predicted structure. All 200+ million structures will also be available for bulk download via Google Cloud Public Datasets, making AlphaFold even more accessible to scientists around the world. This is the actual announcement. UniProt is a large database of protein structure and function. The inclusion of the predicted structures alongside the experimental data makes it easier to include the predictions in workflows already set up to work with the other experimental and computed properties. It's not completely clear from the article whether any of the 200+ million predicted structures deposited to UniProt have not be previously released. Protein structure determines function. Before AlphaFold, experimental structure determination was the only option, and that's very costly. AlphaFold's predictions appears to be good enough to jumpstart investigations without an experimental structure determination. That has the potential to accelerate many areas of science and could percolate up to therapeutics. One area that doesn't get much discussion in the press is the difference between solid state structure and solution state structure. It's possible to obtain a solid state structure determination (x-ray) that has nothing to do with actual behavior in solution. Given that AlhpaFold was trained to a large extent on solid state structures, it could be propagating that bias into its predicted structures. This paper talks about that: > In the recent Critical Assessment of Structure Prediction (CASP) competition, AlphaFold2 performed outstandingly. Its worst predictions were for nuclear magnetic resonance (NMR) structures, which has two alternative explanations: either the NMR structures were poor, implying that Alpha-Fold may be more accurate than NMR, or there is a genuine difference between crystal and solution structures. Here, we use the program Accuracy of NMR Structures Using RCI and Rigidity (ANSURR), which measures the accuracy of solution structures, and show that one of the NMR structures was indeed poor. We then compare Alpha-Fold predictions to NMR structures and show that Alpha-Fold tends to be more accurate than NMR ensembles. There are, however, some cases where the NMR ensembles are more accurate. These tend to be dynamic structures, where Alpha-Fold had low confidence. We suggest that Alpha-Fold could be used as the model for NMR-structure refinements and that Alpha-Fold structures validated by ANSURR may require no further refinement. https://pubmed.ncbi.nlm.nih.gov/35537451/ https://pubmed.ncbi.nlm.nih.gov/35537451/
- crispyambulance 4y agoI got a 5th grader question about how proteins are used/represented graphically that I've never been able to find a satisfying answer for. Basically, you see these 3D representations of specific proteins as a crumple of ribbons-- literally like someone ran multi-colored ribbons though scissors to make curls and dumped it on the floor (like a grade school craft project). So... I understand that proteins are huge organic molecules composed of thousands of atoms, right? Their special capabilities arise from their structure/shape. So basically the molecule contorts itself to a low energy state which could be very complex but which enables it to "bind?" to other molecules expressly because of this special shape and do the special things that proteins do-- that form the basis of living things. Hence the efforts, like Alphafold, to compute what these shapes are for any given protein molecule. But what does one "do" with such 3D shapes? They seem intractably complex. Are people just browsing these shapes and seeing patterns in them? What do the "ribbons" signify? Are they just some specific arrangement of C,H,O? Why are some ribbons different colors? Why are there also thread-like things instead of all ribbons? Also, is that what proteins would really look like if you could see at sub-optical wavelength resolutions? Are they really like that? I recall from school the equipartition theorem-- 1/2 KT of kinetic energy for each degree of freedom. These things obviously have many degrees of freedom. So wouldn't they be "thrashing around" like rag doll in a blender at room temperature? It seems strange to me that something like that could be so central to life, but it is. Just trying to get myself a cartoonish mental model of how these shapes are used! Anyone?
- dekhn 4y agoThe ribbons and helices you see in those pictures are abstract representations of the underlying positions of specific arrangements of carbon atoms along the backbone. There are tools such as DSSP https://en.wikipedia.org/wiki/DSSP_(hydrogen_bond_estimation_algorithm) https://en.wikipedia.org/wiki/DSSP_(hydrogen_bond_estimation... which will take out the 3d structure determined by crystallography and spit out hte ribbons and helices- for example, for helices, you can see a specific arrangement of carbons along the protein's backbone in 3d space (each carbon interacts with a carbon 4 amino acids down the chain). Protein motion at room temperature varies depending on the protein- some proteins are rocks that stay pretty much in the same single conformation forever once they fold, while others do thrash around wildly and others undergo complex, whole-structure rearrangements that almost seem magical if you try to think about them using normal physics/mechanical rules. Having a magical machine that could output the full manifold of a protein during the folding process at subatomic resolution would be really nice! but there would be a lot of data to process.
- codedokode 4y agoToday I learned that there are bacteria that have a protein helping to form ice on plants [1] to destroy them and extract nutrients (however I didn't understand how bacteria themselves survive this). Machine learning typically uses existing data to predict new data. Please explain: Does it mean that AlphaFold can only use known types of interactions between atoms and will mispredict the structure of proteins that use not yet known interactions? And why we cannot just simulate protein behaviour and interactions using quantum mechanics? [1] https://pubs.acs.org/doi/10.1021/acs.jpcb.1c09342 https://pubs.acs.org/doi/10.1021/acs.jpcb.1c09342
- flobosg 4y ago> And why we cannot just simulate protein behaviour and interactions using quantum mechanics? QM calculations have been done in proteins, but they’re computationally very expensive. IIRC, there are hybrid approaches where only a small portion of interest in the protein structure is modelled by QM and the rest by classical molecular mechanics.
- beanwood 4y ago>And why we cannot just simulate protein behaviour and interactions using quantum mechanics? If you wanted to simulate the behaviour of an entire protein using quantum mechanics, the sheer number of calculations required would be infeasible. For what it's worth, I have a background in computational physics and am studying a PhD in structural biology. For any system (of any size) that you want to simulate, you have to consider how much information you're willing to 'ignore' in order to focus on the information you would like to 'get out' of a set of simulations. Being aware of the approximations you make and how this impacts your results is crucial. For example, if I am interested in how the electrons of a group of Carbon atoms (radius ~ 170 picometres) behave, I may want to use Density Functional Theory (DFT), a quantum mechanical method. For a single, small protein (e.g. ubiquitin, radius ~ 2 nanometres), I may want to use atomistic molecular dynamics (AMD), which models the motion of every single atom in response to thermal motion, electrostatic interactions, etc using Newton's 2nd law. Electron/proton detail has been approximated away to focus on overall atomic motion. In my line of work, we are interested in how big proteins (e.g. the dynein motor protein, ~ 40 nanometres in length) move around and interact with other proteins at longer time (micro- to millisecond) and length (nano- to micrometre) scales than DFT or AMD. We 'coarse-grain' protein structures by representing groups of atoms as tetrahedra in a continuous mesh (continuum mechanics). We approximate away atomic detail to focus on long-term motion of the whole protein. Clearly, it's not feasible to calculate the movement of dynein for hundreds of nanoseconds using DFT! The motor domain alone in dynein contains roughly one million atoms (and it has several more 'subunits' attached to it). Assuming these are mostly Carbon, Oxygen or Nitrogen, then you're looking at around ten million electons in your DFT calculations, for a single step in time (rounding up). If you're dealing with the level of atomic bonds, you're probably going to a use time steps between a femto- (10^-15 s) or picosecond (10^-12 s). The numbers get a bit ridiculous. There are techniques that combine QM and AMD, although I am not too knowledgeable in this area. Some further reading, if you're interested (I find Wikipedia articles on these topics to generally be quite good): DFT: https://en.wikipedia.org/wiki/Density_functional_theory https://en.wikipedia.org/wiki/Density_functional_theory Biological continuum mechanics: https://doi.org/10.1371/journal.pcbi.1005897 https://doi.org/10.1371/journal.pcbi.1005897 Length scales in biological simulations: https://doi.org/10.1107/S1399004714026777 https://doi.org/10.1107/S1399004714026777 Electronic time scales: https://www.pnas.org/doi/10.1073/pnas.0601855103 https://www.pnas.org/doi/10.1073/pnas.0601855103
- COGlory 4y agoBefore my comment gets dismissed, I will disclaim I am a professional structural biologist that works in this field every day. These threads are always the same: lots of comments about protein folding, how amazing DeepMind is, how AlphaFold is a success story, how it has flipped an entire field on it's head, etc. The language from Google is so deceptive about what they've actually done, I think it's actually intentionally disingenuous. At the end of the day, AlphaFold is amazing homology modeling. I love it, I think it's an awesome application of machine learning, and I use it frequently. But it's doing the same thing we've been doing for 2 decades: pattern matching sequences of proteins with unknown structure to sequences of proteins with known structure, and about 2x as well as we used to be able to. That's extremely useful, but it's not knowledge of protein folding. It can't predict a fold de novo, it can't predict folds that haven't been seen (EDIT: this is maybe not strictly true, depending on how you slice it), it fails in a number of edge cases (remember, in biology, edge cases are everything) and again, I can't stress this enough, we have no new information on how proteins fold. We know all the information (most of at least) for a proteins final fold is in the sequence. But we don't know much about the in-between. I like AlphaFold, it's convenient and I use it (although for anything serious or anything interacting with anything else, I still need a real structure), but I feel as though it has been intentionally and deceptively oversold. There are 3-4 other deep learning projects I think have had a much greater impact on my field. EDIT: See below: https://news.ycombinator.com/item?id=32265662 https://news.ycombinator.com/item?id=32265662 for information on predicting new folds.
- mupuff1234 4y ago> There are 3-4 other deep learning projects I think have had a much greater impact on my field. Don't leave us hanging... which projects?
- COGlory 4y ago1) Isonet - takes low SNR cryo-electron tomography images (that are extremely dose limited, so just incredibly blurry and frequently useless) and does two things: * Deconvolutes some image aberrations and "de-noises" the images * Compensates for missing wedge artifacts (missing wedge is the fact that the tomography isn't done -90° --> +90°, but usually instead -60° --> +60°, leaving a 30° wedge on the top and bottom of basically no information) which usually are some sort of directionality in image density. So if you have a sphere, the top and bottom will be extremely noisy and stretched up and down (in Z). https://www.biorxiv.org/content/10.1101/2021.07.17.452128v1 https://www.biorxiv.org/content/10.1101/2021.07.17.452128v1 2) Topaz, but topaz really counts as 2 or 3 different algorithms. Topaz has denoising of tomograms and of flat micrographs (i.e. images taken with a microscope, as opposed to 3D tomogram volumes). That denoising is helpful because it increases contrast (which is the fundamental problem in Cryo-EM for looking at biomolecules). Topaz also has a deep learning particle picker which is good at finding views of your protein that are under-represented, or otherwise missing, which again, normally results in artifacts when you build your 3D structure. https://emgweb.nysbc.org/topaz.html https://emgweb.nysbc.org/topaz.html 3) EMAN2 convolutional neural network for tomogram segmentation/Amira CNN for segmentation/flavor of the week CNN for tomogram segmentation. Basically, we can get a 3D volume of a cell or virus or whatever, but then they are noisy. To do anything worthwhile with it, even after denoising, we have to say "this is cell membrane, this is virus, this is nucleic acid" etc. CNNs have proven to be substantially better at doing this (provided you have an adequate "ground truth") than most users. https://www.ncbi.nlm.nih.gov/pmc/articles/PMC5623144/ https://www.ncbi.nlm.nih.gov/pmc/articles/PMC5623144/
- jakosz 4y agoNow we can start guessing what futures they are betting on: these, in which open-sourcing the whole thing commoditises critical complements. --- https://www.gwern.net/Complement https://www.gwern.net/Complement
- naves 4y agoJust imagine if the tech world puts all programatic advertising development on hold for a year and the collective brain power is channeled to science instead…
- inspirerhetoric 4y agoDoes anyone know what it would cost to download this whole dataset? Google Cloud Datasets only allow 1 TB/month for free to download, I believe
- inspirerhetoric 4y agoTo answer my own question it looks like for folks who don’t want to wait 21 months for 21 terabytes, that it might cost approximately 1600 USD to download the full approx 20TB dataset assuming egress costs of .08 USD per GB as mentioned here: https://cloud.google.com/storage/pricing#network-egress https://cloud.google.com/storage/pricing#network-egress It’s a pity it’s so expensive to download
- donut2d 4y agoA fun way I've been thinking about all this is what nanotech/nanobots are actually going to look like. Tiny little protein machines doing what they've been doing since the dawn of life. We now have a library of components, and as we start figuring out what they can do, and how to stack them, we can start building truly complex machinery for whatever crazy tasks we can imagine. The impact goes so far beyond drugs and treatments.
- sabujp 4y agoMany thanks to Deepmind for releasing predicted structures of all known protein monomers. What I'd like next is for Alphafold (or some other software) to be able to show us multimeric structures based on the single monomer/subunit predictions and protein-protein interactions (i.e. docking). For example the one I helped work on back in my structural biology days was the circadian clock protein KaiC : https://www.rcsb.org/structure/2GBL https://www.rcsb.org/structure/2GBL, that's the "complete" hexameric structure that shows how each of the subunits pack. The prediction for the single monomer that forms a hexamer is very close to the experimental https://alphafold.ebi.ac.uk/entry/Q79PF4 https://alphafold.ebi.ac.uk/entry/Q79PF4 and in fact shows the correct structure of AA residues 500 - 519 which we were never able to validate until 12 years later (https://www.rcsb.org/structure/5C5E https://www.rcsb.org/structure/5C5E) when we expressed those residues along with another protein called KaiA which we knew binds to the "top" CII terminal (AAs 497-519) of KaiC. If we would have had this data then, it would have allowed us to not only make better predictions about biological function and protein-protein interactions but would have helped better guide future experiments. What we can do with this data now is use methods such as cryo-em to see the "big picture", i.e. multi-subunit protein-protein interactions where we can plug in the Alphafold predicted structure into the cryo-em 3d density map and get predicted angstrom level views of what's happening without necessarily having to resort to slower methods such as NMR or x-ray crystallography to elucidate macromolecular interactions. A small gripe about the alphafold ebi website: it doesn't seem to show the known experimental structure, it just shows "Experimental structures: None available in PDB". For example the link to the alphafold structure above should link to the 2GBL, 1TF7, or any of the other kaic structures from organism PCC7942 at RCSB. This would require merging/mapping data from RCSB with EBI and at least doing some string matching, hopefully they're working on it!
- arolihas 4y agoYou might be interested in https://www.biorxiv.org/content/10.1101/2021.10.04.463034v2 https://www.biorxiv.org/content/10.1101/2021.10.04.463034v2
- roscoebeezie 4y agoI haven’t had a chance to look through some of the new predictions, but I know there were some issues with predicting the structure for membrane bound proteins previously. PDB hardly contains any. Does the new set of predictions contain a bunch of membrane bound protiens?
- djenendik 4y agoThe many body problem remains unsolved. So the question is, is this approach useful?
- epicquest 4y agoCome play biotech with us and let's figure out EVERYTHING and not just protein folding, yay! https://epicquest.bio https://epicquest.bio