8 ms·
Biological Function Emerges from Unsupervised Learning on 250M Protein Sequences
- lucidrains 7y agoLanguage, music, and now amino acid sequences. Attention is all you need.
- mfatica 7y agoI would say you also need a fair bit of data too...
- bearmcbearsly 7y agoWell, yes. But I think lucidrains was referring to: https://arxiv.org/abs/1706.03762 https://arxiv.org/abs/1706.03762
- return1 7y agoand transformers
- nl 7y agoThe Attention Is All You Need paper is where Transforms were introduced: We propose a new simple network architecture, the Transformer, based solely on attention mechanisms, dispensing with recurrence and convolutions entirely.
- return1 7y agoYup, and the state of the art BERT and gpt-2 are both based on transformers.
- superfx 7y agoSee also: https://www.biorxiv.org/content/10.1101/589333v1 https://www.biorxiv.org/content/10.1101/589333v1
- tepal 7y agoThis blog post seems to anticipate this happening: https://moalquraishi.wordpress.com/2019/04/01/the-future-of-protein-science-will-not-be-supervised/ https://moalquraishi.wordpress.com/2019/04/01/the-future-of-...
- dnautics 7y ago> It does a surprisingly good job of predicting protein function across a diverse set of tasks, including ones structural in nature, like the induction of a single neuron that is able, with some degree of accuracy (ρ = 0.33) to distinguish between α helices and β strands (I suspect the network as a whole is far more performant at this task than the single neuron we’ve identified, but we didn’t push this aspect of the analysis as the problem is well tackled using specialized approaches.) I hate to be that guy, but distinguishing between alpha helices and beta strands is not really that hard. It's a good start though. I would propose the following test: Let's see if we can use the activations from the neurons to predict the luminosity of a 'base' GFP molecule (under a fixed set of experimental conditions). Train the set on 10,000 mutations (this could maybe be done in very high throughput by tethering the XNA to a bead, synthesizing, and then measuring the beads one by one), and see if can extrapolate the effects of 10k more, or heck, just by doing it brute-forcedly, we've got high throughput robots, right?
- jostmey 7y agoAnd predicting protein function is not that hard either. The ground truth labels are often determined by sequence alignment similarity, not by experiment. So the results are far from profound
- jerven 7y agoDoing it right is quite hard. Doing it usefully is even harder [1]. Getting a good training set without to many biases is the really hard part. Generating a ground truth that is actually a truth is very expensive. I have to read the paper carefully again. But for the contact point prediction I think the training set will cover most of the data used in the validation. Due to they way PDB "sequences" are distributed over UniParc as well as how PDB 3D structures are generated experimentally. i.e. there are 120,000 pdb related sequences in UniParc, but they cover 45,000 ones in UniProtKB. Because PDB derived sequences are rarely full length, often mutated and highly duplicative in coverage. [1] predicting the root GO terms will give you and insane TP/FP rate but is completely useless.
- gigantum 7y agoLike some of the other ML/AI posts that made it to the top page today, this research too does not give any clear way to reproduce the results. I looked through the pre-print page as well as the full manuscript itself. Without reproducibility and transparency in the code and data, the impact of this research is ultimately limited. No one else can recreate, iterate, and refine the results, nor can anyone rigorously evaluate the methodology used (besides giving a guess after reading a manuscript). The year is 2019, many are finally realizing it's time to back up your results with code, data, and some kind of specification of the computing environment you're using. Science is about sharing your work for others in the research community to build upon. Leave the manuscript for the pretty formality.
- Havoc 7y ago>any clear way to reproduce the results. Given that it's evolved I'd imagine this is a given? Or more accurately you could probably duplicate some kind of emergent behaviour but it would be different given different randomized parameters
- lysium 7y agoUsually you use an RNG for which you can publish the seed. So, although it’s random, you can reproduce the results.
- tastroder 7y agoGlancing through the paper it seems like they use the recent Transformer model. Does whatever underlying stack they use expose something to share RNG seeds and the exact hardware optimizations your environment applies during training? Otherwise "publishing the seed" sounds nice but might not be as trivial as the phrase suggests.
- arthur_pryor 7y agoreproducibility should be something that's baked into an experiment's design. so, if their experiment was designed such that reproduction is inherently difficult, they should have designed it in a better way, and they should've used a toolset that wouldn't run into that problem. a non-reproducible experiment isn't necessarily completely without value, but it's a thing that everyone should look askance at till it proves its worth. (apologies if my comments don't apply to this experiment and if it is reproducible -- i didn't have time to read through the OP, but i thought this reply was still a worthwhile response to its specific parent comment)
- cellular 7y agoI find these emergent behaviours fascinating: https://youtu.be/gaFKqOBTj9w https://youtu.be/gaFKqOBTj9w
- cellular 7y agoIt would be neat to run this with a few more rules, on a larger world, and for a longer time to see what emerges!
- jakeogh 7y agoFPGA do interesting things when allowed to exploit sidechannels/analog effects: http://citeseerx.ist.psu.edu/viewdoc/download?doi=10.1.1.50.9691&rep=rep1&type=pdf http://citeseerx.ist.psu.edu/viewdoc/download?doi=10.1.1.50....
- ArtWomb 7y agoFergus Lab at NYU. I believe he's across the hall from Yann LaCunn as well ;) Still a long way from a Theory of Biogenesis. But a good next step is using a differentiable model to predict novel proteins which have no analogue in Nature. Much like Materials Genome researchers searching for stable phases of matter! "Training ever bigger convnets and LSTMs on ever bigger datasets gets us closer to Strong AI -- in the same sense that building taller towers gets us closer to the moon." --François Chollet
- eganist 7y ago> I believe he's across the hall from Yann LaCunn as well ;) I'm having a hard time processing what the wink might possibly mean in this context. No sarcasm intended.
- mkolodny 7y agoI'd guess that wink is hinting that Yann LeCun might've had something to do with this research. Whether that's true or not, I have no idea. (Yann LeCun is a Turing award winner for his work in deep learning)
- slashcom 7y agoYann LeCun did not, otherwise he’d be a coauthor. As it is, this was a collaboration between NYU and Facebook AI Research, with multiple authors working at both institutions.
- drb91 7y agoMy understanding is that academic authorship credit is political: authors don’t always contribute, and contributors don’t always get credit. Is this not the case?
- nl 7y agoNot really. Usually the politics goes the other way - people getting an author slot because they are the head of the department or something.
- shpongled 7y agoThis is cool, but would be significantly cooler if they did some kind of biological follow up. Perhaps getting their model to output an "ideal" sequence for a desired enzymatic function and then swapping that domain into an existing protein lacking the new function.
- inciampati 7y agoBingo. That would be really interesting. And useful. There are probably already enzymes in this data set that have measurements of their behavior. Could this modelling approach be coaxed to find the one with the highest processivity? Or do we need more labeled data?
- shpongled 7y agoI'm sure they have a bunch of enzymes in their dataset for which kinetic measurements have been published. Another interesting follow up study would attempting to improve kinetic behavior. They could, for instance, analyze some of the catalytically perfect enzymes out there (TIM, SOD, catalase, etc) and see if the model could project improvements onto existing orthogonal protein classes.
- jerven 7y agoNot in a structured way that is easily useable. Swiss-Prot has most of this data but it is not quite normalized in units. If you did this annoying work I would like to talk to you so we can plug it into Swiss-Prot.
- obviuosly 7y ago> The resulting model maps raw sequences to representations of biological properties without labels or prior domain knowledge. A couple of questions: 1. What are those representations? 2. Also what is "biological function"? 3. What kind of information does the learned representation extract that is not already in the "biological properties" it is trained to map to?
- andbberger 7y agoI find this paper to be so steeped in hype and dogma so as to be nearly incomprehensible. Which is a shame, because it's a reasonable approach. I just wish they just frickin described what they did instead of spending the whole paper monologuing and showcasing unconvincing experiments. No need to justify what you're doing, just do it.
- a_bonobo 7y agoHere's a very cool GitHub repository which uses unsupervised learning (ULMFiT) in the genomics space: https://github.com/kheyer/Genomic-ULMFiT https://github.com/kheyer/Genomic-ULMFiT Very impressive accuracies on hard tasks, and it's open source!