10 ms·
Keras-based molecular autoencoder
- jostmey 10y agoThis is the future of computational chemistry. The field of "Molecular Dynamics" could easily be swept aside by machine learning. It is why I switched from running "Molecular Dynamic" simulations as a graduate student to modelling genomic data with TensorFlow as a Postdoc.
- dkural 10y agoWhat kind of genomic data out of curiosity?
- dnautics 10y agoIs that actually working? I tried rolling my own neural nets and running genomic data through them and instantly realized the problem was low N, high autocorrelation in the data. Bayesian forests seem like a better choice, but that's me.
- deleted 10y ago[deleted]
- alextheparrot 10y agoThere have been some recent papers in subsets of genomics with a bit more data (Transcription factor binding, for example [1]). You're correct though, definitely depends on your problem. [1] https://www.ncbi.nlm.nih.gov/pmc/articles/PMC4908339/pdf/btw255.pdf https://www.ncbi.nlm.nih.gov/pmc/articles/PMC4908339/pdf/btw...
- cing 10y agoHow do you suppose the questions you've answered working on ion channels would be addressed with machine learning? I get the impression by "swept aside" that you feel the field of molecular biophysics is not meaningful in the study of health and medicine.
- jostmey 10y agoThe neural network community uses proper "training"/"validation"/"test" datasets to asses model performance, and has developed algorithms to fit complex models to large amounts of data using relatively little computing power. I think it is possible to build completely new and accurate models of biophysical systems with neural networks with relative ease. Using models like variational autoencoders it might be possible to draw truly independent samples in a single step, instead of relying lots of MD steps to find novel conformations. I could go on, but I have to get to bed now :-)
- cing 10y agoI like the idea, but I wouldn't call it "relative ease". Are we talking about making better force fields, or training a recurrent neural network for generative dynamics that obeys the laws of physics (can be used to calculate ensemble averages) for arbitrary length proteins/folds? How do you construct a VAE for protein dynamics when you only have a single structure? There's a serious lack of data that prevents these things.
- momeara 10y agoMolecular dynamics simulations can be used to answer a range of structural biology questions, but abstractly many of them can be phrased as evaluating the difference in free energy between different conformational states. In molecular dynamics this is done by thermodynamic integrating the energy of over the state space volume for each of the conformational states. An alternative approach is to directly map conformational states to their free energy. This leads to a problem of searching for candidate conformational states (e.g. the folded state, transition states etc.) and scoring them. Usually for a given computational budget there is a trade off between better conformational sampling or higher accuracy energy scoring. Historically, searching and scoring methods have been designed separately. For example [1] improves sampling while [2] improves energetics. This is done because they historically involved different aspects of the simulation and each is lot of work. But searching and sampling are not really separable, in that the deeper one samples the more challenging the task of the scoring function becomes--discriminating stable from unstable conformations. Another application that can be thought of as searching and scoring is the game of GO. My impression is that one of the major breakthroughs with AlphaGo is that they were able to integrate models for searching and scoring together and learn the models simultaneously. It would be awesome if similar architectures could be applied to molecular modeling. A remaining challenge in applying GO models to molecular biology is that while the representation and scoring rules for GO are fixed and quite easy, the ground truth for molecular simulations comes from heterogenous experimental data (X-ray crystal structures, small molecule activities, directed evolution antibody screens etc.) and higher levels of theory QM simulations, which have their own challenges. However, I think the principles carry over--complicated scoring functions (e.g. free energy) over large state spaces (e.g. protein conformation space or chemical space) can be learned by combining models for searching and scoring. I think deep learning is poised to tackle these problems. [1] (Conway, et al., 2013, DOI: 10.1002/pro.2389) Relaxation of backbone bond geometry improves protein energy landscape modeling [2] (Park, 2016, PMID: 27766851) Simultaneous optimization of biomolecular energy function on features from small molecules and macromolecules.
- zump 10y agoSo DE Shaw Research is barking up the wrong tree?
- rgbombarelli 10y agoWho says you cannot do deep learning and molecular dynamics? There are a couple of recent examples of neural networks to predicted energies and gradients allowing faster computation and larger time steps in MD simulations. https://arxiv.org/abs/1609.08259 https://arxiv.org/abs/1609.08259
- Xcelerate 10y agoI agree. I did molecular dynamics simulations for my PhD. The last year I spent some time exploring feature representations that capture atomic environments up to their symmetry invariants (rotations, reflections, permutations of identical atoms, etc.) The use of machine learning for MD and quantum chemistry has exploded in the last year and a half. I'm a bit sad that I just finished my degree right as this kind of work is heating up.
- pizza 10y agoh-suppressed graphs? VESPR graph stuff? compressive sampling / measuring of typical states? ligand complices? ligand transport? membrane dynamics? predictive toxicology? cheaper hartree-fock approximations? quark shit??? time to daydream
- pizza 10y agoPlease, feed me more specifics while I salivate. In as well-defined a problem/mission statement/objective as possible, what are you doing (and with what / how), and where do you intend to reach?
- dkural 10y agoThe author of the github repo doesn't seem to be an author on the ArXiv pre-print. Anyone knows why?
- frisco 10y agoI'm not part of the group that wrote the paper. I'm just a guy on the internet.
- deleted 10y ago[deleted]
- rgbombarelli 10y agoOne of the authors here. Mostly he was just faster than us! Most co-authors are busy right now starting jobs at new places (plus trying to get published at an old school journal for the comfort of chemists out there). Max has done a very good job writing a neat implementation of our autoencoder. I encourage everyone to go invent some new molecules!
- duvenaud 10y agoOne of the authors of the original paper https://arxiv.org/abs/1610.02415 https://arxiv.org/abs/1610.02415 here. From a machine learning point of view, we simply glued together two techniques: text autoencoders and Bayesian optimization. That is, we trained an autoencoder to transform a text representation of chemicals (SMILES) to and from continuous vectors. Then “chemical design” is just maximizing a function of a continuous variable, something that we already know a lot about. We also showed off some of the nice things that one can do with continuous latent representations, such as interpolation. This had already been done for images by many people, and for text in https://arxiv.org/abs/1511.06349 https://arxiv.org/abs/1511.06349 Of course, our paper is just a proof of concept. For instance, instead of encoding to and from SMILES, it would be much better to encode to and from graphs directly. We know how to encode graphs into vectors, but I don’t know of a good way to decode a vector into a graph. Another open problem is that it’s hard to know what to optimize. Our initial experiments optimizing for specific chemical properties produced suggested molecules with crazy structures, such as giant rings. Human chemists have a great intuition for what is easily synthesizable or stable, and it’s hard to articulate all the properties we want the molecule to have programmatically. Alternatively, we could enforce the optimizer to only look at molecules similar to ones we’ve already seen, but this is unsatisfying too - after all, the best result of exploration is when you find something unlike what you’ve seen before.
- brilee 10y agoSMILES strings are not necessarily unique; each molecule can be encoded multiple ways as a SMILES string. As a sanity check, you could pass multiple representations of a molecule into the autoencoder to see if they generate the same continuous representation. Otherwise, you may end up essentially training to the peculiarities of whatever algorithm generated your SMILES strings.
- deleted 10y ago[deleted]
- rgbombarelli 10y agoYou are right about SMILES. We trained on the canonical SMILES output by the RDKit, so as far as our AE is concerned, there is only one way to write SMILES for a given molecule (http://www.rdkit.org/docs/GettingStartedInPython.html#writing-molecules http://www.rdkit.org/docs/GettingStartedInPython.html#writin...). Of course, the choice of what the canonical SMILES is is somewhat arbitrary.
- pizza 10y agoIt might be interesting if there were some routine that could 'reverse-decode' drug-like properties, i.e. producing Lipinski's rule of 5 from nothing but "drug-like" and "not-drug-like" labeled training set.
- rgbombarelli 10y agoWe did something like that for a more continuous property, like logP, which is easy to predict with cheminformatics. We are working on metrics that reflect drug-likeness better.
- pizza 10y agoInteresting. If I'm not mistaken, logP is a type of property where it is easy to look up constituents' "contribution" for the purpose of predicting logP with substitutions/alterations of the example molecule.
- rgbombarelli 10y agoExactly, we have a good understanding that logP is additive (as a matter of fact, one predicts it using group contributions). The AE noticed this and started adding halogens to already high logP molecules. The interesting part is that it stops before going totally crazy and per-halogenating the molecule. This is probably because it has an intuition about how molecules look like and hasn't really seen that kind of substitution pattern.
- Houshalter 10y agoWait, what is logP? I assumed you meant log probability, since that is the standard error metric for text prediction tasks like this.
- rgbombarelli 10y agoIt's the water-octanol partition coefficient, a basic molecular descriptor of the physiological distribution of a drug. It is one of the multiple targets one aims to optimize in a novel drug-like compound. https://en.wikipedia.org/wiki/Partition_coefficient https://en.wikipedia.org/wiki/Partition_coefficient
- michaelmwangi 10y agoI did Computer Aided Drug Design for my undergrad. I wish I knew this
- ecesena 10y agoThe chart in the readme makes the text completely unreadable on mobile. I'd suggest putting it on top or bottom of the text.