6 ms·
Hey David, author of the linked repo here. I thought the paper was pretty neat and I'm a fan of the graph convolution work being done by both you and Patrick Ri
by frisco 10y ago
Hey David, author of the linked repo here. I thought the paper was pretty neat and I'm a fan of the graph convolution work being done by both you and Patrick Riley's and Vijay Pandes' groups. A few questions:
- In the paper you trained the property decoder separately but mention in passing that it could be a good idea to train the whole thing jointly. I haven't implemented the property decoder part yet and it seems like a pretty easy extension to just add that to the same loss function. Where is your thinking on independent vs joint training of the autoencoder/property prediction network now? Do you think non-obvious tricks will be required to get nice smoothness of the latent representation for things more complex than logP?
- Though learning directly to/from SMILES is weird, it's also kind of awesome. Do you think the future is an "inverse weave" module or something of that kind, or going further on learning more complex input representations?
- How do you feel about the synthetic accessibility score you used in the paper overall? Obviously it fell short on carbon ring sizes, but do you think we're close on this or we basically need to start over? What datasets would you think about using for "learning" synthetic accessibility as you imply in your post here?
Thanks! My vague intuition had been that this is a domain where I'd expect physics-based simulation should work better than machine learning but I have to say that the recent literature is changing my perspective.
- duvenaud 10y agoThank you very much, Max! To answer your questions: - As you say, jointly training the autoencoder for reconstruction and prediction accuracy would be easy to code, but it might be tricky to get the tradeoff right since we usually have relatively little labeled data. A massively multi-task approach probably makes sense here. I'm not sure how hard it will be to get smoothness w.r.t. the latent representation, but the logP result was mildly encouraging. - I agree that it's probably time to bring some of the recent augmented recurrent network architectures to bear on graph decoders. What makes is hard is that it requires differentiating through a cascading series of discrete choices, but people are making progress on this sort of problem. - I do think we need to completely start over on the synthetic accessibility. There are just too many ways for molecules to be weird to try to write an explicit function for it. In fact I'd go so far as to say that this is the 'missing half' of this method. - I view this method as a complement to physics-based simulation. A neural net is almost always going to be cheaper to evaluate than a physical simulation, so it's not a bad way to compile all the simulation results together.