4 ms·
I wonder how effective AI will be at enabling us to navigate chemical space for certain desired compounds. Knowing nothing about the problem, is it something ak
by adt2bt 5y ago
I wonder how effective AI will be at enabling us to navigate chemical space for certain desired compounds. Knowing nothing about the problem, is it something akin to the protein folding challenge that AlphaFold[0] recently did well at?
Side note: I love Derek Lowe's writings. I don't know what it is, but every time I see a chemistry related link bubble up in HN, I have a gut feeling it was written by him. And I'm usually impressed. His Things I Won't Work With[1] series is amazingly well written.
[0] https://deepmind.com/blog/article/alphafold-a-solution-to-a-50-year-old-grand-challenge-in-biology https://deepmind.com/blog/article/alphafold-a-solution-to-a-...
[1] https://blogs.sciencemag.org/pipeline/search/Things+I+Wont+Work+With https://blogs.sciencemag.org/pipeline/search/Things+I+Wont+W...
- timr 5y agoDerek is right about the vastness of chemical space, but I go back and forth on the claim (frequently made by those in drug discovery) that AI cannot possibly extrapolate to spaces of this size, for at least three reasons: * Image space and text space are also vast, and yet we've had good success applying AI in these areas. I have yet to see a convincing argument that these spaces aren't equally large. * It's a bit of a red herring: actual drug discovery programs are not exploring "all of chemical space". They're usually focused on "lead series" of much more constrained molecules. * There are actually context-independent signals that can be used to generalize AI methods. The generalization is far from perfect, but it's not like every one of those 10^60 molecules are entirely different from every other molecule in the set. There are clusters and patterns and trends that can be exploited for gain -- this is what makes "medicinal chemistry" an academic field, and not merely an exercise in fortune-telling. Personally, I think the bigger problem applying AI and ML to drug discovery is less the "vastness of chemical space" (a proposition that makes med-chemists feel secure about their jobs), and more that the datasets in drug discovery suck. There's tons of siloing of data, none of it is consistent, and you can't even depend that two assays for the same target, measured in the same lab, years apart, will yield consistent data. It's a total mess.
- dnautics 5y agoSo text space is trivially vectorizable, at the character level and even for difficult languages like Russian chunk-vectorisable with some care. How do you encode the difference between houamine A and atrop-houamine A, while keeping the similarities, without resorting to empirical measurements and classification, which could yield reasonable vectors, but will take 2-5 years of a highly trained grad student's labor to obtain and put into the training corpus
- deleted 5y ago[deleted]
- timr 5y ago> How do you encode the difference between houamine A and atrop-houamine A, There are now lots of ways of encoding molecules. So many, in fact, that it's not really worth debating the merits of any particular method. ECFP fingerprints shoved into a fully connected NN work surprisingly well for a large class of problems. Molecular graph convolutions (of which there are now many flavors) also work well. The field is to the point where people are doing ensembles of different encodings, and seeing what works for any particular problem. > without resorting to empirical measurements and classification, which could yield reasonable vectors, but will take 2-5 years of a highly trained grad student's labor to obtain and put into the training corpus Well, you're sort of touching on my last paragraph with this. The classifier, featurization, etc., usually matters less (a lot less?) than the quality of the assay data. So I agree in that respect.
- dnautics 5y ago> Molecular graph convolutions "also work well" What is the metric for "working well" here? My point is a graph convolution will have a hard time distinguishing the haouamine atropisomers, because they have the same graph, but very different activities. This sort of weird shit is the norm. Like some molecule that isn't the actual active form but needs cyp450 to be activated. But it won't be if you add a methyl group that reduces it toxicity, because that methyl blocks cyp (not your target even). But only in college student pre-phase I volunteers, who happen to be predominantly white male college students looking for beer money In the end your corpus is going to have to fall back to experiment, which is just so much slower in chemistry than the trivially parallel "scrape every Wikipedia article, twitter tweet, and reddit comment" that you have access to for text corpora; it's more like trying to use ml to decode Chinese, if we only had 3000 Chinese documents to work with.
- krab 5y agoI think that the biggest issue isn't the chemical space but the complexity of biological systems. It's hard to tell what the molecule will do. We just don't have a good enough simulator. AlphaFold is definitely a helpful step but more are needed in the same direction. (In 2010 - 2012 I worked in a laboratory that did small compounds screening and I was building some tools to explore the chemical space)
- derefr 5y agoThat’s specifically about chemicals presumed for physiological use/purpose, though, no? Let’s take biology out of the equation. Can we predict chemical structures will give rise to interesting physical properties when you have a large amount of the compound around? For some examples, could we ever potentially have a formula or model that would allow us to predict: • whether any other molecules will possess the “infectious amalgamation” effect that gallium has on other metals it touches? • what chemicals would create a “non-stick” surface like Teflon does? • what types of chemicals would be both strongly colored in the visible spectrum and also chemically stable, i.e. what chemicals would make especially good pigments? • whether something will be a room-temperature superconductor? • whether something would make a better non-reactive glassware than borosilicate glass? • what types of chemicals will be especially good at being electrolytes? And I’d especially like to know, whether there’s some property that unites all “chemicals with interesting physical properties” like this, i.e. whether it’d be possible to have a model that tells you that a chemical is likely to have unusual physical properties in some way, despite not being able to predict the specific effect it’ll have.
- deleted 5y ago[deleted]