5 ms·
I know nothing of this industry but, if I had to guess , a few companies hold large amounts of data . They will be able to train on proprietary data, develop th
by pylua 3y ago
I know nothing of this industry but, if I had to guess , a few companies hold large amounts of data . They will be able to train on proprietary data, develop the proper trial and error automation then really automate out a lot of chemists, much like what could happen to the art scene.
- version_five 3y agoI know nothing of the industry either, but I know a fair amount about corporate data landscapes generally, and if I had to bet, I'd say nobody is sitting on a treasure trove of data that could readily be used for training. In any event, the kind of self or loosely supervised training we're seeing making an impact in the current generation of "AI" is very different anyway from what I picture here as some kind of supervised task. There will need to me some "large chemistry models are zero shot learners" breakthrough with an appropriate pretext task to get to something parallel.
- bloaf 3y agoThey do. Internally, the gatekeepers of this data are mostly old-school chemists whose ideas of AI are stuck in the "maybe AI can do some property predictions better than our first principles FORTRAN 77 models, but I doubt it" mindset.
- jleyank 3y agoNah, first principal models (ie, QM) are now in Hi-Performance Fortran, C, C++ and/or Python: https://psicode.org/ https://psicode.org/
- bloaf 3y agoNo. I'm talking about things like polymer rheological or film properties, or industrial scale reactor performance. That's not to mention one-offs like cloud point predictors, or specialized thermo packages. See, for example, https://www.technipenergies.com/sites/energies/files/2021-11/Brochure_PT_Ethylene_SPYRO_WEB.pdf https://www.technipenergies.com/sites/energies/files/2021-11...
- jleyank 3y agoSorry, my chemical experience is qm and biotech/pharma. Molecular not large-scale…
- shoubidouwah 3y agoThat is just false. The gatekeepers of this data see a new algorithm, get excited, try it, test it against the aforementioned model, and then keep the old model because it's just better /s. Real talk now, I'm an "AI person" from big pharma, and we're quite up to date. topological neural networwks, diffusion models, QM neural potentials, large scale meta-learning, systematic active learning, we're doing it all. We also know that most of the time, a small bayesian GLM or random forest on run-of-the-mill descriptors actually works very well and fails predictably, which is important. Data in pharma in particular tends to be sparse and shallow: a hundred datapoints clustered tightly in chemical space because that's what the process generates. Sharing the data can lead competitors to the precious IP you're protecting, which is why we also invest a lot in blind federated learning etc. Anyway, the going's tough, but everyone is doing their best. nobody's dismissing AI at all... we're just more aware of the domain-specific pitfalls.
- ta988 3y agoThey are already doing that. But you still need chemists because they can make the things and check them and find alternative routes and produce analogues etc There are quite a few tools using DL models that work extremely well to devise synthetic pathways for compounds. But you still need someone to make them. A lot of the easy to automate chemistry (combinatorial chemistry) didn't really give good results compared to the amount of money it gobbled. And these days in chemistry we are seeing a lot of what is happening in the electronics world as well. With a set of companies producing different materials and executing different parts of the process for another one (think Apple cpus with the whole chain from the Swiss EUV mirror makers, the wafers producers, the machines producers, TSMC that orchestrate the whole thing, etc). Pharma companies are externalizing a these days for the chemistry, the analysis etc.