4 ms·
Thank you for the feedback dalke! Yes, I store those pre-computed values as a header in front of the standard RDKit binary molecule, not the SMILES string. I
by rtvp- 2y ago
Thank you for the feedback dalke!
Yes, I store those pre-computed values as a header in front of the standard RDKit binary molecule, not the SMILES string.
I first convert the SMILES string to a RDKit molecule. Then I get those values (number of atoms, number of rings, etc.) from the RDKit molecule object. Then I serialize those values as a header in front of the rest of the binary molecule object.
Like so: Pre-computed values header (20 bytes) + binary RDKit molecule object (n bytes)
I wanted to store the binary molecule object in my database so that other functionality like substructure searches, or converting to a different molecule format could use that molecule object, like the RDKit Postgres extension `mol` type.
But, a separate, dedicated column containing a hash just for exact searches might be even faster with duckdb's string capabilities.
Although a SMILES string may be 60 bytes on average, the binary molecule from RDkit may be bigger. I took a look at chembl, and although it may contain some strange structures, I found that the minimum byte length of a pickled/binary molecule via the RDKit Postgres extension is 65 bytes, with the max being 15,762 bytes [1].
But, I did not calculate the average size of the binary molecule, and your comment helped me to realize that. I should take a look at that. If the average is much bigger, 20 bytes is, perhaps, not too much more overhead.
I am assuming that getting these header/prefix values (number of atoms, rings, etc) from the RDKit molecule is not too expensive. Since the more expensive part of parsing the input (SMILES) and converting it to the RDKit molecule is being done anyways (in order to get the molecule object), I'm assuming I can get those header values without much additional computational cost.
[1]: https://bodowd.github.io/2024/06/18/chembl-column-properties https://bodowd.github.io/2024/06/18/chembl-column-properties
- rtvp- 2y agoFollowing up: The average size of the binary molecule in chembl33 is 454.9 bytes
- dalke 2y agoMy comment about SMILES strings size was meant to highlight how the 20 bytes you use is quite a lot of data. Yes, the binary RDMol is larger than the SMILES - it can also store more data, including coordinates and extended stereochemistry, which cannot be expressed as a SMILES. If you are just worried about structure equivalence, and you consider the canonical SMILES to be the best way to check for that (which your snippet does), and you don't mind the additional space and creation time, then put the canonical SMILES into the database along side the binary, and index on the string. If you don't want any extra space or creation time overhead, then you end up with the design you've seen, where it tries hard to work with values easily computed from the molecule. It seems like you're looking for an intermediate, with fixed space. If you just want fast identity tests, and are willing to take a small fixed space, and aren't so concerned about database creation time, then you can generate the SMILES and hash them to a fixed width. Even 64 bits should be good enough to filter all but a handful of mismatches. You can also use one of the molecule hashes, which are faster to compute, but will have a higher false-positive rate. (But still less than what your current scheme uses.) It looks now like you want something which can also be used as a substructure screen. That's in general what substructure screens are designed for. There's a long and dusty history of developing substructure screens, which is mostly forgotten these days now that compute power and memory are so cheap. Consider that the MACCS keys were developed in the 1970s as a substructure screen, with 166 bits, or just slightly higher than the 160 bits you use in your prefix. Coming up with a good 128-bit or even 64-bit screen is a topic I've wanted to investigate for a long time. The furthest I got was http://www.dalkescientific.com/writings/diary/archive/2012/06/11/optimizing_substructure_keys.html http://www.dalkescientific.com/writings/diary/archive/2012/0... from about 12 years ago. I mention a 54 bit fingerprint. It's available at https://www.mail-archive.com/rdkit-discuss@lists.sourceforge.net/msg02078.html https://www.mail-archive.com/rdkit-discuss@lists.sourceforge... . I'm told it's pretty decent, but I haven't put it to practice myself. If you do more in this area, I've managed to gather some queries (identity, substructure, and similarity), in my Structure Query Collection, at https://hg.sr.ht/~dalke/sqc https://hg.sr.ht/~dalke/sqc .
- rtvp- 2y agoI think I'm getting it now... If I'm going to put a prefix, I could use something that gets more elimination power out of those bytes than what I'm currently using. Do I have the right understanding? I hadn't even thought of using the prefix for substructure queries, but that is an interesting idea. Thank you for the links and the new ideas! You've given me a few new ideas to explore
- dalke 2y agoYes, you have the right understanding. I plan to be at the RDKit User Group meeting in Germany this fall. You might think of participating, even remotely.
- rtvp- 2y agoI've been reading your post and looking at the 54 bit fingerprint you mentioned. I was wondering if you could help clear up any confusion on my end; I am not sure if I am implementing and applying it correctly. Regarding the 54 bit fingerprint, if I do a substructure match between a target molecule and the query molecule: "O", and I find there are 3 matches, then if the list of the patterns were: "0 2 times O largest 55458" and then later "16 3 times O largest 653" "17 1 times O largest 466" does that mean I should set bit 0, 16, and 17 to 1? If I've understood correctly up to this point, I would have some additional questions, but if I am totally off, my additional questions would not make sense. Edit: I had a mistake in my logic with the second question. Removed it