4 ms·
Show HN: PyMOL-RS – Rust reimplementation of PyMOL with modern rendering
Well, it happened. After endless release candidates, we've finally made it to v0.1.0.
What's inside:
GPU-accelerated rendering with WebGPU shaders, shadows, and goodies like silhouette edges and a special soft-light mode
Core operations run up to 1000x faster than the original PyMOL. Surface generation that used to send you on a coffee run now finishes the moment you hit the button
Full PyMOL selection algebra support — 95+ keywords, boolean logic, distance/expansion operators, slash-macros
Distance, angle, and dihedral measurements, atom labels — everything you need for structural analysis
Python API — from pymol_rs import cmd and you're right at home
PDB, mmCIF, BinaryCIF, SDF/MOL, MOL2, XYZ, GRO — read and write, automatic format detection, transparent gzip decompression
Scenes, movies, ray tracing — all on the GPU
Kabsch superposition, CE alignment, RMSD, DSS, symmetry across all 230 space groups
Out of PyMOL's 798 original settings, some number of them actually work. Nobody knows exactly how many, but it's definitely in the hundreds. Plus we've added new settings that the original never had — like per-chain surface generation
13 independent crates — if you're writing Rust, you can use just the selection parser, the file readers, or the full GUI. No monolith
- dalke 7mo agoYou might mention in other forums, like the RDKit mailing list (though that's almost moribund). I looked at the SDF reader, since that's what I know best. I see a few things which look like they need revisiting. Line 75 has 'if name == "$$$$" {return self.parse_molecule();}' This isn't correct. This means the record name is "$$$$" (if you are RDKit), or it means the record is in the wrong format (if you are the CTFile specification, which explicitly prohibits that). Also, does Rust have tail recursion? If not, the recursive nature of the code makes me think parsing a file containing 1 million lines of the form "$$$$\n" would likely blow the stack. In principle the version number test for V2000 or V3000 should look at the specific column numbers, and not presence somewhere in the line. Someone like me might place a "V3000" in the obsolete fields, with a "V2000" in the correct vvvvvv field. ;) The "Skip to end of molecule" code will break on real-world datasets. One classic problem is a company which used "$", "$$", "$$$" and "$$$$" to indicate cost, stored as tag data like: > <price> $$$$ $$$$ where the first "$$$$" is part of the data item, and the second "$$$$" is the end of the SD record. This ended up causing a problem when an SDF reader somewhere in their system didn't parse data items correctly. (Another common failure in data item parsing is to ignore the requirement for a newline after the data item.) I talk about "$$$$" more at http://www.dalkescientific.com/writings/diary/archive/2020/09/18/handling_the_sdf_record_delimiter.html http://www.dalkescientific.com/writings/diary/archive/2020/0... . Then there's the "S SKP" field, which you'll almost certainly never see in real life! I've only seen it used in a published example of a JICST extended MOLfile. See http://www.dalkescientific.com/writings/diary/archive/2020/09/21/molfile_s__skp.html http://www.dalkescientific.com/writings/diary/archive/2020/0... Please don't let these comments get you down! These details are hard to get, and not obvious. It took me years to learn the rare corner cases. I also haven't done molviz since the 1990s, or used PyMol (I was VMD person), so can't say anything about the overall project. We started with GL, and had to port to OpenGL. :) PS. A bit of history for you. PyMol's and VMD's selection syntax look similar because both drew on the syntax in Axel Brunger's X-PLOR. Warren DeLano came out of Brunger's lab, and VMD was from Schulten's group, which were X-PLOR users. (Schulten was Brunger's PhD advisor.)
- dalke 7mo agoI looked at the PDB parser. Will you be adding support for using duplicate CONECT records to store bond type information? That's a RasMol extension that PyMol supports. You'll also need to support the pdb_conect_nodup option in the writer. I see you interpret atom/hetatm serial numbers as an integer. Will you be using the base-36 or hybrid-36 variants (see https://cci.lbl.gov/hybrid_36/ https://cci.lbl.gov/hybrid_36/) which is a common way to handle more than 100,000 atoms? Again, these are corner cases which come with experience. I've no expectation that a new program would handle them. I want you to know about them since they will be issues if you expect long-term uptake.
- dalke 7mo agoThe README says "PyMOL-RS is a clean-room rewrite" but when I look at ./pymol-mol/src/dss.rs I see things like: //! - PyMOL's layer3/Selector.cpp - SelectorAssignSS function //! - PyMOL's layer2/ObjectMolecule2.cpp - ObjectMoleculeGetCheckHBond function //! - PyMOL's layer1/SettingInfo.h - Default angle thresholds and "matching PyMOL's cSS* flags from Selector.cpp" While the Rust code is cleaned up and easier to read, I can see that it preserves similar data flow, uses similar variable names, and of course identical constants. For example, this is PyMol layer3/Selector.cpp: /* look for antiparallel beta sheet ladders (single or double) ... */ if((r + 1)->real && (r + 2)->real) { for(b = 0; b < r->n_acc; b++) { /* iterate through acceptors */ r2 = (res + r->acc[b]) - 2; /* go back 2 */ if(r2->real) { for(c = 0; c < r2->n_acc; c++) { if(r2->acc[c] == a + 2) { /* found a ladder */ (r)->flags |= cSSAntiStrandSingleHB; (r + 1)->flags |= cSSAntiStrandSkip; (r + 2)->flags |= cSSAntiStrandSingleHB; (r2)->flags |= cSSAntiStrandSingleHB; (r2 + 1)->flags |= cSSAntiStrandSkip; (r2 + 2)->flags |= cSSAntiStrandSingleHB; /* printf("anti ladder %s %s to %s %s\n", r->obj->AtomInfo[I->Table[r->ca].atom].resi, r->obj->AtomInfo[I->Table[(r+2)->ca].atom].resi, r2->obj->AtomInfo[I->Table[r2->ca].atom].resi, r2->obj->AtomInfo[I->Table[(r2+2)->ca].atom].resi); */ } } } } and this is pymol-rs's pymol-mol/src/dss.rs // Antiparallel ladder: i accepts j, (j-2) accepts (i+2) if a + 2 < n_res && res[a + 1].real && res[a + 2].real { for &acc_j in &acc_list { if acc_j < 2 || !res[acc_j].real { continue; } let j_minus_2 = acc_j - 2; if !res[j_minus_2].real { continue; } let acc_jm2_list: Vec<usize> = res[j_minus_2].acc.clone(); for &acc_k in &acc_jm2_list { if acc_k == a + 2 { res[a].flags |= SsFlags::ANTI_STRAND_SINGLE_HB; res[a + 1].flags |= SsFlags::ANTI_STRAND_SKIP; res[a + 2].flags |= SsFlags::ANTI_STRAND_SINGLE_HB; res[j_minus_2].flags |= SsFlags::ANTI_STRAND_SINGLE_HB; if acc_j >= j_minus_2 + 2 { res[j_minus_2 + 1].flags |= SsFlags::ANTI_STRAND_SKIP; } res[acc_j].flags |= SsFlags::ANTI_STRAND_SINGLE_HB; } } } } That's close enough that I really think you should include the PyMol license info, before Schrödinger's lawyers notice.