4 ms·
Muyuan Chen is one of (maybe the?) primary developer of the sub-tomogram averaging portion of the EMAN2 software package (linked below in another comment). Typi
by COGlory 3y ago
Muyuan Chen is one of (maybe the?) primary developer of the sub-tomogram averaging portion of the EMAN2 software package (linked below in another comment). Typically what you do is take a 3D tomogram (think like a scan) using a microscope, but it's extremely noisy. Then you go through and extract all the particles that are identical, but in different orientations, in the tomogram. So if the same protein is there multiple times, you can align them to each other, and average them together to increase the signal. Then you clone back in the higher signal averaged volume at the position and orientation that you found them in originally.
The one-line command to go from EMAN2 coordinates to Unreal Engine 5 is kind of crazy.
As usual on these (rare) threads, I'm happy to answer any questions about structural biology or cryo-EM.
- frodo8sam 3y agoSounds pretty cool, how does EMAN2 deal with dynamic structures? I assume you'll get garbage if sufficiently different conformations get averaged together. Is there some kind of clustering to find similar conformations as is sometimes done in cryo-EM?
- COGlory 3y agoYes, at multiple levels. You can do a heterogeneous refinement that takes the structures and solves for n number of averages, trying to use something like PCA to maximize the distances between the averages. Particles get sorted into the average they contribute most constructively to. That works well for compositional heterogeneity or large conformational differences. For minor differences, there's something called the guassian mixture model (lots of software packages have similar, but GMM is EMAN2's version). https://blake.bcm.edu/emanwiki/EMAN2/e2gmm https://blake.bcm.edu/emanwiki/EMAN2/e2gmm What you can get out of the other end is a volume series, that shows reconstructed 3D volumes along various axes of conformational variability. This quickly turns into a multi-dimensional problem, but it has been very successful in, for instance, seeing all the different states of an active ribosome.
- dalke 3y agoDo you know what the author means by "Current visualization software, such as UCSF ChimeraX6 , can only render one or a few protein structures at the atomic level." I haven't used VMD for about 30 years, but even in the 1990s I was using it to visualize the full poliovirus structure (4 proteins in 2PLV * 60 copies, as I recall). It took about 6-10 seconds per update on our SGI Onyx, but again, that was 25 years ago.
- COGlory 3y agoI can only guess, but I believe that ChimeraX's rendering pipeline is single threaded (just an empirical guess based on my CPU usage when using it). Additionally, loading that many atom positions requires a huge amount of memory (I routinely use > 32 GB memory just loading a few proteins) and things start to slow down quite a bit. Loading a 60-fold icosahedral virus has used > 100 GB memory on my workstation, and resulted in a 0fps experience. It might render OK from the command line, but now imagine a few dozen of those, plus a cell, plus all the proteins in the cell...
- dalke 3y agoOdd. I can't see why. I think we had 128 MB on that IRIX box, and I know I loaded a 1 million atom structure with copies of 2PLV (full capsid plus a bit more to get to a million.) Each atom record has ~60 bytes (x, y, z, occupancy, bond list, resid, segid, atom name, plus higher-level structure information about secondary structure, connected fragments, etc.) We had our own display list, so another (x, y, z, r, color-index) per atom, giving 20 more bytes. We probably used a GL/OpenGL display list for the sphere, and immediate mode to render that display list for each point, so all-in-all about 100 bytes per atom, which just barely fits in 128 MB. That was also all single-threaded, with a ~0.1 Hz frame rate. Again, in the 1990s. I wanted to see what more recent projects have done. Google Scholar found "cellVIEW: a Tool for Illustrative and Multi-Scale Rendering of Large Biomolecular Datasets" (2017) at https://www.ncbi.nlm.nih.gov/pmc/articles/PMC5747374/ https://www.ncbi.nlm.nih.gov/pmc/articles/PMC5747374/ which says > The most widely known visualization softwares are: VMD [HDS96], Chimera [PGH04], Pymol [DeL02], PMV [S99], ePMV [JAG11]. These tools, however, are not designed to render a large number of atoms at interactive frame-rates and with full-atomic details (Van der Walls or CPK spherical representation). Megamol [GKM15] is a state-of-the-art prototyping and visualization framework designed for particle-based data and which currently outperforms any other molecular visualisation software or generic visualization frameworks such VTK/Paraview [SLM04]. The system is able to render up to 100 million atoms at 10 fps on commodity hardware, which represents, in terms of size, a large virus or a small bacterium. Following that is a section on Related Work: > With their new improvement they managed to obtain 3.6 fps in full HD resolution for 25 billion atoms on a NVidia GTX 580, while Lindow et al. managed to get around 3 fps for 10 billions atoms in HD resolution on a NVIDIA GTX 285. Le Muzic et al. [LMPSV14], introduced another technique for fast rendering of large particle-based datasets using the GPU rasterization pipeline instead. They were able to render up to 30 billions of atoms at 10 fps in full HD resolution on a NVidia GTX Titan Checking up on VMD, in "Atomic detail visualization of photosynthetic membranes with GPU-accelerated ray tracing" from 2016: > VMD has achieved direct-to-HMD rendering rates limited by the HMD display hardware (75 frames per second on Oculus Rift DK2) for moderate complexity scenes containing on the order of one million atoms, with direct lighting and a small number of ambient occlusion lighting samples. Those citations are 6-7 years ago, which make me scratch my head wondering why ChimeraX can't handle a picornavirus.
- Vt71fcAqt7 3y agoSlightly off topic but what do you think of the potential growth for GWAS? How much will it improve in the next 10 years? 2x? 10x?
- COGlory 3y agoNot an expert in that field. I can speculate wildly that I'm not super optimistic. I can say I've seen the narrative around the field start to give way slightly - shifting from "genome is all you need" to "genome is not enough". The problem is that many of the interesting or urgent pathologies have no obvious (or weak) associations. Or maybe the noise level is too high. So there's got to be a piece of the puzzle we're missing, or something is getting lost in the noise. Whether a neural network can pull something out of the noise remains to be seen, but if it can't, I'm not super optimistic about our chances. Overall I'd say trying to tackle the problem at the outcome level is probably more promising right now. Even if we can find good associations, we're still lacking therapies.