4 ms·
> R is also thriving in bioinformatics with Python trailing behind as an afterthought. Maybe this is sub-field dependent, I'm a bioinformatician who hasn't tou
by _Wintermute 2y ago
> R is also thriving in bioinformatics with Python trailing behind as an afterthought.
Maybe this is sub-field dependent, I'm a bioinformatician who hasn't touched R in about 5 years, and everything is now in python.
- enmce 2y agoSame expirience,about 13 years in the field, most tasks done in python.
- tetris11 2y agooh wow, I made a sweeping generalization - apologies. In single cell, bulk RNA-seq, ATAC, ChIP, I'd say that there are more R packages for the analysis of these omics than Python packages.
- flobosg 2y agoPython is catching up on single-cell transcriptomics; see AnnData, Scanpy, et al.
- tetris11 2y agoScanPy/AnnData has been dead in the water for a while now, and most people use Seurat due to its operability with many many downstream extensions
- flobosg 2y ago> has been dead in the water for a while now Both are under active development and are used in several transcriptomics atlas projects, as far as I can tell.
- tetris11 2y agoThose atlases were established back when Scanpy and Seurat were relatively beta and were still fighting out the tool space. Look at the packages now for integration, pseudo time, pseudobulk - R (and therefore Seurat) dominates heavily
- valarauko 2y agoDisagree - Seurat had first mover advantage with single cell but sucked with larger datasets that Scanpy could handle till the big change in Seurat 5. The preference for either Seurat/Scanpy is incredibly lab specific. That said, Seurat is better documented for sure, but the ecosystem for both is incredibly rich and flourishing.
- cd4plus 2y agoyeah, I also disagree with this. It's true Seurat is still heavily used for scRNA/scATAC but I see most new models increasingly being written/tooled for python and based on anndata. Geneformer, scGPT, scVI etc. I wish there was better operability between the scverse stuff and Seurat, but Seurat went their own way from SCE/bioconductor so that's probably not going to happen.
- tetris11 2y agoTo be fair to them, getting anything submitted to buoconductor requires a ton of effort, and the pay off is often less concise code
- _Wintermute 2y agoI have no dog in the Scanpy/Seurat argument, but AnnData is becoming very popular as a data format even outside of single-cell omics.
- mbreese 2y agoIt’s also training dependent. It took a long time for Perl to go away from daily use in bioinformatics, solely dependent upon how common it was twenty years ago. In our lab, there is me, a (mid career) polyglot who mixes Python, Go, Java, and R daily. I use tab delimited text files to transfer data. I also grew up coding in C++ and like learning new languages. We also have a (mid career) staff scientist who grew up in Perl, but switched 100% to R and an (early career) postdoc who has always used 100% R. For both of these people, if work can be done in R, it is. If it can’t be done in R, they figure out how to do it in R anyway (even if that is shelling out to another program). We also have a (young) grad student that is 95% Python. They try to keep to the Python tooling, even though they are quite aware of the R ecosystem. There is a generational shift in the field, and it is more apparent each year. I find it interesting that Python took over from Perl first and now it’s trying to take over from R.