19 ms·
Rosalind: A genomics toolkit in Rust running whole-genome pipelines on a laptop
- p4ul 5mo agoThis is interesting; thanks for sharing! I have been curious about the adoption of Rust in computational biology. I know that the folks at Saint Jude's [1] are also using Rust for their 'omics research. [1] https://github.com/stjude-rust-labs https://github.com/stjude-rust-labs
- the__alchemist 5mo agoI'm building a structural bio crate system in rust (na_seq, bio_files, bio_apis, dynamics and some more specialized). No one is using it AFAIK other than myself. I am using it to build a GUI multi-purpose structural bio GUI program called Molchanica. Note that this doesn't have much overlap with the traditional bioinformatics workflows like the OP (Rosland), or the one you linked to seem to be focused on.
- clmcleod 5mo agoThanks for the shout out!
- p4ul 5mo agoOh, thank you, @clmcleod! We've been following all your work closely in my team! I'm very bullish on the long-term prospects of Rust in computational biology—as well as research computing more generally.
- croemer 5mo agoWe rewrote Nextclade in Rust and are very happy. Works nicely both for CLI and client side browser with wasm. https://github.com/nextstrain/nextclade https://github.com/nextstrain/nextclade
- shpongled 5mo agoThere is a relatively widely adopted tool (100+ citations, >500k invocations collected via telemetry) for mass spectrometry-based proteomics written in Rust, and quite a few others in the works. [1] https://github.com/lazear/sage https://github.com/lazear/sage
- samuell 5mo agoYeah, there is actually a pretty big shift towards Rust in the comp bio / bioinformatics community. Nature even wrote a feature article about it a couple years ago: Why scientists are turning to Rust https://www.nature.com/articles/d41586-020-03382-2 https://www.nature.com/articles/d41586-020-03382-2 They mention the Rust-Bio [1] project by well known Snakemake author Johannes Köster & co, and there are some other widely used libraries like needletail [2] and noodles [3]. A cool smaller tool developed by performance wiz Ragnar Groot Koerkamp which was just published is Sassy [4] [5]. He has also been involved in developing some high performance SIMD based stuff (minimizers) [6]. [1] https://github.com/rust-bio/rust-bio https://github.com/rust-bio/rust-bio [2] https://github.com/onecodex/needletail https://github.com/onecodex/needletail [3] https://github.com/zaeleus/noodles https://github.com/zaeleus/noodles [4] https://github.com/RagnarGrootKoerkamp/sassy https://github.com/RagnarGrootKoerkamp/sassy [5] https://academic.oup.com/bioinformatics/article/42/5/btag244/8691840 https://academic.oup.com/bioinformatics/article/42/5/btag244... [6] https://github.com/rust-seq/simd-minimizers https://github.com/rust-seq/simd-minimizers
- qzgrid37 5mo ago[dead]
- peterfirefly 5mo agoShould have called it Raymond.
- flobosg 5mo agoOr rather Margaret: https://en.wikipedia.org/wiki/Margaret_Oakley_Dayhoff https://en.wikipedia.org/wiki/Margaret_Oakley_Dayhoff
- cmpb 5mo agoI'm not familiar with Margaret Oakley Dayhoff, but I am aware that Rosalind Franklin [1] was extremely important for our understanding of DNA, comparable to Watson/Crick, with whom she co-discovered the structure of DNA. So it seems "Rosalind" is at least very appropriate as a name for a genomics tool such as this. Not to say the other names mentioned aren't also deserving of similar honors [1] https://en.wikipedia.org/wiki/Rosalind_Franklin https://en.wikipedia.org/wiki/Rosalind_Franklin
- flobosg 5mo ago> I'm not familiar with Margaret Oakley Dayhoff Then you’re one of today’s lucky 10,000. Any time!
- philipallstar 5mo agoRosalind Franklin was the team lead of the research team that photographed DNA. The actual team member that took the key photo[0] was Raymond Gosling. That team didn't interpret the double helix structure of DNA that the photograph had captured - that was Watson and Crick working it out from the photograph. [0] https://en.wikipedia.org/wiki/Photo_51 https://en.wikipedia.org/wiki/Photo_51
- groby_b 5mo agoIt's not quite that clear-cut. Franklin was pretty clear on the helical structure in both research notes and papers, but she didn't quite nail the overall structure (2 strands with opposing winding, complementing bases). Fundamentally, she suffered the curse of the experimental scientist - waiting for actual data before being willing to build a model. Watson & Crick postulated ahead based on partial data.
- shauniel 5mo agoI would love to hear about what the sacrifices are, but this project really looks amazing.
- bonsai_spool 5mo agoDidn't see a publication or preprint for this - is there one?
- boron1006 5mo agoLots of bad smells in this repo.
- the__alchemist 5mo agoDo you have some examples to look at? I am curious.
- semiinfinitely 5mo agobioinformaticians have been making these useless bioinformatic-toolkit-in-my-favorite-programming-language repos for years
- gilleain 5mo agoHate to agree, but it is true. For a while, I think, the main sequencing framework was in perl (Bioperl). Not sure what was best for structures - possibly Biojava? It is very tempting, though - 'just' make a nice, clean API in your favourite language (eg Haskell, Ruby, ...) and everyone will flock to use it! Maybe.
- alice-fishr 5mo agoWhy don't you mention Biopython? Bioperl is already too old and not much up-to-date with newest data.
- flobosg 5mo agoHe’s talking about the past (“For a while, …”). Up to early 2010s, I would say.
- maxall4 5mo agoWell, what else are we going to do while waiting for the bench scientists to finish collecting data?
- asdff 5mo agoDissertationware is common in a lot of fields, honestly.
- _f78k 5mo ago> A deterministic genomics engine with a compact memory footprint. Uhh... are there stochastic genomics pipelines?
- flobosg 5mo agoA quick search gave me for example this one: https://genome.cshlp.org/content/26/1/36 https://genome.cshlp.org/content/26/1/36
- pazimzadeh 5mo agoi think these are more relevant examples https://www.cs.cmu.edu/~ckingsf/software/sailfish/ https://www.cs.cmu.edu/~ckingsf/software/sailfish/ https://www.nature.com/articles/nmeth.4197 https://www.nature.com/articles/nmeth.4197
- mfld 5mo agoI guess the author refers to the fact that many well-known tools have some randomness built-in. The most obvious one is differences due to the order of parallel processing. But these differences are often small and have no significant downstream effects. They are mostly inconvenient for regression testing.
- vatsachak 5mo agoLooking at the commenting pattern, it seems like AI unfortunately
- croemer 5mo agoThose are all the tests for alignment. They don't even check the alignment is correct. Just that there are no errors. This is a joke: https://github.com/logannye/rosalind/blob/main/tests/alignment_pipeline.rs https://github.com/logannye/rosalind/blob/main/tests/alignme... Looks like total slop to me. All code in one commit, then a bunch of commits polishing the Readme. No release, no updates in half a year.
- stelsmind 5mo ago[flagged]
- logannyeMD 5mo agoHey guys, this is my github repo. Glad it's received some interest - I figured HN might be the culprit when it suddenly jumped ~100 stars despite not working on the code base since last year. I prototyped this out of personal curiosity last year and moved on abruptly so there's a lot of gaps I still need to close and knobs that need to be optimized. But if people genuinely find "deterministic genomics workloads on edge devices" proposal useful, I'll begin refining the code tonight and try to make it as useful as possible. If you have any particular bioinformatics tasks or use cases that you want to be feasible on edge devices, lmk and I'll work on integrating new capabilities. Always happy to be helpful
- croemer 5mo agoYour website bio and LinkedIn don't match at all. Is the LinkedIn link on your website wrong? Update: yes it is. This is the correct one: https://www.linkedin.com/in/logan-nye https://www.linkedin.com/in/logan-nye You're doing too much vibe coding and not enough checking/testing. LinkedIn link on your website points to: https://linkedin.com/in/logannye https://linkedin.com/in/logannye Website bio: https://www.logannye.io/about https://www.logannye.io/about
- whateveracct 5mo agooh wow lol never seen that one before
- woodrowbarlow 5mo agothey weren't expecting to receive attention out of the blue today. it seems rude to attack someone's engineering skills because an online profile is out of date.
- Tuna-Fish 5mo agoIt's not out of date, it's pointing at the wrong person.
- whateveracct 5mo agothis isn't an attack. this is a data point that they don't review the slop coming out of the LLLM
- a_bonobo 5mo agoThere has been a bit of a 'trend' to rewrite common bioinformatics/comp-bio into faster languages (Rust) via LLMs, OP's repo seems to be an early example. Seqera Labs has a bit of a manifesto: https://rewrites.bio/ https://rewrites.bio/ Heng Li has an overview here too: https://lh3.github.io/2026/04/17/the-ai-rewrite-dilemma https://lh3.github.io/2026/04/17/the-ai-rewrite-dilemma IMHO it's... OK? Bioinformatics code quality is generally poor, untrained biologists writing functioning code that is poor in scoping, but works. (Unguided) LLMs write on that level, too, so not much harm done.
- ahartman00 5mo agoHow well tested would you say these libraries are? It doesn't sound promising, sadly. If there are comprehensive test suites, that would go a long way to ensuring new, faster tools arent producing subtly wrong answers. That's a pretty big deal, just because the code compiles or there is no exception thrown doesnt mean the analysis was correct.
- Gethsemane 5mo agoIt's very context-dependent - the seqera rewrites so far seem to be pretty reliable, most of the work was spent merging the functions of multiple data QC tools into a single program (previously, there was a lot of redundancy that wasted compute). The success of other rewrites that I've seen tends to depend on the author's care/experience and usefulness. In my experience, bioinformaticians are fairly slow on the uptake of new software which might actually be an advantage here :-) In defense of a lot of these bioinformatics-specific rewrites, there are some really dodgy coding practices and bugs that exist in well used tools, so there is scope for genuine improvement. The most recent release of minimap2 fixed some bugs identified in a rewrite, for example: https://github.com/lh3/minimap2/releases/tag/v2.31 https://github.com/lh3/minimap2/releases/tag/v2.31
- mriet 5mo agoRealistically, without data from a large testset that compares this thoroughly to Samtools (and others?), I wouldn't touch this. Note to the OP: specify a focus please? short, long, mega-long read and bacterial, human, small plant or large plant genome? Alignment heuristics and performance differ significantly across those axes.
- byrohitrajan 5mo ago[flagged]
- devlovstad 5mo agoI work with genomics pipelines in my day job. This repo does not seem quite ready for serious usage until a comparison is made with existing tools such as Bowtie 2/samtools/Strelka or similar. For cancer genomes, it's also a bit limiting that it does not call structural variants instead of just SNVs/indels.
- samuell 5mo agoI shared this since it seems to address a somewhat similar niche that I have had hopes to one day develop, based on FlowBase [1]; A library of streaming processing components based on basic operations, that can be easily stitched together into larger pipelines in a compiled language that can run on smaller hardware too. FlowBase or I didn't have much of ideas about how to keep data structures compact, as the linked library does, and I was mostly aiming to make it really easy to build streaming pipelines. I haven't yet got my head around how the composability story is in rosalind though, so would be interested in any pointers or examples on how this would be done using it. [1] https://github.com/flowbase/flowbase https://github.com/flowbase/flowbase
- danborn26 5mo agoRust is a great fit for genomics. Processing whole genomes locally on a laptop is a huge step up from typical Python pipelines.
- vfalbor 5mo agoHave you tested with other similar softwares such as Blast, which is the most common?
- Jerry2 5mo agoAwesome piece of software! Quick side question... does anyone have a recommendation for a DNA genotyping service that prioritizes privacy? I'm looking for a company that provides private results and doesn't add them to any sort of database (dystopian or otherwise). I'd love to get my DNA profile, but I'm concerned about privacy issues. :\
- samuell 5mo agoYou might try sequencing your DNA at home :) https://iwantosequencemygenomeathome.com/ https://iwantosequencemygenomeathome.com/ (Well, the guy offers to do it for you too, delivering the data on an USB stick).
- Jerry2 4mo agoThanks!
- Gethsemane 5mo agoUltimately you're not going to find a service that can guarantee privacy, but your best bet might be to extract DNA at home (though tricky without a centrifuge etc...) and submit it to a standard sequencing provider novogene, plasmidsaurus etc. Realistically, they'll hold onto the data for a couple of months as part of the order, then delete it to clear up space. A bunch of discordant sets of DNA sequence without metadata isn't exactly useful for nefarious purposes! I wouldn't recommend sequencing at home unless you are very enthusiastic...
- Jerry2 4mo agoThank you so much!
- penciltwirler 5mo agoblatant copyright infringement of https://rosalind.info/problems/locations/ https://rosalind.info/problems/locations/
- sspoisk 4mo ago[flagged]