4 ms·
I recall that the group that created Spark had a bioinformatics project on Spark but I don't know what happened to it. All I could find now is a paper[1] hosted
by east2west 5y ago
I recall that the group that created Spark had a bioinformatics project on Spark but I don't know what happened to it. All I could find now is a paper[1] hosted by databricks.
[1]https://databricks.com/wp-content/uploads/2018/08/SSE15-40-Danford.pdf https://databricks.com/wp-content/uploads/2018/08/SSE15-40-D...
- heuermh 5y agoWe're here, still plugging along. ADAM is a genomics analysis platform with specialized file formats built using Apache Avro, Apache Spark, and Apache Parquet. Apache 2 licensed. https://github.com/bigdatagenomics/adam https://github.com/bigdatagenomics/adam
- dekhn 5y agoYep, that's the one I was thinking of (along with GNOMAD, which IIRC uses ADAM or some similar tech). My main complaint with ADAM was that they came up with their own file format (which had some flaws). But the general idea is the right one.
- heuermh 5y agoI'm interested in chatting with you about this, and genomics on Spark more generally, feel free to reach out on Github or via my username at the usual suspects.
- dekhn 5y agoI left this field, actually. I cofounded Google Cloud Genomics, and when I proposed that we pivot from working with the GA4GH (very stupid APIs) to working with ADAM (real data processing) I got kicked off the team. Since then I've come to see genomics as a minefield of bad practices and don't really work in the field any more, except to help scientists run their workflows in the cloud.