3 ms·
Author here. It's clear that you genuinely believe my intent here is to be deceptive, so I think your comment deserves a thoughtful response. For context, I wor
by lebovic 4y ago
Author here. It's clear that you genuinely believe my intent here is to be deceptive, so I think your comment deserves a thoughtful response. For context, I work on a computational biology infrastructure company that only uses cloud computing; my incentive is for more scientific computing to be on the cloud.
Responses to your points below. Azure does have some better HPC infrastructure than AWS, so maybe some of my response to your comment will be wrong. I'm happy to talk about this more if you want, you can reach me using the email address in my profile.
Spot/pre-emptible instances vs. the prices in this post: large CPU/RAM instances have availability issues, especially for spot instances. I spent a lot of time trying to exploit spot instance pricing, and for standard (run command-line tool that reads in file and writes out files) bioinformatics programs, spot instances haven't made sense averaging their performance over a long period after factoring in restarts and their corresponding data transfer vs. other options for decreasing instance cost (reserved instances, negotiating, etc).
Sending S3 links: yeah, AWS makes sending data easier! Although if the destination is not within same same cloud provider (or region), you get hit with a surprising large charge for sufficiently large files.
Input size vs. output size and storing the results: generally, I agree. Cloud storage costs aren't unreasonable for S3, but I want to note how significantly the pipeline can differ. A new Illumina NovaSeq sequencer (about the size of a copy machine) with dual S4 flow cells produces 6Tb every couple days. Some pipelines are definitely inefficient, but others have more raw data. Storing that data in an infrequent access or archive tier decreases the restore speeds and increases the restore cost. If you have 100 TB of data, that increases the cost of re-running data – especially if it's in cold storage or archived.
Improved algorithms and re-running large sets of data: sure, it's a trade-off between cost and the bandwidth of a queue than you can run. For some use cases, the cloud does make sense.
Hardware improvement cycles in bioinformatics: what software are they using in bioinformatics that uses a GPU, AlphaFold? From what I've seen, most computational genomics still happens on a CPU, although fields like computational chemistry use more GPUs.
Infrastructure components and easily installable components: yeah, this is a definite value-add of cloud services, and the off-AWS/GCP/Azure analogues aren't as good yet.
Cost of networking equipment vs. by-the-hour in a cloud: yeah, if you want results quickly and occasionally, this makes sense.
Overall, this post is about most of scientific computing, not all. For this to work, you need a smoothable queue of jobs. Most computational science (by % of compute) run in this context, in universities, larger/growing co's, and government research institutions. If you want instant scalability, the math is different.
- jiggawatts 4y ago> Azure does have some better HPC infrastructure than AWS That's somewhat surprising to hear, I just assumed AWS has equivalent products. My customer is 80% AWS and 20% Azure, so it's a useful data point to know that some HPC workloads are better off in Azure. > large CPU/RAM instances have availability issues, especially for spot instances. I've had two different ~128 vCPU instances running for days and days in my lab environment, but that's probably because my region tends not to have a lot of HPC workloads that would compete for spot instances. I've noticed that "popular" sizes in Azure such as D4, D8, E4, and E8 are pre-empted regularly, but the "special" sizes like HPC not so much. > surprising large charge for sufficiently large files. Both Azure and AWS use this as a "roach motel" to encourage vendors and partners to co-locate in the same cloud. It's unfortunate that they charge on the order of $100 per terabyte, but it is what it is. However, bioinformatics files compress well, and compared to getting something out of an on-prem traditional network, the egress fees are a bargain. My customer has a rural site with a "WAN" link. Their non-cloud option is to buy a NAS, replicate it to another NAS in a data center, and then build a permanent "file sharing solution". This is going to cost tens of thousands of dollars. They might share a few terabytes annually, which makes cloud egress fees look practically free in comparison. > A new Illumina NovaSeq sequencer (about the size of a copy machine) with dual S4 flow cells produces 6Tb every couple days. The scientists I talked to raised this, and to be honest I'm also concerned, especially as some of the groups I deal with are in rural areas "far from the cloud." (They analyse samples from cattle ranchers to try and prevent foot and mouth disease.) Let's say the machine generates 6 terabytes in 2 days, so 3 TB daily. Assuming that's the uncompressed data, it'll be about 1 TB after compression. The location I'm thinking of has a 500 Mbps link, but that can transfer that data volume in just 4 hours: https://www.wolframalpha.com/input?i=%28+1+TB+%29+%2F+500+Mbps https://www.wolframalpha.com/input?i=%28+1+TB+%29+%2F+500+Mb... One trick I discovered recently is that both the s3cmd and azcopy tools can take pipeline input. Combine that with a parallel compression tool like 'pigz' that can output to the pipeline, and you can have a workflow where the "raw" input files are compressed and streamed to the cloud storage at the same time. That alone can cut hours off the transfer time! > If you have 100 TB of data, that increases the cost of re-running data – especially if it's in cold storage or archived. Not necessarily. Azure Cold storage access has no special charges associated with it. It's a bit slower, but streaming reads were quite fast in my experience. Archive tier of course has some additional costs, but it's not a drama in most cases. For example, "high priority" retrieval costs extra, but normal priority appears to be free. > what software are they using in bioinformatics that uses a GPU From: https://developer.nvidia.com/blog/nvidia-hopper-architecture-in-depth/ https://developer.nvidia.com/blog/nvidia-hopper-architecture... "An example is the Smith-Waterman algorithm for genomics processing". Admittedly, GPU usage for genomics is still rare, but it is becoming more common. > If you want instant scalability, the math is different. In my example, think of 10-20 scientists doing semi-regular gene sequencing workloads and running related analyses. Sometimes needing a single machine with 2TB of memory, other times running 10,000 trivial jobs. In principle, the flexibility of the cloud is nearly optimal for a scenario like this.