4 ms·
Yep, GPUs were the first thing we looked at. :) (And to be fair, our current approach is pretty close to a brute-force search!) The main reasons we decided aga
by joshma 11y ago
Yep, GPUs were the first thing we looked at. :) (And to be fair, our current approach is pretty close to a brute-force search!)
The main reasons we decided against a GPU-based approach were cost and scalability:
- We support a dozen+ reference genomes (eg for difference species), and plan to support a lot more (including eventually supporting custom genomes that users provide). Assuming we want to support a few concurrent searches against the same genome, we'd need a few GPUs per genome, and this gets expensive pretty quickly on AWS.
- Our fleet is now non-homogeneous, and now if machine X fails we need to restore machine X' with the same set of genomes.
- If certain genomes are more popular than others, we'll likely have GPUs spun up that aren't being used much (only one lab might be investigating a certain genome, for example). I suppose you could swap genomes in and out of memory as they're accessed, but again it's more complex to manage resources.
- Our current approach allows us to add genomes ad-hoc - hypothetically, you could point us to your own genome on S3 and we'd be able to work with it.
We hint at it towards the end, but we're actually switching to AWS Lambda soon - based on early calculations, it could cost us as little as $50 a month to run everything!