3 ms·
I think you are correct, but there is a case where saturating the cluster is maybe not the ideal situation. I think there's a latency trade-off there. So imag
by nulltype 11y ago
I think you are correct, but there is a case where saturating the cluster is maybe not the ideal situation.
I think there's a latency trade-off there. So imagine you have a job where latency matters, like a user wants to run a bigquery-style interactive query across 1TB of data, and ideally it should complete within 1 minute. Lets say it takes 10 instance hours (600 instance minutes) to complete this query.
With per-minute billing you could launch 600 instances and complete it in a minute. You could also do the same with per-hour billing, but you would overpay by a factor of 60x or you would launch 10 instances and wait an hour for the query to complete.
Assuming you have a bunch of these queries coming in, you could queue them up, but then the latency will suffer, because you'll have to delay starting a job until you get to that position in the queue.
If you wish to trade some wait time for some cluster efficiency, yes, you can just queue them up and then slowly scale the cluster up and down to keep 100% utilization.
However, it would be nice if you really could scale up and down at per-minute increments and let Google figure out how to get their cluster utilized efficiently and let them take advantage of their different customers with different workloads and preemptible instances.