4 ms·
It's interesting how big of a difference it makes when you allow local disk I/O, and you need to schedule around it. At Google they disintermediate storage and
by romed 8y ago
It's interesting how big of a difference it makes when you allow local disk I/O, and you need to schedule around it. At Google they disintermediate storage and all (not really all) I/O goes to their cluster FS (Colossus, sort of). They don't have to schedule around I/O resources because every process has access to the full I/O resources of the entire cluster at any time from any node. By contrast as soon as you let some open source or off-the-shelf commercial thing leak into your operations, it will demand ordinary POSIX disk I/O and then you've got big problems. I propose that some companies would actually be better off concentrating on disintermediated storage more, and I/O workload scheduling less.
- nemothekid 8y ago>I propose that some companies would actually be better off concentrating on disintermediated storage more, and I/O workload scheduling less. I feel like this isn't possible outside of Google. I've never had good performance (well any performance that you would feel comfortable running a database on) on any storage medium that wasn't POSIX disk I/O. I may be ignorant, but I've never seen a good disintermediate storage solution that is performant on something like AWS.
- makmanalp 8y ago> all (not really all) I/O goes to their cluster FS (Colossus, sort of). Well - sorta. Cluster FS just means disk I/O is now network I/O (EBS is a bit like this). Now you push the hairy problem to a different area: your network topology starts to matter a lot. You might end up having to think about not scheduling too many I/O heavy jobs behind the same switch or something. AND disk I/O is now competing with network I/O. AND it might now be hard to predict the destination of where all that I/O is going to be when scheduling the job, whereas previously it used to just be local. So the problem doesn't go away. But maybe you say it's just easier to wire everything up with gigE / infiniband so that you're massively under-capacity, and call it a day, and that might be a fair tradeoff.
- closeparen 8y agoThe “Storage as a Service” needs to run somewhere. You can’t escape from running a storage aware scheduler. Either it will be the same as your primary scheduler, or storage services will be special snowflakes.
- romed 8y agoThere are always daemons on the machines that aren’t under the control of the cluster scheduler (init or its equivalent, etc). The thing that exports the storage assets of the box can simply be one of those things.