4 ms·
First. Seastar & Scylla are really impressive work. Props Avi & team. Doing disk IO well from userspace is hard. There's obvious topics about durability that h
by mtanski 10y ago
First. Seastar & Scylla are really impressive work. Props Avi & team.
Doing disk IO well from userspace is hard. There's obvious topics about durability that have been covered on HN for years. Getting good performance out of modern drives is one of those things that doesn't get covered enough.
Take a prosumer drive like the Samsung 950 Pro (M2 form factor). It can 1GB to 2GB of streaming transfer and anywhere from 100k to 300k iops. All for about $180.
The system (kernel) interfaces and filesystems haven't really kept up. The only async interface is via libaio and the io_submit syscalls. If you ever worked with you know the limitations, it only works with O_DIRECT, has all sorts of requirements on your ops and very few guarantees. Random class / filesystems will just block on submit. XFS probably does the best here (if you have a recent kernel).
Once you went down this rabbit hole you're implementing your own page caching and replacement algorithms. And finally you get to the point where you need to worry about scheduling your IO because if you push down too many ops down to the kernel your response times become unpredictable (see: https://lwn.net/Articles/682582/ https://lwn.net/Articles/682582/ [paid till next week]).
Anyways, fascinating work & fascinating write up. Much nicer then another rehash about another async framework that only handles small async network requests.
- hendzen 10y agoAvi has an earlier post [0] where he shows that (recent) XFS is the only filesystem that actually executes io_submit asynchronously. http://www.scylladb.com/2016/02/09/qualifying-filesystems/ http://www.scylladb.com/2016/02/09/qualifying-filesystems/
- glommer 10y agoActually, what Avi has demonstrated in this article is that XFS is the only filesystem that executes it mostly asynchronously. Before we got started with the implementation of the I/O Scheduler (which we eventually wanted anyway for prioritization), I saw await time as reported by iostat as bad as 7s (truth be told, those weren't the best disks in the planet). That was basically XFS sleeping during io_submit due to the problem I have briefly mentioned in this article, with the allocation groups. If you limit the amount of requests the filesystem is consuming, then it is gone to the point that we started focused our attention in other areas. But it still has a couple of places in which it will resort to synchronous behavior. No Linux filesystem can execute io_submit completely asynchronous.