6 ms·
As I mentioned in the other thread yesterday database and database like applications are going to be hit particular hard. Even more so on fast flash storage. Do
by mtanski 9y ago
As I mentioned in the other thread yesterday database and database like applications are going to be hit particular hard. Even more so on fast flash storage. Double whammy compared to apps just doing network IO.
And while databases try to minimize the number of syscalls they still end up doing a lot of them for read, writeout, flush.
- surajrmal 9y agoDon't flash devices use nvme (eg userspace queues) now and avoid the kernel all together for read and write operations? Shouldn't they have no impact?
- kllrnohj 9y agoNot by default they don't. That only works if you're willing to dedicate that entire physical drive to a single application anyway.
- mtanski 9y agoIn the future this possible on Linux with the filesystems that support DAX. Currently this all pretty experimental with lots of work being done in this space in the last two years. But this will require you to have the right kind of flash storage, right kind of fs, right kind mount options, and probably a different code path in userspace for DAX vs traditional storage. So we're a little ways away from this.
- kllrnohj 9y agoDAX doesn't appear related here at all. That is about bypassing the page cache for block devices that don't need one. That doesn't move anything from kernel land into userspace, certainly not in the app's process in userspace anyway.
- mtanski 9y agoIf you bypass the page cache you do not have read()/write() and mmap you avoid the syscall overhead. This matters a lot for high IOPs devices. Also these new fangled devices claim support word cache line sync using normal cpu flush instructions. Also avoiding fsync syscall.
- kllrnohj 9y agoOne does not follow the other. Where are any references to how this will let you bypass read & write? User-space applications are still interacting with a filesystem, which they access via read/write and not a block device. There's no talk in the DAX information about how this results in a zero-syscall filesystem API, and I'm not seeing how that would ever work given there would then be zero protections on anything. You need a handle, and that handle needs security. All of that is done today via syscalls, and DAX isn't changing that interface at all. So where is the API to userspace changing?
- mtanski 9y agoPlease re-read my above comment. There is no new API. The DAX userspace API is mmap. This work is experimental but you can mmap a single file on a filesystem on this device using new DAX capabilities. Most access will not longer require a syscall. This comes with all the usual semantics and trappings of mmap plus some additional caveats as to how the filesystem / DAX / hardware is implemented. Most reads/writes will not require a trip to the kernel using the normal read()/write() syscalls. Additionally, there is no RAM page cache baking this mmap instead the device is mapped directly at a virtual address (like DMA). Finally, flush for these kinds of devices is at the block level implemented using normal instructions and not fsync. Flush is going to be done using the CLWB instruction. See: https://software.intel.com/en-us/blogs/2016/09/12/deprecate-pcommit-instruction https://software.intel.com/en-us/blogs/2016/09/12/deprecate-... LWN.net has lots of articles and links in their archives from 2016/2017. It's a really good read. Sadly I do not have time to dig more of them up for you. Do a search for site:lwn.net and search for DAX or MAP_DIRECT.
- wtallis 9y agoNVMe means each drive can have multiple queues, but they're still managed by a kernel driver. You may be thinking of SPDK, which includes a usermode NVMe driver but requires you to rewrite your application. And many systems are still using SAS or SATA SSDs.
- 3pt14159 9y agoDo we have a performance estimate? I can eat 20 or 30%, but I can't eat 90%.
- mtanski 9y agoThis comment further down thread mentions it's 20% in Postgres. https://news.ycombinator.com/item?id=16061926 https://news.ycombinator.com/item?id=16061926
- eropple 9y ago...when running SELECT 1 over a loopback socket. The reply to that comment is accurate: that's a pathological case. Probably an order of magnitude off.
- mtanski 9y agoWe're still learning, but it looks like pgbench is 7% to 15% off: https://www.postgresql.org/message-id/20180102222354.qikjmf7dvnjgbkxe@alap3.anarazel.de https://www.postgresql.org/message-id/20180102222354.qikjmf7...
- eropple 9y agoI've seen that message. It acknowledges the same problems: do-nothing problems over a local unix socket. Real-world use cases introduce much more latency from other sources in the first place. I'm sticking with an expectation in the 2%-5% range.
- cookiecaper 9y agoYep, this is getting blown way out of proportion by all of these tiny scripts that just sit around connecting to themselves. Even pgbench is theoretical and intended for tuning; you're not going to hit your max tps in your Real Code that is doing Real Work. In the real world, where code is doing real things besides just entering/exiting itself all day, I think it's going to be a stretch to see even a 5% performance impact, let alone 10%.
- panarky 9y agoHow would you trade this knowledge? Intel has already dropped and AMD is up. Maybe there's more to move, but first-order effects are at least partially priced in already. But what about second-order effects? Seems like virtualization should be vulnerable (VMWare and Citrix), but maybe they actually benefit as customers add more capacity. Software-defined networking and cloud databases should also suffer though it's unclear how to trade these. AWS, Google Cloud and Azure might benefit as customers add capacity but there's no way to trade the business units. So what about cloud customers where compute costs are already a large percentage of total revenue? Netflix should be OK but Snap and Twilio could get squeezed hard. Akamai and Cloudflare might have higher costs they can't pass through to customers. And where's the upside? Who benefits? If the performance hit causes companies to add capacity, maybe semiconductor and DRAM suppliers like Micron would benefit.
- rufugee 9y agoNot sure why you are being downvoted. It's an interesting question. I'm in on AMD for the time, just to see how to flows.
- mtanski 9y agoI think the model for cloud vendors would be quite complicated. Not every version of the CPU and not every application is impacted as much (new intel processors with PCID will suffer less). Add on top of that the fact that a lot cloud customers over provision (there's good scientific papers on how much spare CPU capacity there is). Cloud service providers that sell things on a per request / real CPU usage model (vs reserved capacity) prob benefit more. Also, you can't just separate trading in AWS or GCE from the rest of the core business. Potentially business units of DELL, HP, IBM, ... should do better as people use this as a justification to upgrade overdue hardware they should cover 5% to 10% lower performance (needing more units to cover that).
- user5994461 9y agoAgree on that last paragraph. The only reasonable thing people can do is buy more hardware to cover the performance loss and/or buy more hardware that's not needed using the bug as a pretext to get the budget approved now.