15 ms·
I/O is no longer the bottleneck? (2022)
- wmf 9mo agoPrevious related discussion: https://news.ycombinator.com/item?id=33751266 https://news.ycombinator.com/item?id=33751266
- tomhow 9mo agoI/O is no longer the bottleneck - https://news.ycombinator.com/item?id=33751266 https://news.ycombinator.com/item?id=33751266 - Nov 2022 (326 comments)
- gnabgib 9mo agoThat's not actually this post (that post is by this OP.. confusingly) Harmen/@stabbles just chose the same title. What's here is a retort to that post from a week later (still from back in 2022). Edit: Ah, I saw @wmf has edited their comment to include "related"
- akoboldfrying 9mo agoThe author recently posted an addendum that describes a clever trick to go even faster, which I highly recommend reading: https://news.ycombinator.com/item?id=46503612 https://news.ycombinator.com/item?id=46503612
- pvorb 9mo agoBut I/O being the bottleneck never was about sequential reads, was it? I get the point of the article, though.
- geoctl 9mo agoWith modern CXL/PCIe, I guess it's not going to be that stupid to claim that RAM/memory controller is slowly becoming I/O on its own.
- throwaway94275 9mo agoOld IBM's term for RAM was "storage."
- geoctl 9mo agoI wonder whether the current huge funding in AI will ever lead to a revolution in computer architecture. Modern PCIe/CXL is already starting to blur the difference between memory and I/O. Maybe the future is going to be that CPUs, RAM, storage devices, GPUs and other devices are going to directly address one another like a mesh network. Maybe the entire virtual memory model will change to include everything to be addressed via a unified virtual memory space from a process/CPU perspective with simple read/write syscalls that translate into network packets flowing between the CPU and the destination (e.g. RAM/GPU/NVMe) and vice versa.
- fc417fc802 9mo agoDon't we already almost (but not quite) have that? PCIe devices can talk directly to each other (still centralized AFAIK though) and from the programmers perspective everything you mentioned is mapped into a single unified address space prior to use (admittedly piecemeal via mmap for peripherals). Technically there's nothing stopping me from mmaping an entire multi-terabyte nvme at the block level except for the part where I don't want to reimplement a filesystem from scratch in addition to needing to share it between lots of different programs.
- 9mo ago
- eliasdejong 9mo agoIncreasingly the performance limit for modern CPUs is the amount of data you can feed through a single core: basically memcpy() speed. On most x86 cores the limit is around 6 GB/s and about 20 GB/s for Apple M chips. When you see advertised numbers like '200 GB/s' that is total memory bandwidth, or all cores combined. For individual cores, the limit will still be around 6 GB/s. This means even if you write a perfect parser, you cannot go faster. This limit also applies to (de)serializing data like JSON and Protobuf, because those formats must typically be fully parsed before a single field can be read. If however you use a zero-copy format, the CPU can skip data that it doesn't care about, so you can 'exceed' the 6 GB/s limit. The Lite³ serialization format I am working on aims to exploit exactly this, and is able to outperform simdjson by 120x in some benchmarks as a result: https://github.com/fastserial/lite3 https://github.com/fastserial/lite3
- dehrmann 9mo ago> 6 GB/s Samsung is selling NVMe SSDs claiming 14 GB/s sequential read speed.
- johncolanduoni 9mo agoSequential read speed is attainable while still having a (small) number of independent sequential cursors. The underlying SSD translation layer will be mapping to multiple banks/erase blocks anyway, and those are tens of megabytes each at most (even assuming a 'perfect' sequential mapping, which is virtually nonexistent). So you could be reading 5 files sequentially, each only producing blocks at 3GB/s. A not totally implausible access pattern for e.g. a LSM database, or object store.
- eliasdejong 9mo ago> 14 GB/s Yes, those numbers are real but only in very short bursts of strictly sequential reads, sustained speeds will be closer to 8-10 GB/s. And real workloads will be lower than that, because they contain random access. Most NVMe drivers on Linux actually DMA the pages directly into host memory over the PCIe link, so it is not actually the CPU that is moving the data. Whenever the CPU is involved in any data movement, the 6 GB/s per core limit still applies.
- kevmo314 9mo agoThis was my instinct when NVMe SSDs first came out: I'd joke that now we have 2 TB of RAM. The real joke is on me though, some of these GPU servers actually have 2 TB of RAM now. Crazy engineering!
- npn 9mo agoNow? I had found some used epyc servers with 2TB ddr4 ram for around 5k usd yesteryear. Too bad I didn't purchase it.
- sroussey 9mo agoHe said GPU servers
- esjeon 9mo agoThe point is that we did have CPU servers with TBs of RAM. These machines are still pretty much relevant.
- npn 9mo agoHe didn't say VRAM. GPU servers are just servers with GPUs.
- sroussey 9mo agowho cares how much cpu ram a gpu server has? But yeah, if that is what he meant, then that is silly. Although... may be impossible to buy 2TB of RAM later in 2026! ;)
- ThreatSystems 9mo ago*Unless your in the cloud, then it's a metric to nickel and dime with throttling! On a more serious note, the performance of hardware today is mind boggling from what we all encountered way back when. What I struggle to comprehend though is how some software (particularly Windows as an OS, instant messaging applications etc.) feel less performant now than they ever were.
- nine_k 9mo agoThe answer, I suspect, is is the same as always: waiting for I/O in the GUI thread. Both Telegram and FB messenger are snappy; I didn't use anything else seriously as of late. (Especially not Teams, nor the late Skype.)
- fragmede 9mo agoThey could be way faster. They're snappy enough but still, so slow.
- josephg 9mo ago> waiting for I/O in the GUI thread The problem is sloppy programming. We knew how to make small, fast, programs 20+ years ago that would just scream on modern hardware. But now everything is bloated and slow. CPUs can retire billions of instructions per second. Discord takes 10+ seconds to open. I’m simply not creative enough to think up how to keep the cpu busy that long opening IRC.
- eru 9mo agoMoore's law really help you with throughput, but latency still requires good engineering. And you are right, that we got good UI latency even back in the 1980s. You just have make sure that you do the absolute minimum amount of work possible in the UI 'thread' and do gradual enhancement as more time passes. As an example, the Geos word processor on the C64 does nice line breaks at the end of words only. But if you type really fast, it just wraps lines when you hit exactly x letters, and later when it has some time to catch up, it cleans up the line breaks. That way it can give you a snappy user experience even on a comically underpowered system. But you can also see that the logic is much more complicated, than just implementing a single business logic for where line breaks should be. Complication means more bugs, more time spent writing and debugging and documenting etc.
- gary_0 9mo agoNot a new idea, but it's intriguing to think about an architecture that's just: CPU <-> caches <-> nonvolatile storage What if you could take it for granted that mmap()ing a file has the exact same performance characteristics as malloc(), aside from the data not going away when you free the address space? What if arbitrary program memory could be given a filename and casually handed off to the OS to make persistent? A lot of basic software design assumptions are still based on the constraints of the spinning rust era...
- eru 9mo agoYou can get something like this from Linux today. (And mmap is actually how you request memory from the kernel in almost all cases.) It's just that mmap is slower than using read/write, because the kernel knows less about your data access patterns and thus has to guess for how to populate caches etc.
- gary_0 9mo agoYes, I know mmap already sort of allows this (and has for well over a decade). To elaborate: when I want to, say, parse a megabytes-sized file, I don't muck about with mmap(), I just read() into a buffer; it's simple and it's fast enough even though I'm just wasting microseconds waiting for bytes on one fast chip to get copied into another slightly faster chip (and then copied into CPU cache). If I'm dealing with a larger amount of data, I'd be tempted to use database middleware to figure out all the platform-specific shuffling between RAM and disk (designed on the assumption that the disk is spinning rust, cough), but that pulls in yet another chunk of complexity. Instead, imagine if I could just state in one line of system-agnostic code "give me a pointer to /home/user/abc" and it does the right thing--assuming there was some way around mmap's current set of caveats. Imagine if I could turn a memory buffer into a file in one line of code and it Just Worked. Imagine if the OS treated my M.2 SSD as just another chip on the bus instead of still having a good amount of code on the hot path that assumes I'm manually sending bytes to a mechanical drive.
- tmerr 9mo agoThis has me thinking, it could be a fun project to prototype a convenient file interface based on pointers as a C library. I imagine it's possible to get something close to what you want in terms of interface (not sure about performance). I suspect in some cases it will be more convenient and in other cases it will be less convenient to use. The write interface isn't so bad for some use cases like appending to logs. It's also not bad for its ability to treat different types of objects (files, pipes, network) generically. But if you want to do all manipulation through pointers it should be doable. To support appending to files you'd probably need some way to explicitly grow the file size and hand back a new pointer. Some relevant syscalls would probably be open, ftruncate, and mmap.
- atrooo 9mo ago[flagged]
- leentee 9mo agoFrom my experience optimizing an OLAP database with high concurrency; lots of time the bottleneck is memory speed.
- dpc_01234 9mo agoIt's not about memory/CPU/IO, but latency vs throughput. Most software is slow because it ignores the latency. If you program serially waiting for _whatever_ it is going to be slow. If you scatter your data around memory, or read from disk in small chunks, or make tons of tiny queries to the DB serially your software will be 99.9% waiting idle for something to finish. That's it. If you can organize your data linearly in memory and/or work on batches of it at the time and/or parallelize stuff and/or batch your IO, it is going to be fast.
- verdverm 9mo agostill my bottleneck generally speaking, cloudvm/container filesys i/o sucks
- AmazingTurtle 9mo agoI read tons of comments like "It's not [this], it's [that] instead!" which is also wrong. The performance bottleneck is whatever resource hits saturation first under the workload you actually run: CPU, memory bandwidth, cache/allocations, disk I/O, network, locks/coordination, or downstream latency. Measure it, prove it with a profile/trace, change one thing, measure again.
- grayxu 9mo agoThe memory wall is an eternal problem when performing computations on the CPU
- stabbles 9mo agoAuthor here. There is a part 2 to this: https://stoppels.ch/2022/11/30/io-is-no-longer-the-bottleneck-part-2.html https://stoppels.ch/2022/11/30/io-is-no-longer-the-bottlenec...
- imtringued 9mo agoIf this is on a single core then the "6GB/s" guy is disproven not just in theory but also in practice.
- anonymoushn 9mo agoHello, a couple years ago I participated in a contest to count word frequencies and generate a sorted histogram. There's a cool post about it featuring a video discussing the tricks used by some participants. https://easyperf.net/blog/2022/05/28/Performance-analysis-and-tuning-contest-6#upd-july-20th-2022-results https://easyperf.net/blog/2022/05/28/Performance-analysis-an... Some other participants said that they measured 0 difference in runtime between pshufb+eq and eqx3+orx2, but i think your problem has more classes of whitespace, and for the histogram problem, considerations about how to hash all the words in a chunk of the input dominate considerations about how to obtain the bitmasks of word-start or word-end positions.
- stabbles 9mo agoAwesome! The slides with roofline analysis are great! https://docs.google.com/presentation/d/16M90It8nOK-Oiy7j9Kw27o9boLFwr6GFy55XFVzaAVA https://docs.google.com/presentation/d/16M90It8nOK-Oiy7j9Kw2...
- pjdesno 9mo agoSince no one else seems to have pointed this out - the OP seems to have misunderstood the output of the 'time' command. $ time ./wc-avx2 < bible-100.txt 82113300 real 0m0.395s user 0m0.196s sys 0m0.117s "System" time is the amount of CPU time spent in the kernel on behalf of your process, or at least a fairly good guess at that. (e.g. it can be hard to account for time spent in interrupt handlers) With an old hard drive you would probably still see about 117ms of system time for ext4, disk interrupts, etc. but real time would have been much longer. $ time ./optimized < bible-100.txt > /dev/null real 0m1.525s user 0m1.477s sys 0m0.048s Here we're bottlenecked on CPU time - 1.477s + 0.048s = 1.525s. The CPU is busy for every millisecond of real time, either in user space or in the kernel. In the optimized case: real 0m0.395s user 0m0.196s sys 0m0.117s 0.196 + 0.117 = 0.313, so we used 313ms of CPU time but the entire command took 395ms, with the CPU idle for 82ms. In other words: yes, the author managed to beat the speed of the disk subsystem. With two caveats: 1. not by much - similar attention to tweaking of I/O parameters might improve I/O performance quite a bit. 2. the I/O path is CPU-bound. Those 117ms (38% of all CPU cycles) are all spent in the disk I/O and file system kernel code; if both the disk and your user code were infinitely fast, the command would still take 117ms. (but those I/O tweaks might reduce that number) Note that the slow code numbers are with a warm cache, showing 48ms of system time - in this case only the ext4 code has to run in the kernel, as data is already cached in memory. In the cold cache case it has to run the disk driver code, as well, for a total of 117ms.
- obogobo 9mo agoWhat metrics does saturating memory bandwidth manifest as? ...iowait? 100% system CPU? How does one isolate memory as the bottleneck specifically?
- zozbot234 9mo agoIn process monitoring you just see 100% "cpu" use with the processor cores running in their low-medium frequency range and no real thermal issues (fans aren't spinning up). You can use perf indicators to specifically look at whether memory bandwidth is the issue.