11 ms·
How fast are Linux pipes anyway? (2022)
- chris_armstrong 3y agoAbsolutely amazing, I know about page tables and the like but tying it to performance analysis with `perf` makes it clear how central it is to throughput
- epistasis 3y agoFantastic article, I learned a lot despite pipes being a bread and butter user tool for me for a quarter century.
- bloopernova 3y agoPipes are fast enough for iterating and composing cat sed awk cut grep uniq jq etc etc.
- mg 3y agoOne surprising fact about Linux pipes I stumbled across 4 years ago is that using a pipe can create indeterministic behavior: https://www.gibney.org/the_output_of_linux_pipes_can_be_indeter https://www.gibney.org/the_output_of_linux_pipes_can_be_inde...
- jstimpfle 3y agoNot surprising, the pipe you've created doesn't transport any of the data you've echoed. (echo red; echo green 1>&2) | echo blue This creates two subshells separated by the pipe | symbol. A subshell is a child process of the current shell, and as such it inherits important properties of the current shell, notably including the open file descriptor table. Since they are child processes, both subshells run concurrently, while their parent shell will simply wait() for all child processes to terminate. The order in which the childs get to run is to a large extent unpredictable, on a multi-core system they may run literally at the same time. Now, before the subshells get to process their actual tasks, file redirections have to be performed. The left subshell gets its stdout redirected to the write end of the kernel pipe object that is "created" by the pipe symbol. Likewise, the right subshell gets stdin redirected to the read end of the pipe object. The first subshell contains two processes (red and green) that run in sequence (";"). "Red" is indeed printed to stdout and thus (because of the redirection) sent to the pipe. However, nothing is ever read out of the pipe: The only process that is connected to the read end of the pipe ("echo blue") never reads anything, it is output only. Unlike "echo red", "echo green >&2" doesn't have stdout connected to the pipe. Its stdout is redirected to whatever stderr is connected to. Here is the explanation what ">&2" (or equivalently, "1>&2") means: For the execution of "echo green", make stdout (1) point to the same object that stderr (2) points to. You can imagine it as being a simple assignment: fd[1] = fd[2]. For "echo blue", stdout isn't explicitly redirected, so it gets run with stdout set to whatever it inherited from its parent shell, which is (probably) your terminal. Seeing that both "echo green" and "echo blue" write directly to the same file (again, probably your terminal) we have a race -- who wins is basically a question of who gets scheduled to run first. For one reason or other, it seems that blue is more likely to win on your system. It might be due to the fact that the left subshell needs to finish the "echo red" first, which does print to the pipe, and that might introduce a delay / a yield, or such.
- tuatoru 3y agoThank you for taking the time to write this very detailed and lucid explanation.
- jcrites 3y agoFor additional clarification, `echo` doesn’t read from stdin, so `… | echo xyz` doesn’t do what you probably assume. Try running `echo a | echo b` and you’ll see that only “b” is printed. That’s because `echo b` doesn’t read the “a” sent to it on stdin (and also doesn’t print it). If you want a program to read from stdin and write to stdout, you can use the `cat`, e.g. `echo a | cat` will print “a”. Lastly, be aware that `echo` is usually a shell builtin that functions like `print`. I’m not sure of all the ways that it might behave differently, but something to be aware of (that it’s not a child process like `cat`).
- dietrichepp 3y agoThe way that shell builtins behave differently here is that SIGPIPE can take out the whole shell on the left side when echo is built-in. When you /bin/echo red, then it's a subprocess, and its parent shell continues on, so you always get green somewhere in the output.
- paulddraper 3y agotl;dr Piped commands run in parallel not in serial. (The data "runs" in serial.)
- rixed 3y agoI don't think your message (or others) does justice to the original blogpost. Yes the pipe runs two subcommands in parallel but that is not why the blogpost is interesting (or its author surprised). It's because 'echo red' is supposed to block, thus introducing synchronization between the two branches of the pipe, yet it doesn't! And I must confess, when reading the command my first though was: "Ok so that first echo will die with a SIGPIPE and stderr will be all about the broken pipe." And I was wrong, because of that small buffer. I wonder what other unices do allow a write to a broken pipe to complete successfully?
- arp242 3y agoAre there cases where his causes real-world problems? Because to be honest this example seems rather artificial.
- 4death4 3y agoThat may have been surprising, but, if you think about it a little deeper, it makes perfect sense. Programs in a pipeline execute concurrently. If they didn’t, pipelines wouldn’t be useful. For instance a pipeline that downloads a tar file with curl and then untars it. If you wait for curl to finish before running tar, you run in to all sorts of problems. For instance, where do you store the intermediate tar file if it’s really large? Tar needs to run while curl is running to keep buffers small and make execution fast. The only control flow between pipeline programs is done via stdin and stdout. In your example program, you write to stderr so naturally that’s not part of the deterministic control flow.
- eru 3y ago> If they didn’t, pipelines wouldn’t be useful. Pipes would still be a useful way to structure your program. They would just be less useful.
- psd1 3y agoPowershell implements pipelines deterministically and without concurrency, and you can be very precise about it. Of course, it will use OS pipes if you include binaries in your pipeline. Nushell looks like it also has an internal implementation of pipelines. But I can't read rust so that's just my assumption.
- 4death4 3y agoWhat do you mean “without concurrency”? One program runs entirely before the other starts?
- psd1 3y agoPowershell pipelines are an engine construct rather than OS pipes or file descriptors. (If you include OS binaries in a PS pipeline, it will map the internal pipeline to OS pipes for that element of the pipeline, of course.) Every Powershell command has a begin, process, and end block. (If you don't write these explicitly, your code goes in an implicit end block.) When a pipeline is evaluated: 1. From left to right, the begin block of each command is run, sequentially. No process or end blocks are run until every begin block has run. 2. Each command's process block is run, once per object piped in. A process block can output zero, one or many objects; I'd have to check on a computer, but IIRC this is "breadth-first" - each object that a process block outputs is passed to the next process block before returning control to the current process block. 3. After all process blocks are exhausted, from left to right, each command's end block is run. Commands that did not declare a process block receive all piped objects as a single collection. Any output from the end block triggers the process block to the right. 4. When all end blocks have completed, the pipeline is stopped 5. Errors in Powershell can be terminating or non-terminating. When a terminating error is thrown, the pipeline is stopped 6. There is a special StopPipeline error which stops the pipeline but is handled by the engine so the user never sees it. That's how `select -First 5` works (for PS `select`, not gnu select). Pipelines only operate on streams 0 and 1, as with OS pipes. The other streams (ps has 7) are handled immediately, modulo some buffering behaviour intoxicated for performance reasons. Broadly speaking, the alternate streams are suppressed or enabled by defaults and by switches on each command individually and are rendered by the engine and given to the console to display. But they can also be redirected or captured in variables. You can do asynchrony in Powershell; threading is offered by a construct called "runspaces". These are not inherently connected to the pipeline, but pipelined commands can implement them, e.g. `foreach -Parallel {do-stuff}`
- xorcist 3y agoIt that surprising? What would you have guessed output would look like, and why? Perhaps that information would help straighten out any confusion. The command, perhaps intentionally, looks unusual (any code reviewer would certainly be scratching their head): There's an "echo red" in there but it's never sent anywhere (perhaps a joke with "red herring"?). There's an "echo green" sent to stderr, that will only be visible if it terminates before "echo blue". The exact order would be dependent on output buffering, which will depend on which time slice is sorted first, which will vary with number of cpus and their respective load. So yes, it will be indeterministic, but in the same way "top" is.
- Racing0461 3y agoChatgpt was able to figure this out with a simple "what does the following do". But it could also be a case of chatgpt being trained on your article. >>> Note: The ordering of "green" and "blue" in the output might vary because these streams (stdout and stderr) might be buffered differently by the shell or operating system. Most commonly, you will see the output as illustrated above.
- leodag 3y agoThat's wrong though, it's got nothing to do with different buffering (which is usually done at the application level, by the way).
- heavyset_go 3y agoI'm genuinely curious, how else could this work? It's like spawning threads, it's inherently indeterministic.
- psd1 3y agoMy shell throws an error if I try to pipe to a command that doesn't accept piped input. It's just better design. This is also why python sucks - if you feed it garbage, the error may surface a long way away and it may do a lot of damage while it's underwater
- oldbbsnickname 3y agoIf one enjoys fast, 0-copy I/O on Linux, here's an article.[0] PS: Precision of language to avoid confusion: "Indeterministic" is a philosophy term, while the CS term is "nondeterministic". 0. https://blog.superpat.com/zero-copy-in-linux-with-sendfile-and-splice https://blog.superpat.com/zero-copy-in-linux-with-sendfile-a...
- sbjs 3y agoI remember using linux pipes for a shell-based irc client like 12 years ago. For most application uses, they're plenty fast enough. Kinda wish I had the source code for that still.
- nh2 3y ago(2022) Previous discussion: https://news.ycombinator.com/item?id=31592934 https://news.ycombinator.com/item?id=31592934
- dang 3y agoThanks! Macroexpanded: How fast are Linux pipes anyway? - https://news.ycombinator.com/item?id=31592934 https://news.ycombinator.com/item?id=31592934 - June 2022 (200 comments)
- mannyv 3y agoHow fast are they compared to raw memory throughput? It's interesting that memory mapping is so expensive. I've often wondered the price that everyone pays for multiple address spades. Is isolation really worth it?
- formerly_proven 3y agoThe relative performance cost of virtual memory was way higher in days past, but people considered it worth it for the increased system reliability.
- DiabloD3 3y agoTL;DR: Maximum pipe speed, assuming both programs are written as optimally as possible, is approximately the speed of what one core in your system can read/write; this is because, essentially, the kernel maps the same physical memory page from one program's stdout to the other's stdin, thus making the operation a zerocopy (or a fast onecopy in slightly less optimal situations). I've known this one for awhile, and it makes writing shell scripts that glue two (or more) things together with pipes to do extremely high performance operations both rewarding and hilarious. Certainly one of the most useful tools in the toolbox.
- NortySpock 3y agoI assume for heterogenous cores (power vs efficiency cores) it bottlenecks on the throughput of the slowest core?
- DiabloD3 3y agoSurprisingly no. I'd expect similar performance. In these designs, the actual memory controller that talks to the RAM is part of an internal fabric, and the fabric link between the core and the memory controller is (technically) your upper limit. For both Intel and AMD, the size of the fabric link remains constant to the expected performance of the different cores, as the theoretical usage/performance of the load/store units remain otherwise constant in relation, no matter if it is a big core or a little core. Also, notice: the maximum performance of load-store units is your actual upper limit, period. Some CPUs historically never achieved their maximum theoretical performance because the units were never engaged optimally; sometimes this is because some ports on the load/store units are only accessible from certain instructions (often due to being reserved only for SIMD; this is why memcpy impls often use SSE/AVX, just to exploit this fact). That said, load-store performance usually approaches that core's L2 theoretical maximum, which is greater than what any core generally can get out of its fabric link. Ergo, fabric link is often governing what you're seeing in situations like this. On Intel and AMD's clusters, the memory controller serving their respective core cluster designs requires anywhere from 2 to 4 cores saturating their links to reach peak performance. Also, sibling threads on the same core will compete for access to that link, so it isn't merely threads that get you there, but actual core saturation. On a dummy benchmark like proposed in the linked article, the performance of a single process being piped to another process, either in the situation of "both processes are actually on the same big core, simultaneously hyper-threading", or "two sibling little cores in the same core cluster, being serviced by the same memory controller", the upper limit of performance should approximate optimal usage of memory bandwidth, but in some cases on some architectures this will actually approximate L3 bandwidth (a higher value). Also, as a side note: little cores aren't little. For a little bit more silicon usage, and a little bit less power usage, two little cores approximate one big core /w two threads optimally executing, even in Intel's surprisingly optimal small core design, but very much true in Zen4c. As in, I could buy a "whoops, all little cores" CPU of sufficient size for my desktop, and still be happy (or, possibly, even happier).
- Borg3 3y agoHah, nice article :) I remember fighting with Cygwin pipe implementations to have decent performance from them. They are hella slower compared to Linux, but still usable, just tricky to pass data in/out.
- whalesalad 3y agoLove the Edward Tuftian aesthetic of this site. Although above a certain viewport width I would imagine you want a `margin: 0 auto` to center the content block. On a 27" display it is tough to read without resizing the window.
- ldoughty 3y agoI have to agree... I really like the side-notes to get more details/explanation. You can skip the side-notes to keep reading and stay on the main story, but get what normally would be included in parenthesis or otherwise as an in-line comment.... Best of both worlds here I think. If I actively maintained a blog, I'd probably steal this design! :-)
- emmelaich 3y agoIs there some standard css/html way of pushing side notes or pics into the first column if viewing width is too small? That would be the best of both worlds!
- whalesalad 3y agoresponsive design concepts would enable this
- deleted 3y ago[deleted]
- jcrites 3y agoAre there good data handling libraries that provide abstractions over pipes, sockets, files, and memory and implement optimizations like these? I'd be interested in knowing if there are such libraries in C, C++, Rust, or other systems languages. I wasn't familiar with some of the APIs mentioned in the article like splice() and vmsplice(), so I wondered if there are libraries that I might use when building ~low-level applications that take advantage of these and related optimizations where possible automagically. (As another commenter mentioned: these APIs are hard to use and most programs don't take advantage of them) Do libraries like libuv, tokio, Netty handle this automatically on Linux? (From some brief research, it seems like probably they do)
- jeromegn 3y agoThere’s a crate for tokio, so it’s not automatic but might still be interesting: https://lib.rs/crates/tokio-splice https://lib.rs/crates/tokio-splice
- duped 3y agoThis may go against the grain but this isn't really worth abstracting over since it's not portable. You'll probably want to implement it by hand everywhere you need it. Higher level code only uses them rarely because they're pretty special purpose and they have to be specialized for Linux. If you're shuffling data around without looking at it only on Linux, splice is useful. There's not that many applications that have that property (something like say, TCP/UDP proxies definitely need it - but your bog standard HTTP server? Not so much). And if you are writing these apps then the buzzwords like "zero copy" come up often, and splice is one of the first results you'll see.
- NavinF 3y agoThe main reason why people write abstractions over stuff like this is to make it portable. I'm sure there's something similar to vmsplice on every relevant OS. The library can also fallback to write_read if you're targeting some ancient platform
- 3y ago
- codercowmoo 3y agoAnyone see the stonks image hidden quite well behind the first table? I could only see it because of my dark mode extension, otherwise I guarantee I wouldn't have caught it.
- deleted 3y ago[deleted]
- nathants 3y agopipes are great. is the other process on another cpu or another machine? honestly who cares. https://github.com/nathants/s4/blob/master/examples/nyc_taxi_bsv/count_rides_by_date.sh https://github.com/nathants/s4/blob/master/examples/nyc_taxi...
- Too 3y agoSo if I understand correctly, vmsplice is more of a mini shared memory mechanism between two processes, if used on both the reader and writer end simultaneously? Meaning both processes need to be exceptionally careful in when they read and write to the buffers and how it is returned after use. Hot, yet scary at the same time. Other main takeaway, it’s a bit sad that the naive implementation everybody will write, is 20x slower than what is possible. Exceptionally written article btw.
- winternewt 3y agoAnd if you try to write the 20x faster version, your coworkers will think you are over-complicating and not being a team player.
- db48x 3y agoNot necessarily. Good comments go a long way.
- crabbone 3y agoI have a very long list of things that were good, worked well, and ended up rejected because the team didn't want to put in effort to learn how they work. My conclusion so far is that if you want to make things work well, you shouldn't be working on a commercial project, use a unpopular language with a steep learning curve to filter out those who'd be a drag on your project. Maybe you don't have to be a jerk, but being blunt helps. Below are some examples of initiatives that were meant to improve things and how they failed due to other programmers being lazy and / or ignorant. When ActionScript was a thing it competed with HaXe. A similar (also ECMAScript-related) language with a small but dedicated community, a compiler that was hugely superior to to MXMLC (official Adobe compiler for AS3) and a bunch of features intended to improve code correctness and performance. I was hired by a company making a "PowerPoint online" kind of product. The main system component was a large Flex (AS3) applet that was hugely inefficient especially in terms of how it utilized network. It had to load huge shared libraries with assets (mostly clip art) every time users wanted to either edit or watch a presentation. The AWS bill was growing dangerously big. My mission was to find a solution to reduce the network activity. My idea was to create a separate player component that would extract the relevant assets from the libraries server-side, compile them into individual SWFs. The reason was that the final presentations were loaded a lot more often and by first-time users (i.e. no caching). HaXe was the ideal language because it already had a library that could generate a large subset of SWF, and it could compile both to AS3 and to C++, so that generation could also be done on a server using a more efficient implementation. After several month of work, I produced a set of programs that could generate SWFs both server-side and client side and showed how this would improve the network activity. The other programmers on the AS3 team, who earlier promised to get familiar with HaXe, since they had to incorporate the new player component into the existing Flex applet... didn't hold their part of the bargain. No matter the amount of help I provided, they simply wouldn't do anything to incorporate the new component, instead making claims that grew more bizarre and more untrue as time went by. Having spent more time trying to convince the team to adopt my code rather than writing it, I decided to look for a different place to work at. In the end, this entire effort went down the drain. ---- In a very similar way, I had to solve a problem created by using Google's Protobuf Python bindings which required generating Python modules in order to function. We needed to have an API server that could simultaneously serve multiple versions of the same Protobuf API from similarly named modules. Since Google's implementation didn't allow this, I wrote my own (while my manager was on maternity leave). I improved parsing speed, network load, reduced the amount of maintenance the component needed by making it possible to add new Protobuf message definitions at run time... The problem was I wrote the parser in C. This is what enabled good performance. When my manager came back to work, she realized she declared that she doesn't know C and will never learn (even though she wasn't related directly to the project), and the project was thrown to the dogs. ---- I have a similar story about extracting and aggregating Web interface from a RoR app, producing a Swagger definition... written in Prolog, which was also thrown away because Prolog. Similarly, had written an I/O tester for distributed filesystem in Prolog, which was thrown away for the same reason... And this answer will eventually hit the character limit if I keep listing things that were discarded simply because programmers didn't want to learn how to do their job.
- loondri 3y agoThis article talks about making Linux pipes faster, but other methods like shared memory or message queues might still be quicker. For example, in systems that need to move a lot of data quickly, the extra steps with pipes could slow things down. Also, when many threads are sharing data, pipes might cause more problems than other methods. So, the improvements in the article might not help much in real-world situations where speed is crucial.
- 127 3y agoAlso the benefit of using a message queue library is that you don't have to worry about multi-platform incompatibilities as much.
- 0d0a 3y agoCan you give some examples? When batching data, you benefit from picking something like io_uring. But for two-way communication, you still need to notify either side when data is ready (maybe you don't want to consume cpu just polling), and it isn't clear to me how those options handle that synchronization faster than pipes.
- _trackno5 3y agoThe main thing io_uring gives you is avoiding multiple syscalls. With a pipe you can’t really avoid that. With a shared memory queue/ring buffer you can write to the memory without any syscalls. But you need to build synchronisation yourself (e.g., using semaphores for example). You don’t necessarily need to poll.
- bakul 3y agoWhy not simply use mmap judiciously in a program managed shared memory ring buffer? Then you can copy at roughly memory speed.
- rostayob 3y agoThis post is an excuse to explain VM concepts, rather than a tutorial, something I maybe could have made clearer.
- qweqwe14 3y agoYou can probably still make it faster by avoiding libc and using syscalls directly. From looking at the final perf output, it looks like there's some overhead from using libc functions