4 ms·
I realize one needs a catchy title and some storytelling to get people to read a blog article, but for a summary of the main points: * This is not about a buil
by masto 3y ago
I realize one needs a catchy title and some storytelling to get people to read a blog article, but for a summary of the main points:
* This is not about a build step that makes the app perform better
* The app isn't 10x faster (or faster at all; it's the same binary)
* The author ran a benchmark two ways, one of which inadvertently included the time taken to generate sample input data, because it was coming from a pipe
* Generating the data before starting the program under test fixes the measurement
- meowface 3y agoAnother semi-summary of the core part of the article: >"echo '60016000526001601ff3' | xxd -r -p | zig build run -Doptimize=ReleaseFast" is much faster than "echo '60016000526001601ff3' | xxd -r -p | ./zig-out/bin/count-bytes" (compiling + running the program is faster than just running an already-compiled program) >When you execute the program directly, xxd and count-bytes start at the same time, so the pipe buffer is empty when count-bytes first tries to read from stdin, requiring it to wait until xxd fills it. But when you use zig build run, xxd gets a head start while the program is compiling, so by the time count-bytes reads from stdin, the pipe buffer has been filled. >Imagine a simple bash pipeline like the following: "./jobA | ./jobB". My mental model was that jobA would start and run to completion and then jobB would start with jobA’s output as its input. It turns out that all commands in a bash pipeline start at the same time.
- anonymous-panda 3y agoThat doesn’t make sense unless you have only 1 or 2 physical CPUs with contention. In a modern CPU the latter should be faster and I’m left unsatisfied by the correctness of the explanation. Am I just being thick or is there a more plausible explanation?
- DougBTX 3y agoIt depends on where the timing code is. If the timer starts after all the data has already been loaded, the time recorded will be lower (even if the total time for the whole process is higher).
- anonymous-panda 3y agoI’m not following how that would result in a 10x discrepancy. The amount of data we’re talking about here is laughably small (it’s like 32 bytes or something)
- deleted 3y ago[deleted]
- masklinn 3y ago> The amount of data we’re talking about here is laughably small So is the runtime.
- DougBTX 3y agoI’ll admit to not having looked at the details at all, but a possible explanation is that almost all the time is spent on inter process communication overhead, so if that also happens before the timer starts (eg, the data has been transferred, just waiting to be read from a local buffer) then the measured time will be significantly lower.
- masklinn 3y agoThe latter is faster in actual CPU time, however note that TFA the measurement only starts with the program, it does not start with the start of the pipeline. Because the compilation time overlaps with the pipes filling up, blocking on the pipe is mostly excluded from the measurement in the former case (by the time the program starts there’s enough data in the pipe that the program can slurp a bunch of it, especially reading it byte by byte), but included in the latter.
- anonymous-panda 3y agoMy hunch is that if you added the buffered reader and kept the original xxd in the pipe you’d see similar timings. The amount of input data is just laughably small here to result in a huge timing discrepancy. I wonder if there’s an added element where the constant syscalls are reading on a contended mutex and that contention disappears if you delay the start of the program.
- vlovich123 3y agoGood hunch. On my machine (13900k) & zig 0.11, the latest version of the code: > INFILE="$(mktemp)" && echo $INFILE && \ echo '60016000526001601ff3' | xxd -r -p > "${INFILE}" && \ zig build run -Doptimize=ReleaseFast < "${INFILE}" > execution time: 27.742µs vs > echo '60016000526001601ff3' | xxd -r -p | zig build run -Doptimize=ReleaseFast > execution time: 27.999µs The idea that the overlap of execution here by itself plays a role is nonsensical. The overlap of execution + reading a byte at a time causing kernel mutex contention seems like a more plausible explanation although I would expect someone better knowledgeable (& more motivated) about capturing kernel perf measurements to confirm. If this is the explanation, I'm kind of surprised that there isn't a lock-free path for pipes in the kernel.
- rofrol 3y agoThis @mtlynch
- vlovich123 3y ago
- karmakaze 3y agoI would definitely classify the title as clickbait because the app didn't go "10x faster".