4 ms·
This is such an ignorant and low-effort post. The author did not bother to check whether the delay comes from the library loading time vs. the actual string fo
by matrix_overload 4y ago
This is such an ignorant and low-effort post.
The author did not bother to check whether the delay comes from the library loading time vs. the actual string formatting.
He is complaining about a single millisecond difference (that is near the measurement error) and haven't bothered to run the entire thing in a loop and get the difference to something meaningful (where the terminal performance can eat away any noticeable differences).
He picked a completely unrealistic metric (console output is usually meant to be read by humans, so it was never optimized for +- millisecond performance).
He doesn't distinguish between the language (C++) and the I/O API (printf() vs iostream).
The main advantage of C++ is that if you bother to understand how it works, you can let the compiler/debugger do a lot of mind-numbing and error-prone work that C developers proudly do by hand (collections, RAII, inline templates with concepts) at exactly 0.0 runtime overhead. The key is "if you bother to understand how it works". If you randomly pick an arbitrary API and complain how it's worse than a different arbitrary API from C for an arbitrary use case, it only shows your own incompetence.
- jenscow 4y agoThe post is "Hello world" is slower in C++, not "C++ is slower" or anything else. The speed of the Hello world application is the metric and they used hyperfine - a benchmarking tool.
- __ryan__ 4y agoFrom the article: > I do not believe that printing ‘hello world’ itself should be slower or faster in C++, at least not significantly. What we are testing by running these programs is the overhead due to the choice of programming language.
- MaulingMonkey 4y ago> The post is "Hello world" is slower in C++, not "C++ is slower" or anything else. The post pretty clearly blames "C++" if you read more than the title - note the "due to C++" here: >> Yet if these numbers are to be believed, there is a significant penalty due to C++ for tiny program executions, under Linux. > The speed of the Hello world application is the metric And it's a poor metric, mostly testing process setup/teardown. > and they used hyperfine - a benchmarking tool. And benchmarking tools are only as good as their usage and application. A more apples-to-apples comparison would be to use `\n` instead of `std::endl`, and `std::ios_sync_with_stdio(false)` to avoid excessive syncronization with C I/O. Admittedly, neither particularly helps here: I observe most of the overhead the article does merely by switching out clang/gcc for clang++/g++ on the same C source code, with ~6ms-9ms total runtime under wsl, or ~9ms (no measurable difference between the C++ or C code) when compiled with cl 19.15.26732.1 and run on windows. This is mostly measuring process setup/teardown and OS I/O overhead: redirecting stdout to null, I don't see a significant perf impact until I put the printing in some rather large loops: #include <stdio.h> #include <stdlib.h> int main() { for (int i=0; i<1000000; ++i) printf("hello world\n"); return EXIT_SUCCESS; } #include <iostream> #include <stdlib.h> int main() { std::ios::sync_with_stdio(false); for (int i=0; i<1000000; ++i) std::cout << "hello world\n"; return EXIT_SUCCESS; } Guess which is faster? The C++ version, oddly enough, both on wsl ubuntu linux: /mnt/c/local/ben$ clang++ -Os hello_world.cpp && cargo run --bin bench --quiet --release average runtime: 36.78434ms /mnt/c/local/ben$ clang -Os hello_world.c && cargo run --bin bench --quiet --release average runtime: 86.33667ms /mnt/c/local/ben$ clang++ -O3 hello_world.cpp && cargo run --bin bench --quiet --release average runtime: 36.34467ms /mnt/c/local/ben$ clang -O3 hello_world.c && cargo run --bin bench --quiet --release average runtime: 86.75787ms And on windows: C:\local\ben>cl /nologo /O2 hello_world.cpp && cargo run --bin bench --quiet --release hello_world.cpp C:\Program Files (x86)\Microsoft Visual Studio\2017\Community\VC\Tools\MSVC\14.15.26726\include\xlocale(319): warning C4530: C++ exception handler used, but unwind semantics are not enabled. Specify /EHsc average runtime: 1.44849203s C:\local\ben>cl /nologo /O2 hello_world.c && cargo run --bin bench --quiet --release hello_world.c average runtime: 1.75088561s Take these numbers with a massive grain of salt: Launching the first subprocess is extra expensive (on the order of 90-100ms) and included in the above averages, distoring them, presumably from delayed loading or initialization of a library in the benchmarking process as subsequent runs of the benchmarking process show the same overhead for the first subprocess[1]. Meanwhile, the author of the original article is calling overhead that has to be rounded up to 1ms a "huge penalty". (EDIT[1]: to rule out first-access overhead being from windows defender or similar, I confirmed that launching a sacrificial, unrelated, first subprocess such as "cmd /C ver" eliminates the overhead for the first "./a.out" execution)
- jenscow 4y agoThank you for taking the time. I know. However, you're testing the speed to print something, not the application. I mean, I'm assuming they wanted to include the process start and teardown. But sure, the test isn't fair.
- MaulingMonkey 4y ago> However, you're testing the speed to print something, not the application. I'm testing what impact modifications to the application have, to better understand what is and isn't contributing to the original application's performance and bottlenecks. I've shown how far I have to modify the original to actually start to hit I/O bottlenecks, which by extension, also shows how little they contribute to the bottlenecks of the original application, despite people looking to blame it for the "C" vs "C++" differences elsewhere in the discussion tree. > I'm assuming they wanted to include the process start and teardown. And there are interesting discussions to have about that in terms of breaking down OS overhead, API overhead, first launch vs subsequent launches, etc. - none of which the original post bothered with. We could peek at a more realistic workflow involving, say, an xargs-spammed executable processing a piped stream, if we really wanted to simulate a startup/teardown heavy workload. We could fire up `perf` to prepare to profile - I at least installed the frontend before realizing actually using it on wsl requires compiling stuff. The post didn't bother with any of that. And to be fair, I didn't bother too much either. Low-effort on my part as well ;)
- blastonico 4y agoThe author knows C++ (https://github.com/simdjson/simdjson https://github.com/simdjson/simdjson) and writes a lot about his performance experiments. I don't see why he - or anybody else - shouldn't raise such questions and arguments without having people (like you) getting angry about it. Is it offensive? Anyway, he has a point, `cout` is used extensively as a logging mechanism. If you don't see that "single millisecond" making any difference, you certainly haven't work on a relevant system.
- pjmlp 4y agoThen he should know his "C example" is valid C++ code.