4 ms·
A couple of points: - It's hard to reproduce the benchmarking results, as source code for the benchmarks is not provided. - The original bug report was for 32
by froydnj 16y ago
A couple of points:
- It's hard to reproduce the benchmarking results, as source code for the benchmarks is not provided.
- The original bug report was for 32-bit code, not 64-bit code, as the post assumed throughout.
- If you compile the code given in the original bug report as 64-bit code with and without -fomit-frame-pointer, there's no difference in the generated code.
- It's not clear to me that the "potential pieces of code" in the article are actually generatable with real-world C code. Again, not having actual source code available hurts.
- You shouldn't be using -fomit-frame-pointer on 64-bit code anyway, as you don't need the frame pointer for debugging/unwinding purposes on x86-64 like you do on x86. If the poster had read the x86-64 ABI, this would have been apparent.
- ice799 16y agoChill, son. 1.) I can provide the codez. I'll add a link to the article. 2.) Yeah. If you read the article, it reproduces on 64bit code, too. 3.) Not true. Try gcc (Debian 4.3.2-1.1) 4.3.2. The version I mentioned that I used in my post. 4.) etc 5.) Read the article. I'll add some more shit and reply to you again when its online. I need to eat breakfast and head to the office but I'll make you happy soon.
- masklinn 16y agoMaybe you should provide both graphs on a full scale (from 0 to 4.8) to show just how little, in your benchmark, -fomit-frame-pointer brings to the table in 32b (under half a percent using your mean cycles count) versus how much is lost due to it in 64b (nearly +30% cycles)
- ice799 16y agodone, refresh page and you should see em.
- herdrick 16y agoChill, son He was perfectly calm. No need for that. Let's keep things polite here.
- ice799 16y agobuilding the code in the bug report as a 64bit binary and various system information: http://gist.github.com/483494 http://gist.github.com/483494 and testing harness, scripts to build it, and to run it: http://gist.github.com/483524 http://gist.github.com/483524 -- yes i was too lazy to make a makefile. you will still need to construct some command line fu to separate the results into separate files so you can load it into whatever maths program you want.
- froydnj 16y agoThanks for the information. Your microbenchmark appears to be alignment sensitive. With your assembly code on my machine (quad-core 2.66 Core2 Quad) running for 250 tests I get: test1 usecs: avg 1.16759e+06 stddev 13917.6 test2 usecs: avg 1.31382e+06 stddev 405.725 which are similar results to yours. But if you add .align 8 right before the definition of test2 in the assembly file (i.e. make it be 8-byte aligned, just like test1), I get the following numbers: test1 usecs: avg 1.15972e+06 stddev 17004 test2 usecs: avg 1.1264e+06 stddev 754.44 so the code that "doesn't use frame pointers" is actually slightly faster, as you might expect. Additionally, if I simply modify your testcase to use 16-byte alignment, rather than 8-byte alignment, I get the following numbers: test1 usecs: avg 1.15895e+06 stddev 15764.7 test2 usecs: avg 1.12657e+06 stddev 941.606 I think aligning both test functions by 8 bytes at least makes things fair, but you can see that minor changes in alignment can cause big changes. You can see the assembly sources I used: http://gist.github.com/483840 http://gist.github.com/483840 FWIW, the code that uses movs rather than pushes and pops ought to be faster since (generally speaking for larger prologues and epilogues) you can execute a series of movs in parallel, whereas your pushes and pops are serialized, since they're all updating a common resource (the stack pointer). Empirical testing on benchmarks like SPEC2k has borne this out, both on x86 and x86-64. (You ought to be able to see this effect with gcc, depending on what cpu you use for the -mtune switch.) As you noted, this strategy carries a size penalty, since movs are somewhat larger than pushes and pops. I'll also note that on my machine, with gcc saying it's: @nightcrawler:~$ gcc --version gcc (Ubuntu 4.4.3-4ubuntu5) 4.4.3 I get identical assembly for compiling the testcase from the PR with and without -fomit-frame-pointer (I should have noted the gcc version I was using, just as you did. My bad.) Furthermore, for: @nightcrawler:~$ gcc-4.3 --version gcc-4.3 (Ubuntu 4.3.4-10ubuntu1) 4.3.4 I also get identical assembly. On one of the servers at work, with: @nightcrawler:~$ ssh henry7 gcc --version gcc (GCC) 4.2.4 (Ubuntu 4.2.4-1ubuntu4) I get identical assembly. Finally, also at work, with: @nightcrawler:~$ ssh henry7 /usr/local/tools/gcc-4.3.3/bin/i686-pc-linux-gnu-gcc --version i686-pc-linux-gnu-gcc (Sourcery G++ 4.3-83) 4.3.2 which is a somewhat patched version of GCC circa 4.3.2, I get identical assembly. So with four different flavors of GCC, there's no difference on the testcase in the PR with and without -fomit-frame-pointer. I'd be willing to bet that there's no differences with 4.5.x and mainline GCC as well. It looks like Debian may just have a peculiar set of patches to its version of GCC. EDIT: formatting fixes.