5 ms·
Why was C slower though?
by hyperbrainer 3y ago
Why was C slower though?
- KeplerBoy 3y agoNo optimization flags could be a big part of the reason. Haven't looked closely at the code or tried it, but with -O3, -fopenmp and a well-placed pragma the performance could increase many-fold. Heck, with NVC++ you could offload that thing to a GPU with minimal effort and have it flying at the memory bandwidth limit.
- wzdd 3y agoSimply compiling it with -O3 produces something which completes in half the time of the JavaScript version (350ms for C, 750ms for JS), so perhaps that. Edit for Twirrim: on this system (Ryzen 7, gcc 11): "-O3": 350ms; "-O3 -march=native": 208ms; "-O2": 998ms; "-O2 -march=native": 1040ms. Edit 2: Interestingly, changing the C from float to double produces a 3.5x speedup, taking the time elapsed (with "-O3 -march=native") to 58ms, or about 12x faster than JS. This also makes what it's computing closer to the JavaScript version.
- anon946 3y agoDid you add output of the results to the C code?
- wzdd 3y agoI did not, but I confirmed with objdump that my compiler was not removing the code. (But to be sure, I just ran it again with an output and got the same value.)
- Twirrim 3y agoFor fun and frolics: No flags: 1843ms -march=native: 2183 ms -O2: 423 ms -O2 -march=native: 250 ms -O3: 425 ms -O3 -march=native: 255 ms O3 doesn't seem to be helping in my case.
- Tadpole9181 3y agoThey didn't use any optimization flags.
- Twirrim 3y agoYeah, I was just trying to show the difference. Doing it without optimisation flags is an utterly bewildering decision by the author.
- deleted 3y ago[deleted]
- nickpsecurity 3y agoEspecially given the goal was increasing performance.
- jasonjmcghee 3y agoSeems crazy to me that double would produce that kind of speed up. Is float getting emulated somehow? Don't they end up the same size?
- jlarocco 3y ago> Don't they end up the same size? No. Float is half the size of double.
- suby 3y agoMy intuition would have been that floats were faster because of this. Less memory to iterate through.
- fartsucker69 3y agoit is faster in just about every way. less memory, even the cpu instructions (which are usually not the problem) are faster. there's something fucky going on with code gen here. or it could also simply be the measurement procedure that is doing something weird like working with not properly cold or equally warmed up data or instruction caches.
- pvg 3y agoAuthor mentions they didn't use optimization flags but doesn't include the compilation details. You can sort of guess that (relatively) unoptimized C might perform worse than V8's JIT on short, straight computational code - you're more or less testing two native code generators doing a simple thing except one has more optimizations enabled and wins.
- hyperbrainer 3y agoOh, I assumed he still did -O2, and did not do anything else. Is that bad to assume? PS: I do not use C beyond reading some of its code for inspiration, so kinda unaware
- pvg 3y agoIt just says 'no optimization flags' and that's that, I don't think even the compiler is mentioned so the author is not giving you a lot to go on here - you don't even know what the default is. A modern C optimizing compiler is going to go absolutely HAM on this sort of code snippet if you tell it to and do next to nothing if you don't - that's, roughly, the explanation for this little discrepancy.