5 ms·
IMO, not using any optimization flags with C is somewhat arbitrary, since the compiler writers could have just decided that by default we'll do thing X, Y, and
by anon946 3y ago
IMO, not using any optimization flags with C is somewhat arbitrary, since the compiler writers could have just decided that by default we'll do thing X, Y, and Z, and then you'd need to turn them off explicitly.
FWIW, without -O, with -O, and with -O4, I get 2500ms, 1500ms, and 550ms respectively. I didn't bother to look at the .S to see the code improvements. (Of course, I edited the code to output the results, otherwise, it just optimized out everything.)
- hyperbrainer 3y ago> (Of course, I edited the code to output the results, otherwise, it just optimized out everything.) O(1) :)
- caranea 3y agoThanks for posting your results! Since I was already set on writing in-browser particle life, I didn't benchmark C code with different flags.
- anon946 3y agoCompletely reasonable. I'd probably edit your blog post a bit to indicate that.
- TylerE 3y agoShould also test -Os when doing this sort of thing. Sometimes the reduced size greatly improves cache behavior, and even when not it's often outright competitive with at least -O2 anyway (usually compiles faster too!)
- abainbridge 3y agoOne optimization for the C code is to put "f" suffixes on the floating point constants. For example convert this line: t[i] += 0.02 * (float)j; to: t[i] += 0.02f * (float)j; I believe this helps because 0.02 is a double and doing double * float and then converting the result to float can produce a different answer to just doing float * float. The compiler has to do the slow version because that's what you asked for. Adding the -ffast-math switch appears to make no difference. I'm never sure what -ffast-math does exactly. Minimal case on Godbolt: https://godbolt.org/z/W18YsnMY5 https://godbolt.org/z/W18YsnMY5 - without the f https://godbolt.org/z/oc1s8WKeG https://godbolt.org/z/oc1s8WKeG - with the f
- a1369209993 3y ago> I believe this helps because 0.02 is a double and [...] can produce a different answer In principle, not quite. The real/unavoidable(-by-the-compiler) problem is that 0.02 is a not a diadic rational (not representable exactly as some integer over a power of two). So its representation (rounded to 52 bits) as a double is a different real number than its representation (rounded to 23 bits) as a float. (This is the same problem as rounding pi or e to a double/float, but people tend to forget that it applies to all diadic irrationals, not just regular irrationals.) If, instead of `0.02f` you replaced `0.02` with `(double)0.02f` or `0.015625`, the optimization should in theory still apply (although missed optimization complier bugs are of course possible).
- deleted 3y ago[deleted]
- abainbridge 3y agoOK, right, that's the clarity of thought I was missing. But in this case the compiler still misses the optimization with '(double)0.02f'. https://godbolt.org/z/az7819nKM https://godbolt.org/z/az7819nKM I think this is because the optimization isn't safe. I wrote a program to find a counter example to your claim that "the optimization should in theory still apply". It found one. Here's the code: #include <stdio.h> #include <stdlib.h> float mul_as_float(float t) { t += 0.02f * (float)17; return t; } float mul_as_double(float t) { t += (double)0.02f * (float)17; return t; } int main() { while (1) { unsigned r = rand(); float t = *((float*)&r); float result1 = mul_as_float(t); float result2 = mul_as_double(t); if (result1 != result2) { printf("Counter example when t is %f (0x%x)\n", t, *((unsigned*)&t)); printf("result1 is %f (0x%x)\n", result1, *((unsigned*)&result1)); printf("result2 is %f (0x%x)\n", result2, *((unsigned*)&result2)); return 0; } } } It outputs: Counter example when t is 0.000000 (0x3477d43f) result1 is 0.340000 (0x3eae1483) result2 is 0.340000 (0x3eae1482) What do you think?
- 3y ago
- deleted 3y ago[deleted]