4 ms·
I think the author is making a fairly critical mistake here. If I understand correctly, they missed the fact that the -O2 and -O3 results are several orders of
by topsycatt 5y ago
I think the author is making a fairly critical mistake here. If I understand correctly, they missed the fact that the -O2 and -O3 results are several orders of magnitude faster... Likely meaning the compiler optimized away the whole loop in those cases, rendering their results useless.
Edit: Yeah, looking at the godbolt link they provided, at least the execute_loop_basic function does essentially no work.
- innocenat 5y agoThis. For trivial loop, Both GCC and clang optimize away the entire BasicLoop, but clang failed to optimize away the entire DuffDevice, while GCC did. Thus, the only thing doing the work is clang's Duff Device. But even then, both GCC and clang optimize the BasicLoop function itself into just single add operation (data += loop_size) while it actually loops on the duff device function (though as mentioned, these functions weren't called from the benchmark)
- codebje 5y agoAre those results useless? The question posed early in the blog was whether Duff's Device is still relevant in 2021. The fact that the compiler can completely elide the loop without it but cannot with it is a pretty good indicator that it's likely to cause more harm than good. It might prevent inlining of a function call in the loop body, which would make a loop much, much slower to execute. It will make the loop body larger (even without inlining) and likely cause an increase in cache misses. It will be unpredictable based on build: an unrelated change somewhere else might misalign it with a cache line. Different CPUs might have different cache behaviours and perform differently. The right view on this sort of thing hasn't changed for decades: don't optimize prematurely. Trust your compiler, check it with profiling, look first for algorithmic improvements because you can gain FAR more converting an O(n^2) algorithm to an O(nlogn) than you can by unrolling a loop. Duff's Device is obsolete in all but a vanishingly tiny number of edge cases.
- topsycatt 5y agoYes, they still are useless, at least for the -O3 and -O4 sections. The compiler removing the loop is due to a bad test setup in which the function in the loop does nothing and compiler can figure that out and skip the whole calculation. In a real scenario the function in the loop would (presumably) actually do some work so it could not be elided. Also, note that the compiler does this for BOTH the Duff's device version and the non Duff's device version. Also, I'm neither attacking nor defending Duff's device. I'm just saying that we can't draw any useful conclusions from this article because of poor methodology.
- bluGill 5y agoThe thing is duffs device is about IO, so the loop cannot be optimized away as it isn't a pure function. The example code though is a pure function and so the compile is justified in optimizing it away. We have long known that in general compilers are better at optimization than humans so you should write readable code until a profiler shows otherwise. However there are cases where the compiler doesn't make an optimization. There are sometimes corner cases that the compiler can't figure out doesn't apply to you, so manually doing the optimization might work. These cases need to be re-evaluated with every CPU change, and every compiler upgrade though. Nothing above says if duff's device applies to not though. Duff's device is very likely to hit a corner cases where the compiler isn't sure if it can safely apply optimizations so it won't. As such I wouldn't be surprised if it is still relevant. (though mostly on embedded systems where we don't have nearly as fast of CPUs as more common computers)