4 ms·
I just checked godbolt [0]. gcc only calls strlen once even with -O0. [0]:https://godbolt.org/z/j4o1915vE https://godbolt.org/z/j4o1915vE
by thethirdone 5y ago
I just checked godbolt [0]. gcc only calls strlen once even with -O0.
[0]:https://godbolt.org/z/j4o1915vE https://godbolt.org/z/j4o1915vE
- josefx 5y agoFor me the strlen call appears directly before loops backwards jump when set to -O0, resulting in a call every iteration as far as I can tell. However -O1 already seems to optimize it to a single call at the start of the function.
- still_grokking 5y agoI find it every time funny that when using languages that want to give you "total control" over the execution of your code (mostly C/C++) you actually almost never know what code gets executed in the end. It depends on the compiler, it's version, it's flags, and likely "the position of the moon". Of course the compiler is only allowed to do transformations that the spec permits. But it's impossible for a human being to anticipate the exact outcome. It's more like: "Compiler, do something that has the same outcome as this code I show you here". The output can be than something that doesn't resemble the input even slightly! There's obviously nothing wrong when the compiler is so smart that it sees some patterns and transforms your code into something much more efficient. Only that there's not much difference to what happens when you use a high level language. In both cases you in fact don't control the exact code that gets executed, and in both cases you rely on the smartness of your compiler to produce some efficient code, "whatever" you've written. That's why I think it's mostly a function of the code-style how performant or efficient some language can be (to some extend of course). When you write low-level style code (even in a high level language) a smart compiler will (hopefully) create something like what you would get form writing your code in C/C++.
- bombela 5y agoI think what matters is the intent. Here the intent is to compare against strlen at every turn of the loop. When the compiler is optimizing it might very well realize that the parameter to strlen doesn't change and the output can be saved first. If the intent is to compute the length only once, then we can save the length in a variable. Of course; as others have pointed out; there is no need to call strlen at all in this example.
- pinteresting 5y agoI guess that's true. Calling an expensive function effectively in the body of a loop like in this example is going to be an issue in every language. It is difficult to optimise because you need the compiler to evaluate and prove at compile time that both the loop cannot affect the result of the function call and the function call will not affect the loop. It's a common and easy optimisation to simply move function calls like this out of the loop. For the cost of 1 line of code I've regularly seen 10%, 100%, 1000% speed ups. It's actually one of the most common optimisations to do in non-compiled / "slow" languages if you know how functions are evaluated, you see that the cost of a simple getter function call can be the most expensive part of a loop.
- still_grokking 5y agoSure. But that's not really my point here. It's great and sometimes even astonishing what GCC and LLVM can do. Also it's clear that even the smartest compiler can't magically optimize any code. My point was more about the fact that compilers for lower level languages like C/C++, exactly the two named, use the most "magic" possible and that it's therefore almost impossible to anticipate upfront how their generated code will look like. But C/C++ claim that you have the most possible control over the code. My point was that this is only true to some extend, and that you can get almost equally good generated code using a less low level language just by writing code in a style matching the usual low level languages. (Especially than you need to think about loop invariants and such like you said)! So my point was more: The claim that you have "total control" over what happens at runtime when using a language like C/C++ is false. Seeing this example and at the same time people discussing pages long (while using even de-compilers) given that source snippet how the generated code may or may not look like reminded me of that, like I said "funny", fact about the "total control" C/C++ gives you. It's not an issue, of course. It's just an observation and I was reminded of it.
- pinteresting 5y agoIf you know what a compiler will optimise, might optimise and can't optimise, you can get an intuition for what the generated code will be executing, but even ASM is not "full control" and often not super useful because what you think is happening will shuffled around again by the CPU. The execution time for the same asm will vary significantly depending on architecture. If you have a good mental model of modern CPUs, in a simple loop, you can estimate what you think the bottlneck of the function will be, either by counting the micro ops or the number of stack / heap memory reads or memory allocations, etc, to estimate what is really happening to work out how you can optimise it, otherwise you're just shooting in the dark trying random combinations of flags or code not understanding why something worked or didn't work. At least in C/C++ that model works. In slow languages like python/javascript/etc, doing simple operations doesn't translate down to the very low levels at ALL. Generally if you imagine the worst possible way you can think of for how something will execute in a simple loop and multiply it by 10, it might be close.