7 ms·
very good point on the "addresses compare == iff same object" rule. In that case though, I think clang is right to optimize the callee (but it does introduce a
by obl 8y ago
very good point on the "addresses compare == iff same object" rule.
In that case though, I think clang is right to optimize the callee (but it does introduce a problem in the caller) :
the only place you could do the equality check and observe the rule being broken is before the callee returns since the lifetime of its variable is bound to the call.
It seems that clang will not let the return pointer alias a local in the caller except when the call is the initialization of said local.
So if the caller goes :
foo x; leak(&x); x = returns_foo();
the memory will be temporary stack (and then memcpy), thus upholding the rule. (and it seems to me that this inefficiency is really required to respect the standard if we actually leak the pointer)
in the case :
foo x = returns_foo();
clang will pass the actual address of x down but that's before the object exists (and its address cannot be known yet) so the rule is still fine.
I stand corrected though, this does mean that RVO would be useful for C as well, as a way to relax the aliasing rule.
edit: nevermind that, in the first case it's perfectly legal to read/write foo through the pointer downstream so you cannot make the optimization anyway.
- josephg 8y agoThat seems to match the logic clang-trunk is using: If the assignment is in the initializer then its using the simple C++ RVO-style call to blah(): void zot() { foo f = blah(); } zot: push rbp mov rbp, rsp sub rsp, 1040 lea rdi, [rbp - 1040] call blah add rsp, 1040 pop rbp ret But if the variable has the opportunity to leak then it adds a memcpy in the caller: void zot() { foo f; leak(&f); f = blah(); } zot: push rbp mov rbp, rsp sub rsp, 2080 lea rdi, [rbp - 1040] call leak lea rdi, [rbp - 2080] call blah mov eax, 1036 mov edx, eax lea rdi, [rbp - 1040] lea rcx, [rbp - 2080] mov rsi, rcx call memcpy add rsp, 2080 pop rbp ret Although weirdly the memcpy is removed when optimizations are turned on. This may be a bug.
- BeeOnRope 8y agoYes, the same thought occurred to me (that perhaps clang is careful in the caller in the case the address escapes), but I seemed to find cases where clang optimizes the caller also, so that two distinct objects receive the same pointer and both pointers escape. Here's an example: https://godbolt.org/g/yxFzqT https://godbolt.org/g/yxFzqT This happens on clang versions back 3.6. Note that if you change the caller to: Foo f; f = callee(&f); the code changes and distinct objects are passed. I'm not sure if the first form (all in the definition) has a relevant difference per the standard that lets clang do this.