5 ms·
Frustratingly, I just changed the macro to this: #define BINARY_OP(returnType, type1, type2, op) \ { \ register PTR pRet = pCurEvalStack - sizeof(type1)
by chrisb 15y ago
Frustratingly, I just changed the macro to this:
#define BINARY_OP(returnType, type1, type2, op) \
{ \
register PTR pRet = pCurEvalStack - sizeof(type1) - sizeof(type2); \
pCurEvalStack = pRet + sizeof(returnType); \
*(returnType*)pRet = *(type1*)pRet op *(type2*)(pRet + sizeof(type1)); \
}
And the assembly generated is this:
BINARY_OP(I32, I32, I32, +);
004117BA mov eax,dword ptr [pCurEvalStack]
004117BD sub eax,8
004117C0 mov dword ptr [pRet],eax
004117C6 mov ecx,dword ptr [pRet]
004117CC add ecx,4
004117CF mov dword ptr [pCurEvalStack],ecx
004117D2 mov edx,dword ptr [pRet]
004117D8 mov eax,dword ptr [edx]
004117DA mov ecx,dword ptr [pRet]
004117E0 add eax,dword ptr [ecx+4]
004117E3 mov edx,dword ptr [pRet]
004117E9 mov dword ptr [edx],eax
which is still unimaginably terrible. Why isn't it re-using values that have already been loaded into registers? Why isn't it using a register for pRet? I've even told it to! Although I think it's documented that the MS compiler ignores the 'register' keyword.
And this is with all optimisations turned on. How depressing.
- maximilianburke 15y agoMost compilers completely ignore "register" these days. I believe also that aliasing the stack when you're performing the actual operation is greatly hindering the compiler's ability to optimize.
- chrisb 15y agoIt does look as though no optimisation is being performed at all. I just isolated the use of the BINARY_OP macro, and put it in a simple-ish test function. now it's being optimized excellently. The function that contains the apparently unoptimisable code is in a hugely long and complex function, and I wonder if something in it is preventing all optimisation from occuring within that function. I've quickly looked through the assembly produced in the function and all of it appears unoptimised; whereas code in other functions is optimised ok. What can prevent all optimisation from occuring in a function?
- kenjackson 15y agoHow long is long? Is this a code gened function? I've seen in some compilers that they sometimes have limits where they turn off optimization due to throughput issues. If you could break the function up in to pieces, and see if it still doesn't optimize.
- chrisb 15y agoThe function is just over 2900 lines long. Every line crafted lovingly by hand. The whole source file is here: http://pastebin.com/9L8N3AVF http://pastebin.com/9L8N3AVF The function starts at line 232, and the disassembly I was looking at is from line 1852. This is the function that implements the direct-threaded interpreter. Direct-threading works by using goto's (jmp's) to dispatch the next instruction to be interpreted, which means that I don't think it can be broken up into multiple smaller functions. Please let me know if you think I'm wrong :)
- kenjackson 15y agoThat's the longest handwritten function I think I've ever seen :-) I don't know what limits the compiler imposes, but I wouldn't be surprised if you hit them.
- iam 15y agoThat's definitely the problem, compiler optimizations are easily O(N^2) on the # of operations in a function, so rather than taking forever to compile they'll turn off optimization if it takes too long. Ask yourself, do you even need a function that large? You might already be approaching the limits of the L1 cache, at which point you might as well be using separate functions since the call itself will be negligible to the cache miss.
- swolchok 15y agoWhat if you made BINARY_OP a __forceinline (http://msdn.microsoft.com/en-us/library/z8y1yy88.aspx http://msdn.microsoft.com/en-us/library/z8y1yy88.aspx) function instead of a macro? Things that the function contains that might disable optimization: a lot of gotos, inline assembly in GO_NEXT (this does affect optimization: http://msdn.microsoft.com/en-us/library/5hd5ywk0.aspx http://msdn.microsoft.com/en-us/library/5hd5ywk0.aspx)... It's also not immediately obvious to me that pCurEvalStack is initialized. Having a look at the architecture optimization manual re: the redundant loads; it's not immediately obvious to me that modern processors won't handle this fine (http://www.intel.com/Assets/PDF/manual/248966.pdf http://www.intel.com/Assets/PDF/manual/248966.pdf).
- kenjackson 15y agoSome of those loads are unexpected. Can you verify that /O2 optimization is on (since release builds can, in theory, have optimizations off).