5 ms·
When you look at the bottlenecks in opcode throughput for virtual machines it's almost always lousy branch prediction. Using a branch table instead of a switch
by jd 18y ago
When you look at the bottlenecks in opcode throughput for virtual machines it's almost always lousy branch prediction. Using a branch table instead of a switch block is not going to make much of a difference in any realistic program. The CPU has no idea which function is going to get called next so it can never fully use all the advanced look-ahead mechanisms.
For those interested, look at all the performance improvements made in the ocaml and perl and clisp interpreters. There is a lot of low hanging performance fruit left in the Python interpreter - the question is whether it's worth plucking. For instance, python can get a dedicated accumulator (I don't think it has one now) and you can introduce new opcodes that combine functionality of frequently occurring opcode sequences. For instance, python now has a LOAD_CONST and BINARY_ADD instruction, but not LOAD_COST&ADD instruction like clisp.
By combining the functionality of two instructions into one you're essentially saving an opcode dispatch. Saving a jump is going to make a bigger difference than making a jump faster. So that's where I think the python guys should focus on if they really want to improve performance.
- eru 18y agoSo you argue going from a RISC VM to a CISC VM? Interesting trend.
- sb 18y agojust for the record of translation into research literature for the interested: * dedicated accumulator: stack caching. * LOAD_CONST&BINARY_ADD: superinstructions. (though research literature suggests that dynamic superinstructions fare better w.r.t. performance than the static approach you suggest)
- kragen 18y agoLast I remember, this patch was purported to make a big difference precisely because it improves branch prediction, by giving the CPU a bunch of different possible sites to jump from.