4 ms·
Modern processors do not execute one instruction at a time, which is one of the reasons we rename registers (to recover the original name independent data flow)
by FullyFunctional 4y ago
Modern processors do not execute one instruction at a time, which is one of the reasons we rename registers (to recover the original name independent data flow). However a stack machine isn't a significant hindrance to renaming as long as the ISA is otherwise amendable to super-scalar fetch, decode, and rename.
The biggest reason we do not use stack machines today is that it's actually harder for optimizing compilers to generate good code for and it's awkward when joining control-flow from multiple blocks (stacks have to be in the same state).
However, amusingly, almost all modern processors have a small-ish control stack; it's called the RAS, the Return Address Stack predictor and allows the frontend to fetch through calls and returns way before the actual instructions are executed. There's definitely some redundancy here as everything is done again at execution time (and we check that the frontend got it right).
ADD: Knuth's MMIX and SPARC's register windows implement effectively a coarse stack, but both completely miss out on the instruction density that a true stack computer has.
- pjmlp 4y agoThey still make good targets for anyone getting their feet wet into compiler design and bytecode formats. With a good macro assembler it is super easy to translate those stack upcodes into good enough Assembly. It won't win any performance prices, but it will give a sense of acomplishment, and even provide a path for bootstraping, if wanted. Then the whole SSA generation, register colouring and what not, can come later if the author is then really interested into deep diving into compilers.
- akira2501 4y agoThe basic stack engine is also pretty neat, which tracks the value of the stack pointer separately and has it's own adder for it. If all you use is PUSH and POP you almost never have to wait for the value of the stack pointer to be available in the next instruction. It seems like it would be great for implementing a FORTH if you used the RSP pointer to hold the data stack, but you constantly have to PUSH and POP the return values out of the way, somewhat defeating the RAS. I wonder which is more of a win?
- FullyFunctional 4y agoThe RAS assumption is that your don’t override the return value, which we verify at execution. As long as you don’t, your code will run fast. If you do, then you’ll take some very expensive pipeline restarts. But you are right that x86 processors can (they all do?) track stack pointer movement in the frontend. For RISC you can track affine transformations of registers or even just the ABI designated stack pointer but it’s usually not worth the high cost.