3 ms·
It's always been a mix. Inlining, unrolling, and otherwise choosing longer sequences of instructions that are faster is still a major theme of optimization and
by throwawaylinux 4y ago
It's always been a mix. Inlining, unrolling, and otherwise choosing longer sequences of instructions that are faster is still a major theme of optimization and optimizing compilers.
Even on today's high performance CPUs, which are far more sophisticated than the primitive 5-stage scalar in-order R4300 (not sure it even had any branch prediction actually).
- kimixa 4y ago> Even on today's high performance CPUs, which are far more sophisticated than the primitive 5-stage scalar in-order R4300 (not sure it even had any branch prediction actually). I believe it didn't, with a short scaler pipeline (only 5 stages) the cost of waiting for a conditional branch to be calculated is relatively less, and MIPS of that era actually exposed the idea of a "delay slot" - rather than pausing the pipeline while calculating the jump, it still executed the next few instructions - so they were always executed no matter if the jump was actually taken or not. Filling this with NOPs makes it equivalently the same as a pipeline bubble, but I guess the intention was in some cases it could actually be used for something. I think a few RISC architectures of that era did similar things, their entire goal was simplicity after all, the argument being that something like a branch predictor and the logic required to reverse speculatively executed instructions "should" instead be used to make the CPU smaller or higher performing in the best case (either as implementing better performing but larger logic, or higher frequency due to shorter critical paths). I think this was dropped in later ISAs as it confused a lot of people, it wasn't well used by compilers (which is interesting as much of the MIPS ISA was explicitly designed around what would be easy for a compiler rather than handrolled asm), and improving process technology made transistors to handle the expensive superscaler speculative execution architectures relatively cheaper.