3 ms·
VLIW in reality was actually worse (see Itanium as an example) than the currently favored superscalar out of order execution with register renaming. Both of the
by sharpneli 12y ago
VLIW in reality was actually worse (see Itanium as an example) than the currently favored superscalar out of order execution with register renaming. Both of these techniques try to take advantage of instruction level parallelism inherent in pretty much all code.
The reason for failure is that VLIW is basically stuck with whatever the compiler decides can be done. OoO does it dynamically. In the end the only real drawback with OoO vs VLIW is the chip area and increased power consumption required for the logic.
- tittat 12y agoYeah. One would wonder if we werent so stubborn with the way we write our programs what kind of performance we can get from our processors.
- CyberDildonics 12y agoWell, if programs are created to process simple instructions on lots of data there are huge speedups to be had on current x86 processors. By doing this memory latency is no longer a bottleneck and SIMD instructions can be used much more often. If unnecessary heap allocations have already been taken out, structuring a program like this can result in very substantial speedups.
- deleted 12y ago[deleted]
- graycat 12y agoYou are correct that VLIW, e.g., as in Itanium, flopped and out of order execution, speculative execution, branch prediction, etc. won out. The 9:1 speed up I reported was directly from Ebcioglu: For a while he was in our AI group at Watson. I don't recall just what he did, but he may have had an intermediate step that, for a given program, analyzed the stream of 370 instructions and rewrote them for the 24 way VLIW. An idea I had him consider was very long addresses. Or, who the heck really wants the addresses in main memory to be 0, 1, 2, ...? Of course, no one! That's why we have lots of work, from old link edit relocation to collection classes, memory management (garbage collection), etc. So, what do we actually do? Sure, darned near everything comes out of level 1, 2, 3 or so cache which works with just a hash of the main memory address. So, since we are going to hash the main memory addresses anyway, let's just do that! So, have a very long address, say, 1024 bytes long. So, the 1024 bytes are immediate from the source code! That is, just concatenate, say, address space name, program name, function name, class name, instance name, member name, key name (as in key-value pairs). That's the address. Then hash it. Ebcioglu checked how fast the execution logic could be, and it seemed okay. I know; I know; I left out a lot and need a better explanation! I'm concentrating on my software for my project and have gotten away from such low level hardware issues; heck, I don't even remember if we hash the real or the virtual address! But, the OP was a little light on a lot of history maybe relevant to the work of Soft Machines.
- ggreer 12y agoThose of you wanting to know more about this may be interested in Cliff Click's Crash Course in Modern Hardware.[1] It does a pretty good job of explaining how pipelined, superscalar, OoO CPUs came to be. 1. http://www.infoq.com/presentations/click-crash-course-modern-hardware http://www.infoq.com/presentations/click-crash-course-modern...