3 ms·
I think this was a great comment and discussion thread on OOO vs VLIW: https://news.ycombinator.com/item?id=24466593 https://news.ycombinator.com/item?id=244665
by celrod 6y ago
I think this was a great comment and discussion thread on OOO vs VLIW:
https://news.ycombinator.com/item?id=24466593 https://news.ycombinator.com/item?id=24466593
The biggest takeaway for me was the importance of scalability. As OOO CPUs get better, they can become increasingly parallel when running the exact same assembly.
While VLIW like Itanium was limited to 3 ops/instruction, which is already behind today's OOO CPUs. (Although Mill's 33 ops/instruction sounds like it'd take much longer to beat.)
My takeaway was that NVidea's Volta looked on the right track, where it added 6-bit dependency bitmasks to instructions to make building the dependency graphs (primary cost to OOO) more efficient.
- microtherion 6y agoIt seems to me that earlier RISC designs learned the same lesson with branch and load delay slots. They performed fine when introduced, but did not adapt well. All of this was tried to improve the speed of in-order CPUs. Once OOO became widespread, such hardware assists became pointless (And it would appear to me that the dependency assists you describe will similarly not stay relevant all that long).
- xscott 6y ago> While VLIW like Itanium was limited to 3 ops/instruction, which is already behind today's OOO CPUs Itanium could group 3 instructions per 128 bit "bundle", but multiple bundles could be run in parallel, and the compiler had to insert explicit "stops" when that wasn't allowed. The architecture was designed specifically so that future processors could run old code with more parallelism. https://en.wikipedia.org/wiki/Explicitly_parallel_instruction_computing#Moving_beyond_VLIW https://en.wikipedia.org/wiki/Explicitly_parallel_instructio...
- Quequau 6y agoThanks for the link. I actually saw that when it was posted but there's new comments since then.
- marcosdumay 6y ago> As OOO CPUs get better, they can become increasingly parallel when running the exact same assembly. While VLIW like Itanium was limited to 3 ops/instruction, which is already behind today's OOO CPUs. There's no reason for the VLIW machine not to get wider with time. The fact that Itanium didn't speaks more of its economic unhealthiness than of anything else.