11 ms·
I've see this a lot, but IMO the truth is slightly different: the assumption behind EPIC was that a compiler _could_ do the scheduling which turned out to be _i
by FullyFunctional 4y ago
I've see this a lot, but IMO the truth is slightly different: the assumption behind EPIC was that a compiler _could_ do the scheduling which turned out to be _impossible_. The EPIC effort roots goes way back, but still I don't understand how they failed to foresee the ever growing tower of caches which unavoidably leads to a crazy wide latency range for loads (3-400+ cycles) which in turn is why we now have these very deep OoO machines. (Tachyum's Prodigy appears to be repeating the EPIC mistake with very limited but undisclosed reordering).
OoO EPIC has been suggested (I recall an old comp.arch posting by an Intel architect) but never got green-lit. I assume they had bet so much on compiler assumption that the complexity would have killed it.
It's really a shame because EPIC did get _some_ things right. The compiler absolutely can make the front-end life easier by making dependences more explicit (though I would do it differently) and by making control transfers much easier to deal with (the 128-bit block alone saves 4 bits in all BTB entries, etc). On the balance, IA-64 was a committee-designed train wreck, piling on way too much complexity, and failed both as a brainiac and speed-daemon.
Disclaimer: I have an Itanic space heater than I occasionally boot up for the chuckle - and then shuts down before the hearing damage gets permanent.
- klelatti 4y agoThis is really interesting. Any recommendations for further reading on EPIC and related technologies?
- FullyFunctional 4y agoIn practice EPIC = IA-64 and Itanium is the only implementation, but IA-64 is probably the easier thing to search for. The only book I have is “IA-64 and Elementary Functions: Speed and Precision”. EPIC’s problem is shared with VLIW of which EPIC can be understood as a refinement. VLIW excels in a deterministic world where the compiler can predict latencies and produce a good schedule, but falls apart in face of loads with highly variable latencies (an in-order implementation has no option but to stall when load data doesn’t arrive on time). EPIC patches this a bit by allowing software prefetching and SW exposed speculation, but it comes at a significant code bloat and it can cover only a small fraction of what dynamic scheduling can cover. At Transmeta, our VLIW engine did this too (with some obvious advantages over IA-64) and we suffered from similar problems.
- klelatti 4y agoThanks!
- FullyFunctional 4y agoIt such a fascinating topic and the story is far from over. VLIW (maybe not EPIC though) has tremendous power efficiency so sometimes it’s the right trade-off (say, if you can cover latencies by switching to another hyper-thread). Micro-architecture is still a hot topic and everything is a trade-off. Next I’ll start talking about super scalar OoO stack machines …
- fanf2 4y agoWere there any superscalar stack machines other than the inmos T9000 transputer? (It’s a bit of a cheat, though, because the transputer has a very limited stack, and the T9 worked more like a register machine, treating the very short local addressing mode as a register number)