4 ms·
The current CPU's already are VLIW in a fashion. Its just that they have a hardware jitter for it. That's what superscalar OoO CPU basically is, a piece of HW t
by sharpneli 6y ago
The current CPU's already are VLIW in a fashion. Its just that they have a hardware jitter for it. That's what superscalar OoO CPU basically is, a piece of HW that generates VLIW instructions based on the normal instructions that come in. A superscalar execution port is basically just one subinstruction in VLIW. Explicit VLIW only helps you to save that piece of silicon from the chip. And it's not that much in the grand scheme of things. It used to be, but not anymore.
Static compile time VLIW means one cannot really make it wider anymore, or narrower. Dynamically doing it on runtime means a cheaper and smaller core can just be narrower. Remove an instruction slot for one integer ALU? Fine, no issues. Everything still works, albeit slower by that much. Make a beefier chip? Perfect, it got faster by the amount of instruction level parallelism that was available.
In addition a compiler cannot really see across function boundaries (except if it's statically determinable). A jitter can. Modern chips have reordering window of hundreds of instructions, M1 apparently goes up to over 600. That's quite a lot of stuff there that it can dynamically reorder across. A compiler might not have noticed that due to some weird dynamic call that was not visible during compilation time there is now an FPU instruction that could be inserted here, an OoO processor can.
Due to that explicit VLIW is basically dead outside of some highly specific applications, like DPS and whatnot.
In a similar fashion ridculously wide vector units were obsoleted by the approach pioneered by GPU's. Just add few things to allow masking based on branches and you get SIMT approach. Write as if it was scalar code and it'll run on HW that has vector lengths going from none at all up to whatever, NVidia has 32 wide vector units as an example.
- jleahy 6y agoEven more importantly, and I'm surprised you didn't mention it, static compile time VLIW doesn't know what to do for memory access. A load might take any number of cycles depending on whether the line is in L1, L2 or L3 cache, with static compile time VLIW then the compiler has to guess, if it's optimistic the whole CPU stalls, if it's pessimistic then you're much slower than you would otherwise be. I believe this is the real thing that make superscalar OoO (ie. JIT to VLIW in silicon) win.
- sharpneli 6y agoThat's a great point to emphasize. Because memory access times are inherently more or less nondeterministic. I consider it to be roughly as important as being able to reorder instructions across non statically determined branches and function calls. And both of these expose the fundamental weakness in explicit VLIW, if it's essentially nondeterministic it cannot be taken advantage of. Then we naturally have some other benefits, like hyperthreading. Which basically is just compiling two instruction streams together on the fly.
- mhh__ 6y agoThe performance on particularly memory-bound workloads is why I chose it as an example rather than a prediction. For it to work statically (although it really could be halfway in between), it would probably require a complete paradigm shift away from the current way we think about cpu caches Regardless of whether it'll work or not, I'll be very happy if the mill ever makes it onto a chip.
- alexvoda 6y agoAnd by jitter you mean JIT-er (Just-In-Time Compiler), not jitter as in https://en.wikipedia.org/wiki/Jitter https://en.wikipedia.org/wiki/Jitter Got really confused at first.
- ants_a 6y agoYou're using terms to mean things that they usually do not mean. Calling superscalar OoO scheduling a jit-ed VLIW is misleading at best. The whole point of VLIW is that you don't need the control logic and, more importantly these days, the power cost of scheduling each instruction by itself. From a single thread performance standpoint, the important part is that OoO scheduling is able to dynamically schedule around cache misses to effectively keep more memory accesses in flight, extract more memory level parallelism. In principle you could make an OoO VLIW CPU, but that would negate most of the power benefit while hamstringing the scheduler with unnecessary dependencies. Where in-order VLIW shines is when memory accesses are predictable, like DSP code. There you get an order of magnitude power efficiency gain. GPUs are effectively still in-order CPUs with large SIMD instructions, some useful instructions to make masked execution simpler and a specialized language and compiler to hide this model from the developers. GPU manufacturers calling these separate data lanes threads is just misleading marketing BS. There is no independent instruction pointer for each lane.
- sharpneli 6y agoThat's why I tried to use it more as a methaphor. Because the whole point of explicit VLIW (EPIC was what Itanium folks called it) was to save that scheduling HW. But nowadays that piece of HW is relatively minor part. So it's no longer worth to save it in a general purpose CPU. As this thread is about general purpose CPU's for direct consumer use (not a controller in hard drive or whatnot, but a full fledged CPU you run arbitrary programs in) we're talking about chips like Itanium when it comes to VLIW. I do not disagree that VLIW is a great for things like DSP where the power consumption is of the essence. One can get ridiculously high perf/watt by going explicit VLIW. I just don't see any way we'd see that approach in general purpose CPU's again. GPU's do not need to reorder instructions to hide memory latency. It just happens on a different level. While a single threadgroup (as in that single instruction pointer that controls the SIMD unit) will not get reordered at all, one has multiple threadgroups in flight. So if one group stalls at a memory load the unit will just schedule a different threadgroup. Because one generally has tons of them in flight. It's all about throughput. One could think of this as an in order CPU (from viewpoint of a single thread) but with ridiculous amounts of hyperthreading (one thread stalls, we can pick an instruction from another thread but never one from the stalled thread).