6 ms·
This is drastically over-simplified, but in short, the Mill is a sort of middle-ground between superscalar implementations (x86, ARM) and VLIW (Itanium, DSPs).
by peller 10y ago
This is drastically over-simplified, but in short, the Mill is a sort of middle-ground between superscalar implementations (x86, ARM) and VLIW (Itanium, DSPs).
Superscalars leave a lot of the optimization process to complexity in the hardware. This is seen in stuff like the out-of-order scheduling and large cache hierarchies. In one of the Mill talks, it's guesstimated that almost 90% of the circuitry is dedicated to simply moving data around, as opposed to actually doing any work on that data.
VLIWs take the opposite extreme, leaving the complexity of optimization to the compiler. History has shown so far though, that many computing problems are (again oversimplified) too conditionals-heavy to really benefit from VLIW, and end up running much more slowly.
So the Mill is a bet that they can get more benefits from each approach without the major draw-backs of each. This isn't really ground Intel could simply "move into" without half a decade plus of work, and even then, they'd be cannibalizing their x86 ecosystem, which is not a risk most entrenched corporations are fond of.
- gpderetta 10y agoThe biggest issue with WLIW is the impredictability of the latency of memory loads which make any sort of static scheduling very hard for general purpose workloads. Itanium it seems had similar issues even with specialised hardware to help the compiler. Not sure what's the memory latency hiding story for the Mill. Edit: Symmetry said it better.
- MertsA 10y agoOn the Mill, you tell it the address that you want to read and when you want to read it. So you could issue a load for XYZ as soon as you know what XYZ is and if there's a subsequent store to XYZ that will update the result so when the load returns you have the data that was there when the load returns and not from when the load was issued. If you issue a load that completes in 5 cycles and it needs to go to RAM for that data then it's going to stall waiting for that data but because you can issue loads sooner than on other architectures you can sort of alleviate some of the problems with stalling on some load.
- gpderetta 10y agoBut as far as I understand and hear, Itanium had sort-of similar capabilities (ALAT and speculative loads) and it didn't work great outside of FP workloads as it is still hard to programmatically schedule load early even if you can ignore RAW hazards. What's unique in the Mill?
- Symmetry 10y agoBasically they've used their exposed pipeline design and register metadata to fix the things that didn't work well about ALAT. Or at least that's the theory. EDIT: I think the main problem was that on Itanium issuing a speculative load could potentially trigger a page fault making it potentially dangerous. On the Mill the page fault won't trigger until the load leads to a side effect outside the belt, sort of like it's been wrapped in a Haskell Maybe monad. So if you' have something like for (int i = 0; i < foo; i++) { a[i] = a[i]+1; } you can speculatively load the a[i+n] while you're working on a[i] even if allocated memory stops at the border of the array allocation.
- AlphaSite 10y agoMy thought process of this has been that Superscalar is like a JIT, in that it attempts to optimise these instruction scheduling as it happens, always executing when viable. Bit obviously although this is good for optimising execution does come at a fairly high implementation complexity cost. Whereas VLIW is where you optimise ahead of time with no fore knowledge, if the compiler is good enough, you get the same performance with simpler, cheaper, more efficient hardware (or potentially more performance) but often this doesn't pan out fully.
- XorNot 10y agoBut that's the point: if I want to buy non-Intel at my company, I've got contracts, lawyers fees, upper-level management etc. to convince to do it. Intel is probably sending me support personnel and sample hardware. And I'm still looking at recompiling, debugging and deploying a ton of software to see advantages from the Mill. So whatever advantages it brings, they need to be very substantial (i.e. if the next-gen of Intel x86 chips still out-perform it, I'll buy them) and quite quick - because I can go 5 years not buying the Mill while Intel promises to support me for their awesome new architecture. I mean, probably we buy some Mill machines, but how likely is it that it's game changing on my codebase? That's where I see the problem. There's this whole huge assumption that the Mill will yield a bunch of benefits. If they're clear cut (a huge if) then they still have to beat their competitors being able to brute-force performance improvements until they come up with a new architecture themselves, which can take advantage of the very compensation they're asking from their customers ("switch architectures, it'll be great we promise").
- peller 10y agoVery good points. I agree that many established companies are going to be very wary of making the jump. My hunch is that if they have a real shot, and this is assuming that most of their promises turn out, it's not going to be going after general purpose computing head-on. For instance are there big problems out there that are compute intensive, but branch-heavy enough that GPUs aren't a great fit, in applications that need low power but that don't make sense as "cloud services"? I don't know. The optimist in me sees potential opportunity in opening up new domains. The practicalist agrees with you; it's going to be an uphill battle and a lot of stars need to align just right. But if they really can deliver 10x on general purpose computing workloads, it's hard for me to see that not being game changing.
- Symmetry 10y agoThere are a few places I could see Mill making inroads. The security features could be very useful for the big internet companies, Amazon, Google, Facebook, etc and they can afford to spend a few tens of millions of dollars on something speculative like this. And they do, with things like Arm or OpenPower servers. There's also the high end embedded land where you don't need to run a traditional operating system. Network switches, cell towers, that sort of thing.
- Symmetry 10y agoMore accurate to say, I think, that the Mill is a VLIW with certain hardware facilities to overcome the traditional weaknesses of those - code density and variable memory latency. The solution to conditional branches is exactly the same for VLIW as for Superscalar machines: use a branch predictor. And VLIWs can tolerate a slightly worse branch predictor since they tend to have shorter pipelines. EDIT: To explain a bit more, VLIW has historically worked great in cases like DSP workloads where the memory access patterns are very predictable and you aren't unexpectedly loading things from lower level caches very often. There's anther thread on HR right now about doing deep learning with a Hexagon, which is a sort of VLIW, and it works very well. But as soon as you miss L1 in a VLIW the whole thing comes to a stop, whereas in an OoO (Out of Order) processor you can keep executing subsequent instructions that don't depend on that load and so you don't have to stall. Basically every instruction that isn't a load has deterministic or at least hard bounded latency that the compiler can easily plan for. The other big disadvantage of VLIW is that sometimes you have stretches of code where you can only have one useful instruction at a time. VLIWs often use very RISCy encoding formats that still take up lots of space per bundle in these stretches, leading to potentially very low code density. The Mill gets around with a very CISCy encoding that only takes a small amount of I-cache size for single instructions.