3 ms·
>>All of which is to say, I think the 10x they're talking about is unrealistic. I think that is putting it mildly and is a little strange to somehow claim in t
by cpleppert 13y ago
>>All of which is to say, I think the 10x they're talking about is unrealistic.
I think that is putting it mildly and is a little strange to somehow claim in the first place. That can't possibly be true, their isn't room for a 10X increase in optimization on an Intel core chip and would be impossible to reach based on memory bandwidth and the amount of execution resources available on a chip alone. The ideas they have concretely put forth simply don't work or don't really provide a performance increase.
Take their virtual memoryless implementation. Getting rid of virtual memory doesn't buy you a whole lot especially when you need to add in a protection mechanism that looks a lot like a TLB in the first place(and must have the same general properties to provide protection, you just gain very marginal lookup costs).
If you do the math, this can't add more than 1-2% in performance in common application software at the cost of making every modern operating system unusable and increasing memory consumption(embedded systems anyone?). If getting rid of virtual memory was so great, why didn't someone do it in every other preceding clean room architecture? My answer: It isn't.
Or consider how they want to do branch prediction: add a separate ISA to do static branch prediction that is added by the compiler and loaded asynchronously by another cpu component and then supplied to the main cpu.
First of all, this doesn't work. The CPU can't have performance critical data pushed to it by another component. There is a reason the Branch prediction table and branch target buffers are small and focused and able to be accessed quickly. Secondly, static branch prediction is awful. You simply must be able to modify branch prediction data as the CPU executes to provide optimum performance. So it seems they want to be more power hungry, more complex, and have less performance than a mainstream CPU when it comes to branch prediction.
It is possible I have misinterpreted some elements of this scheme but basic design decisions like putting branch prediction into a separate component of the CPU simple don't make sense at all from a chip layout perspective.
Finally, I'm not really sure where the performance is supposed to come from with the 'belt' in the first place. Data dependency is incredibly complex in a modern pipelined cpu and while it is possible to reduce the cost by precompiling software for an optimized CPU the benefits are all very low level and really don't extend beyond reduced power consumption(assuming compilation cost can be amortized). At some point, to get more instruction level parallelism you simply have to bite the bullet and do dynamic out of order scheduling in the CPU to extract more performance. This has a well defined cost and an upper level limitation on how much total parallelism a CPU can extract from an instruction stream. Think of it another way: a static compiler has less information than a running CPU so one can't expect it to be able to extract more parallelism than the CPU itself.
- willvarfar 13y agoThe Instruction Encoding talk http://millcomputing.com/topic/instruction-encoding/ http://millcomputing.com/topic/instruction-encoding/ was the first talk, so explained these numbers in the first few slides. DSPs are massively faster than your Out-of-order Superscalar Monster, just ... not on general purpose code. The Mill is a DSP-like architecture with secret sauce so it can overcome the gotchas and go DSP-fast on general purpose code.
- cpleppert 13y agoRight, and I'm saying that the data dependency limitations alone in general purpose code pose an upper limit on how much instruction level parallelism can be extracted from the instruction stream. You simply can't make use of execution resources if they need to wait on the results of another one. EDIT: I'm not even considering pipelining, latency or transferring data among execution units, just assume every instruction completes in one cycle and makes its result available instantly.
- willvarfar 13y agoTrue, if you go hopping all over main memory, we'll go at main memory speed just like everybody else. Luckily there is a noticeable performance improvement between an i3 and an i7 precisely because normal app code doesn't go hopping long chains across main memory. This is the code we speed up. An order of magnitude improvement is saying that - hand waving - you can have a hot monster at 10x OoO SS performance or a cool little chip equiv to the OoO SS but low power. However, the faster you crunch the bits between main memory stalls, the more dominating those stalls become. Its diminishing returns. And hot is not good. So we talk more about sweet spots like 3x performance and 3x less power, and temper them appropriately. The numbers are based on sim and experienced estimates. Mill is faster because we can time x86 code and we can sim Mill code and we can compare them. Am typing on a phone, apologises if brief.
- synthos 13y agoI agree, you are always as slow as your critical path of executions. General purpose code, I imagine, often has very long critical path that no amount of parallelism will improve. What I want to know is just how silent the pipelines of this design would be under multi-threaded general purpose code.