3 ms·
Right, and I'm saying that the data dependency limitations alone in general purpose code pose an upper limit on how much instruction level parallelism can be ex
by cpleppert 13y ago
Right, and I'm saying that the data dependency limitations alone in general purpose code pose an upper limit on how much instruction level parallelism can be extracted from the instruction stream. You simply can't make use of execution resources if they need to wait on the results of another one.
EDIT: I'm not even considering pipelining, latency or transferring data among execution units, just assume every instruction completes in one cycle and makes its result available instantly.
- willvarfar 13y agoTrue, if you go hopping all over main memory, we'll go at main memory speed just like everybody else. Luckily there is a noticeable performance improvement between an i3 and an i7 precisely because normal app code doesn't go hopping long chains across main memory. This is the code we speed up. An order of magnitude improvement is saying that - hand waving - you can have a hot monster at 10x OoO SS performance or a cool little chip equiv to the OoO SS but low power. However, the faster you crunch the bits between main memory stalls, the more dominating those stalls become. Its diminishing returns. And hot is not good. So we talk more about sweet spots like 3x performance and 3x less power, and temper them appropriately. The numbers are based on sim and experienced estimates. Mill is faster because we can time x86 code and we can sim Mill code and we can compare them. Am typing on a phone, apologises if brief.
- synthos 13y agoI agree, you are always as slow as your critical path of executions. General purpose code, I imagine, often has very long critical path that no amount of parallelism will improve. What I want to know is just how silent the pipelines of this design would be under multi-threaded general purpose code.
- axman6 13y agoI think that the advantage the mill gives here is more room to compute things speculatively on the critical path. You can happily load from an invalid address and then realise that was wrong later without causing an exception in the CPU; same with FP arith etc. This allows you to parallise more sequential code than say x86. I think the other related advantage here is that they've done as much as they can to remove any idea of core global state (comparison flags and so on) so that more operartions can be run in parallel; do several comparisona at once, and then process all the results together.
- igodard 13y agoYou then assume your own conclusion when you ignore pipelining. If instructions are executed in sequence indian-file then necessarily none will be faster than any other. The traditional rule-of-thumb is that programs have an ILP of two. The Execution talk (millcomputing.com/docs/execution) explains how the Mill turns that into an ILP of six. Then for the 80% or so of code that is in loops, pipelining has unbounded ILP - there will be a talk on pipelines upcoming.