4 ms·
Intel CPUs run one micro-op per execution port per cycle (see section 2.3.4 of http://www.intel.com/content/dam/www/public/us/en/documents/manuals/64-ia-32-arch
by panic 9y ago
Intel CPUs run one micro-op per execution port per cycle (see section 2.3.4 of http://www.intel.com/content/dam/www/public/us/en/documents/manuals/64-ia-32-architectures-optimization-manual.pdf http://www.intel.com/content/dam/www/public/us/en/documents/... for details, though the manual is a bit out of date -- more recent CPUs have more ports: http://www.anandtech.com/show/6355/intels-haswell-architecture/8 http://www.anandtech.com/show/6355/intels-haswell-architectu...). Optimized code has to be written with execution port utilization in mind.
- ant6n 9y agoI'm not disputing that ... I just thought that out-of-order execution was supposed to deal with that to an extend, with tens to even hundreds instruction in flight at a time.
- gpderetta 9y agoThere are so many execution ports compared to the throughput of decoding, renaming and retiring, that they are rarely the bottleneck. As ant6n correctly pointed, memory latency, branch prediction, dependency chain length, decoding and renaming are, in that order, normally the bottleneck even for hand optimized assembler code.