3 ms·
Replacing x86 with a RISC instruction set would do little good. Performance doesn't come from the RISCness of the decoded uops, it comes from superscalar out-of
by WallWextra 12y ago
Replacing x86 with a RISC instruction set would do little good. Performance doesn't come from the RISCness of the decoded uops, it comes from superscalar out-of-order execution, big caches, finely-tuned branch predictors, forwarding networks, a good memory controller, etc. That all takes up more transistors than x86 decode. POWER8 is the only RISC that is comparable to high-end Intel parts, and it is not a particularly small chip.
ARM used to have some ISA extensions for directly-executing JVM bytecodes, but it's deprecated. An optimizing JIT compiler gets much better results.
- hajile 12y agoCISC vs RISC is a little blurred. Most/all x86 instruction sets from SSE and onward are very RISC-like (and theses provide most of the chip's power). If we move back from these problem types to look at something like coremark, then MIPS has the highest coremark/MHz score. A more interesting question is that of the resources required to compete. In 2013, Intel spent 10 Billion in R&D while it's closest competitor (Qualcomm) spent 3 Billion (in fact, Intel spent more than the next four companies put together -- lest you argue about fabs, TSMC spent only 1.6 Billion). When companies using RISC are getting similar results for a fraction of the R&D, is x86 winning or just showing that with enough money, even a bad design can be made to work? ARM's A9 has similar performance to a Core2 T7200 (Tegra3 -- even when you account for the compiler the A9 is faster per clock at several things). A15 is around 60% faster than A9, ARM claims A72 is supposed to be 3.5x faster than A15 in 2016. ARM's 2014 R&D budget was around 400 Million (25x less than Intel). ARM is getting very close very fast on what is a shoestring budget in comparison. http://www.icinsights.com/news/bulletins/Top-10-Semiconductor-RD-Leaders-Ranked-For-2013-/ http://www.icinsights.com/news/bulletins/Top-10-Semiconducto...
- WallWextra 12y agoFirst, I would dispute that a cortex A9 is comparable to a Core 2. Second, if ARM is in fact getting that close to Intel, I would attribute it to diminishing returns, and not some mythical property of the ISA that makes it difficult to implement. I'm not sure what you mean by wide SIMD extensions providing "most of the chip's power." If you mean for vectorized DSP loops or whatever, sure, but I was talking about performance on branchy pointer-chasing spaghetti, which is what most real code is, and where it would be the hardest to catch up to Intel.
- hajile 12y agoThe particular benchmark in question was coremark which tries to test the performance of the core instead of the entire processor. While I agree completely that the uncore parts of the processor are very important, they aren't really a part of the RISC vs CISC discussion (Intel could swap out x86 and keep the uncore parts mostly unchanged). There are other sources comparing ARM and x86, but I leave that up to you. source http://www.vrworld.com/2011/02/21/why-nvidiae28099s-tegra-3-is-faster-than-a-core-2-duo-t7200/ http://www.vrworld.com/2011/02/21/why-nvidiae28099s-tegra-3-... source http://www.edn.com/electronics-blogs/systems-interface/4419918/2/Is-Intel-within-ARM-s-reach--Pedestrian-Detection-shows-the-way http://www.edn.com/electronics-blogs/systems-interface/44199... Programs like cinebench or linpack are used to to test top-end performance. A lot of things factor into these benchmarks, but SIMD, cache, and data throughput factor very heavily into them. This is why Sunway BlueLight supercomputer (a PetaFLOP computer) uses an Alpha design from around 1997 paired with a very large SIMD. The old core was "fast enough" to keep the SIMD units flowing. When it comes to branch prediction, the techniques are well-known. The 13-stage A8 supposedly has 95% branch prediction rates (http://www.ti.com.cn/cn/lit/wp/spry112a/spry112a.pdf http://www.ti.com.cn/cn/lit/wp/spry112a/spry112a.pdf), so I wouldn't say that is an issue (except to say that x86 decoding is much more complex and usually adds several extra stages to the pipeline which requires more sophisticated prediction hardware along with the associated costs). As far as prefetching go, the contract is basically that the programmer puts related data close together and the computer fetches the nearby data. The bigger the caches, the better chances of hitting. This is why lisp's pervasive use of linked lists can be problematic when compared to arrays (in theory) and why a naive b-tree implementation can be less performant than expected. Adding cache takes die and power. When Intel tries for low power, one of the first things it does is cut the cache size. IBM did something very interesting on this front when they used onboard eDRAM for cache (along with some fancy refresh logic to deal with persistence problems) because it was much more dense allowing more cache per chip (32MB L3 at 45nm).
- lispm 12y ago> This is why lisp's pervasive use of linked lists can be problematic when compared to arrays (in theory) The problem is less list vs. array in Lisp. Both tend to get allocated nearby or a copying/compacting GC will move them so. The problem is more that lists and arrays in Lisp usually point to data. For very few data types they will contain the data (fixnums, characters). Thus a list/array is often a list/array of pointers to data...