4 ms·
First, I would dispute that a cortex A9 is comparable to a Core 2. Second, if ARM is in fact getting that close to Intel, I would attribute it to diminishing re
by WallWextra 12y ago
First, I would dispute that a cortex A9 is comparable to a Core 2. Second, if ARM is in fact getting that close to Intel, I would attribute it to diminishing returns, and not some mythical property of the ISA that makes it difficult to implement.
I'm not sure what you mean by wide SIMD extensions providing "most of the chip's power." If you mean for vectorized DSP loops or whatever, sure, but I was talking about performance on branchy pointer-chasing spaghetti, which is what most real code is, and where it would be the hardest to catch up to Intel.
- hajile 12y agoThe particular benchmark in question was coremark which tries to test the performance of the core instead of the entire processor. While I agree completely that the uncore parts of the processor are very important, they aren't really a part of the RISC vs CISC discussion (Intel could swap out x86 and keep the uncore parts mostly unchanged). There are other sources comparing ARM and x86, but I leave that up to you. source http://www.vrworld.com/2011/02/21/why-nvidiae28099s-tegra-3-is-faster-than-a-core-2-duo-t7200/ http://www.vrworld.com/2011/02/21/why-nvidiae28099s-tegra-3-... source http://www.edn.com/electronics-blogs/systems-interface/4419918/2/Is-Intel-within-ARM-s-reach--Pedestrian-Detection-shows-the-way http://www.edn.com/electronics-blogs/systems-interface/44199... Programs like cinebench or linpack are used to to test top-end performance. A lot of things factor into these benchmarks, but SIMD, cache, and data throughput factor very heavily into them. This is why Sunway BlueLight supercomputer (a PetaFLOP computer) uses an Alpha design from around 1997 paired with a very large SIMD. The old core was "fast enough" to keep the SIMD units flowing. When it comes to branch prediction, the techniques are well-known. The 13-stage A8 supposedly has 95% branch prediction rates (http://www.ti.com.cn/cn/lit/wp/spry112a/spry112a.pdf http://www.ti.com.cn/cn/lit/wp/spry112a/spry112a.pdf), so I wouldn't say that is an issue (except to say that x86 decoding is much more complex and usually adds several extra stages to the pipeline which requires more sophisticated prediction hardware along with the associated costs). As far as prefetching go, the contract is basically that the programmer puts related data close together and the computer fetches the nearby data. The bigger the caches, the better chances of hitting. This is why lisp's pervasive use of linked lists can be problematic when compared to arrays (in theory) and why a naive b-tree implementation can be less performant than expected. Adding cache takes die and power. When Intel tries for low power, one of the first things it does is cut the cache size. IBM did something very interesting on this front when they used onboard eDRAM for cache (along with some fancy refresh logic to deal with persistence problems) because it was much more dense allowing more cache per chip (32MB L3 at 45nm).
- lispm 12y ago> This is why lisp's pervasive use of linked lists can be problematic when compared to arrays (in theory) The problem is less list vs. array in Lisp. Both tend to get allocated nearby or a copying/compacting GC will move them so. The problem is more that lists and arrays in Lisp usually point to data. For very few data types they will contain the data (fixnums, characters). Thus a list/array is often a list/array of pointers to data...