4 ms·
Here is more information from official site [1,4,5]: - native "Elbrus" ISA or x86 ISA, - Ebrus ISA is VLIW, can dispatch 23 operations per cycle (33 with SIMD
by spatular 11y ago
Here is more information from official site [1,4,5]:
- native "Elbrus" ISA or x86 ISA,
- Ebrus ISA is VLIW, can dispatch 23 operations per cycle (33 with SIMD), in-order execution,
- it's stated that x86 code translation + register allocation is done in HW, but later they write about a software translator and full-system emulator,
- 6 ALUs (all support integer operations, 4 can do FP),
- 256 x 84-bit register file,
- hardware support for loops, including pipelining,
- some kind of module for async mem preloading,
- speculative execution and branching predicates,
- "4S" model has 4 cores,
- 800 Mhz core clock,
- 64 KB L1, 128 KB L2, 8 MB L3 (shared between cores),
- 3 DDR3-1600 interfaces, ECC support,
- 3 x 12GBytes/s inter-CPU links, support for up to 4 sockets,
- 65nm process, 380 mm^2 die size, 986e6 transistors,
- software is based on Linux 2.6.33 and Debian 5.0 with more than 3000 packages.
There are some benchmarks for older chip model "2S" (overclocked to 500MHz, 2 cores) [2,3]. FP performance is about 1-5x of Pentium M 1GHz (1 core?) depending on benchmark, integer performance is about 1x. New CPU, "4S", should be 3 times faster than "2S".
[1] http://www.elbrus.ru/arhitektura_elbrus http://www.elbrus.ru/arhitektura_elbrus
[2] http://www.elbrus.ru/files/535269/9f0cd8/50606f/000000/2014-04-19_161108.png http://www.elbrus.ru/files/535269/9f0cd8/50606f/000000/2014-...
[3] http://www.elbrus.ru/files/535269/0e0cd8/50586f/000000/2014-04-19_161043.png http://www.elbrus.ru/files/535269/0e0cd8/50586f/000000/2014-...
[4] http://www.mcst.ru/mikroprocessor-elbrus4s http://www.mcst.ru/mikroprocessor-elbrus4s
[5] http://www.mcst.ru/mikroprocessor-elbrus4s-gotov-k-serijnomu-proizvodtstvu http://www.mcst.ru/mikroprocessor-elbrus4s-gotov-k-serijnomu...
--
Edit: loop pipeling, OS information
- agumonkey 11y agoHow many processors have hardware support for loops ? I expect it to be a different, more efficient infrastructure than Comparison/Jump, maybe something similar to DisplayLists in old OpenGL ?
- spatular 11y agoWell, full translation would be "hardware support for loops, including pipelining". How exactly is pipelining implemented is not clear.
- Zardoz84 11y agox86 LOOP
- revelation 11y agoOr just REP.
- ddingus 11y agoWhich CPU is this for?
- Sanddancer 11y agox86/x64 -- http://web.itu.edu.tr/kesgin/mul06/intel/instr/rep.html http://web.itu.edu.tr/kesgin/mul06/intel/instr/rep.html
- ddingus 11y agoThanks!
- kazinator 11y agoDBcc instructions in Motorola 68000.
- rzzzt 11y agoIntel's Itanium line had the following hardware support for loops: * Register rotation [1]: referenced registers inside the loop are cycled on each iteration * Branch prediction [2]: with a dedicated loop count register, the CPU can tell with complete certainty when not to take the jump to the beginning (no mispredictions at the exit) [1] http://www.cs.nmsu.edu/~rvinyard/itanium/register_rotation.htm http://www.cs.nmsu.edu/~rvinyard/itanium/register_rotation.h... [2] http://www.cs.nmsu.edu/~rvinyard/itanium/branching.htm http://www.cs.nmsu.edu/~rvinyard/itanium/branching.htm
- petjuh 11y agoBranch prediction is already extremely efficient for loops, it simply assumes the last comparison will be the same. For loops this is true for all but the first and last iteration.
- SamReidHughes 11y agoEven if it's perfect it still takes up branch prediction overhead in the processor, and hurts the performance of other branches. Also, getting the first or last iteration wrong is pretty bad for a ton of for loops out there.
- jevinskie 11y agoQualcomm Hexagon DSPs have "instruction packets" that have hardware loop support. http://en.wikipedia.org/wiki/Qualcomm_Hexagon#Code_sample http://en.wikipedia.org/wiki/Qualcomm_Hexagon#Code_sample
- userbinator 11y agoThe most interesting thing about the benchmarks is that they show some pretty amazing IPC. The P6-based Pentium M was known for its high IPC, but this 500MHz Elbrus core is more than 50% of the speed of a 1GHz Pentium M. The floating-point results are even better, although perhaps not surprising due to having 4 FP ALUs. Another set of benchmarks containing an Elbrus is at http://www.7-cpu.com/ http://www.7-cpu.com/ and also shows extremely good IPC efficency - it achieves 1.2 MIPS/MHz/core for compression, which is better than all the other non-x86 in that list, and somewhere between Haswell's 1.18 and Ivy Bridge's 1.24. If they could scale this up to a newer process, they'd probably be equal to if not surpassing Intel's current x86 performance.
- microarchitect 11y agoHow good is the compiler? To paraphrase David Patterson, its easy to get really great IPC for code from a shitty compiler, but that doesn't mean your program will run faster.
- Everlag 11y agoBeing limited to in-order execution, won't the overall IPC be limited by Amdahl's law as a result of issues such as pipeline stalls? Out-of-order will take as much space and power as you can throw at it but there's a reason we aren't using 'speed demon' processors for general computation.
- spatular 11y agoIt's quite possible that in SPEC benchmarks they used 2 Elbrus cores and 1 Pentium M core, so IPCs are about the same. And it's confirmed by 7zip benchmark. In theory Elbrus can dispatch 23 non-SIMD ops per cycle, it has only 6 integer units and 4 among them can do FP, so it's not clear what remaining 17 operations can do. Modern x86 CPUs have about the same number of integer/FP ALUs. Current IPC is quite good, but it's hard to tell if core clock can be increased without lowering IPC. If ALU pipelines are short, it may be difficult to increase frequency without adding more stages. If more stages are added then pipeline flushes due to unpredicted jumps will result in heavier penalty. And it looks like Elbrus doesn't have a branch predictor at all. Pipeline stalls caused by memory access will also increase with CPU frequency. In x86 CPUs they are somewhat masked by out-of-order execution and SMT, which Elbrus doesn't have, so it should depend on memory preload instructions inserted by compiler.