6 ms·
RISC vs. CISC: What's the Difference?
- higherpurpose 11y agoIn other words, even if the x86 ISA itself is not bloated anymore, the CPUs can be. Because x86 CPUs still support a lot of 20-year old legacy stuff.
- SixSigma 11y agoHeadline : X found to be Y "X is Y" as an assertion "or that's what researchers claim in new report" as a caveat I hate this style
- struct 11y agoI don't have access to the actual paper, but looking at the linked results[0]: Core Name Performance (MIPS) Energy (J) Power (W) Cortex A8 178 25 0.8 Cortex A9 625 11 1.5 Atom N450 978 16 2.5 i7-2700 6089 28 25.5 So A9 delivers 625/1.5 = 417 MIPS per Watt, whereas the i7 delivers 6089/25.5 = 239 MIPS per Watt and the Atom delivers 391 MIPS per Watt. In addition, their spreadsheet has an "energy" tab calculated from a "normalized" power figure (where Atom comes out on top), but if you multiply the measured figures without the dubious adjustment, it seems that the A9 is actually more efficient (at least when you consider the board power), and MIPS is conspicuously absent from this spreadsheet. So the fundamental conclusion is "either ARM or Intel are better, but it depends on what you measure under what workload". [0] http://research.cs.wisc.edu/vertical/wiki/index.php/Isa-power-struggles/Isa-power-struggles http://research.cs.wisc.edu/vertical/wiki/index.php/Isa-powe...
- cliffbean 11y agoWhy are we measuring performance in MIPS?
- daemonwrangler 11y agoBecause it's easy to calculate? Too bad it's also utterly meaningless.
- Symmetry 11y agoI wouldn't read too much into the virtues of different ISAs from this comparison. The test processors were all built on different process nodes and even if the node is "32nm" that only means that the minimum feature size is 32nm, other sizing rules might be different and the drive current and leakage almost certainly will be.
- daemonwrangler 11y agoSomething else to keep in mind is that you can get significant power savings when you lower the clock rate. So if you measure total power consumed to run a calculation, it may actually be more efficient to run on a fast CPU, finish quickly, and then drop into a low power state than it would be to run it on a low performance CPU for significantly longer.
- dspillett 11y agoThis is the sort of factor that people forgot to include when testing SSDs for power/performance metrics in the early days of them being within reach of the average home users. An SSD (especially some of the older models) can pull more power than a good spinning metal drive when running as full force, but what some people didn't factor in was that the SDDs did more in a given time especially with latency sensitive workloads - so to do the same work as the traditional drive it would need to run at run pelt for far less time meaning quite a saving in power. Another thing modern CPUs do as well as slowing down when under light load is to almost turn parts of themselves off when not needed. These are things that any CPU could potentially do though, it isn't a difference between CISC and RISC designs.
- digi_owl 11y agoThat may work under a synthetic workload where you know the beginning and end of the "heavy" load. But i don't know if it holds up in real life scenarios, in particular on multitasking platforms.
- stephengillie 11y agoI can't find a good reference now, but supposedly the i7 has a set of transistors that calculates if its workload would execute faster on multiple cores, or fewer cores, and can park cores to save heat, and let the electricity be focused into the unparked cores. Intel's marketing material in 2008 mentioned the number of transistors doing the load calculations was about equal to the number of transistors in a 486. So you have a 486 constantly determining thread scheduling load, they claimed.
- wtallis 11y ago
- dspillett 11y agoAnother factor to consider is the rest of the chipset that goes with the CPU. Early Atom based networks and such always used the Atoms along with a chipset that under normal working conditions consumed nearly as much power as the CPU itself, more under certain loads.
- JoachimS 11y agoNone of the CPUs compared have very reduced (as in few) number of instructions. We've come quite far from MIPS1, IBM 801 and the first SPARCs in terms of ISA complexity. The big difference is really that x86 has an ISA->uop decoder, which basically is another decoder in front of the decoder in a RISC.
- Symmetry 11y agoPretty much every deep OoO design is going to have uops inside. Look at the A15 here, for instance: http://regmedia.co.uk/2011/10/20/arm_a15_pipeline_large.jpg http://regmedia.co.uk/2011/10/20/arm_a15_pipeline_large.jpg On ARM most of your splitting is going to be breaking out the predication so that the scheduler only has to work with 2 input uops. The thing about x86 is that the instruction stream isn't self-synchronizing, it's hard and occasionally impossible to figure out where the instruction boundaries are if you don't decode progressively from the front. This means that to decode mulitple instructions per clock cycle x86 has to use complicated voodoo.
- cbd1984 11y agoRISC wasn't about reducing the number of instructions, it was about reducing the instructions themselves, to make them simpler and faster to execute.
- daemonwrangler 11y agoAlso simpler to decode. If you look at how much chip area goes to the frontend decoder in an x86 chip, that's a significant difference.
- Someone 11y agoI don't think simpler to decode necessitates fewer instructions. If all your instructions put the bits describing which registers to use in the same place, use the same way to specify constants instead of registers, to specify that they operate on floating point variables, etc, using 8 bits that select the instruction gives you 256 different discussion with limited added complexity. Of course, it is unlikely that you can make 256 instructions use the exact same format (some operations will not have any use for 3 registers, for instance) but if you can keep things as consistent as possible, decoding becomes easier. Price paid is that you sacrifice instruction space, for example because the instruction format allows you to write the result of an operation to a register hard-wired to contain a zero. And x86 isn't that CISC-y. There were processors that basically could do a simple number to string conversion in one instruction, and there _are_ processors with an instruction that does Unicode conversions (http://publibz.boulder.ibm.com/cgi-bin/bookmgr_OS390/BOOKS/dz9zr003/7.5.49 http://publibz.boulder.ibm.com/cgi-bin/bookmgr_OS390/BOOKS/d...)
- Symmetry 11y agoARM up until the 64 bit transition was always one of the CISCiest RISC designs and x86 wasn't nearly as CISCy as, say, VAX. 64 bit ARM is a much more traditional RISC ISA than the previous encoding. But anyways, here's the link I always post when people talk about RISC versus CISC. http://userpages.umbc.edu/~vijay/mashey.on.risc.html http://userpages.umbc.edu/~vijay/mashey.on.risc.html
- ajross 11y agoAll production ARM hardware retains the old ISA. The 64 bit transition made it more CISCy, not less. Don't confuse an instruction architecture with a CPU.
- Symmetry 11y agoYes, essentially all ARM chips are going to support the old ISA. I'm not sure I understand your criticism though? RISC and CISC are terms that always apply to an ISA rather than a CPU, I hope I didn't accidentally imply anything else.
- ajross 11y agoIt's important because when people have silly internet wars over this stuff, they end up picking one side or the other because of its effect on actual hardware. So to say that AArch64 made things simpler is just wrong: all actual 64 bit ARM CPUs in fact implement a more complicated ISA using more die area and more power than their 32 bit predecessors.
- Symmetry 11y agoIn some sense it's more complicated because it's implementing more instruction sets but that's not as big a deal as you might think. ARMs have been doing that for a while, besides the normal A32 ISA there's also Thumb, Jazelle, etc. But while adding ISAs does take some die area it doesn't take too much since ARM instructions are very easy to decode. And that extra die area doesn't cost any power since when you're using A64 you can turn off the A32 decoder.
- m0skit0 11y agoIMHO author is actually missing the point of RISC architecture: instruction homogeneity allows for a simpler (and cheaper) hardware. Of course for software developers RISC or CISC, it actually doesn't matter, that's what abstract layers are all about.
- TheLoneWolfling 11y agoMy thoughts on the matter: Given that process sizes keep shrinking, and every time you shrink the process size you can fit more on the chip, and heat doesn't scale (as in, the smaller the process size the more heat per in^2), and we're up against a heat wall as it is, we're to the point now where a large chunk of the chip has to be dark at any point in time. As such, CISCs are looking better and better. Because you cannot really scale frequency more (due to heat concerns - freq^2 heat output, to a first approximation), and you have to run most of the chip dark at a time anyways, and you have the space, so you may as well have things that are optimized for rare use cases. And we're already seeing that. The micro-ops on modern x86 processors are getting more and more complex and specialized. This will especially start happening once we get decent CPU caches - the 3d-ish stacks that are being talked about. Where you have a separate chip stacked under or over the CPU that has a process optimized for RAM. Note that this is not talking about ISAs, this is talking about the processor itself. Although it's not done much currently, you can just as easily (or rather, with just as much effort) convert a RISC into CISC-like micro-ops (macro-ops?) as convert a CISC into RISC-like micro-ops. It's looking more and more as though ISAs can be successfully decoupled from the actual processor design. Which is encouraging. Treat the instruction encoding as effectively a compression scheme for the instructions that the actual processor runs.
- VLM 11y agoHundreds of MIPS is interesting, for a certain class of application, but it would be interesting to see the results for sub MIP applications. The microcontroller in a microwave oven.