5 ms·
Modern Microprocessors: A 90 Minute Guide
- TempleOSV2 13y agoKnow what's funny? My x86_64 compiler start RISC. I added addressing modes so that it used displacement off a register instead of an add instruction to get an address. Guess what? It didn't matter much. The CPU converts the CISC to RISC, so don't even bother in the compiler, except for dignity. God says... middle_class Ramsay fancy Dudly_Doright that's_for_me_to_know bad oil illogical abnormal here_now thats_laughable news_to_me doh study are_you_deaf youre_lucky what_a_mess hit You_da_man that's_your_opinion You_fix_it obviously phasors_on_stun horrendous a_likely_story chump_change horrendous bad look_buddy experts My compiler is 20,000 lines of code. I am not interested in adding 5,000 lines to get a 3% improvement. I do not support 32-bit floats, so SIMD is kinda wasted on just 2 64-bit floats.
- djcapelis 13y agoThis is a really good overview of modern processor design. I keep noticing that a lot of people don't seem to understand the CPU designs past the five stage pipeline taught in undergrad architecture courses and this document does a lot of good work in targeting the areas that folks are most likely to have misconceptions about. I do wish they covered trace caches, or as they're known partly in their more modern form, u-op (micro-op) caches, which are back in modern Intel chips again and cause some interesting performance artifacts. (The old trace caches of the P4 chips are different than the u-op caches of the new architectures, since the trace cache actually encoded branch predictions into the actual cache line lookup, which was always pretty wild.)
- solarexplorer 13y agoI agree, this is the best overview on the topic I have seen so far. Of curse, there is only so much you can squeeze into such a short overview. E.g. you probably don't really get VLIW if you don't know about trace scheduling. Or many cores without knowing about cache coherency. Etc. That's where you need to go on and read the references at the end.
- djcapelis 13y agoI wouldn't have minded if they cut out VLIW. That debate is mostly academic at this point and there's good reasons none of them have ever really been successful. If you're writing a guide to how modern chips actually work, it seems unnecessary to include a mostly dead path of work, but I've been hating on VLIW for a long time, so that's easy for me to say.
- solarexplorer 13y agoYes, VLIW is great for DSPs but not very much else. When the guide was written in 2001, Transmeta and Itanic were still around. So it made sense to include it back then.
- cma 13y agoGPUs are somewhat VLIW aren't they?
- zhemao 13y agoAs far as I know, they're SIMD, not VLIW. In VLIW, each ALU can be executing a different opcode. In SIMD, each ALU is executing the same opcode, but with a different operand.
- willvarfar 13y agoThe ATI cards were VLIW which gave them advantages in fixed pipeline, but as more and more programmatic shaders turned up ATI moved more towards CISC afaik.
- erichocean 13y agoGPUs are actually "single program, multiple data" (SPMD) machines, as far as programming is concerned. Internally, I believe they're implemented in a SIMD-like fashion, with extra hardware to handle the single program aspect within each lane.
- 13y ago
- chm 13y agoWhat reading do you recommend for someone who wants to learn about basic CPU design? I'm not a CS major, but I'm interested.
- sliverstorm 13y agoThe old standby, Computer Architecture: A Quantitative Approach, works you up through the basics.
- cfallin 13y agoPast the Patterson/Hennessy books, Shen/Lipasti's "Modern Processor Design" was really helpful to get a better sense of the implementation details of real chips, especially when you move past the 'conceptual' views (e.g. Tomasulo) and start to ask how instruction schedulers or renaming or load/store buffers actually work.
- throwaway_yy2Di 13y agoHold on, don't you mean Computer Organization and Design, by the same authors? The one with these pages: https://books.google.com/books?id=RXARim9cNBIC&pg=PA288&dq=basic+implementation+mips+subset https://books.google.com/books?id=RXARim9cNBIC&pg=PA288&dq=b...
- scott_s 13y agoThey're different books. "Computer Architecture: A Quantitative Approach" is an introduction to the subject for people who will work in the area. "Computer Organization and Design" is for people who need to understand how processors and hardware systems work in order to do their own work. (Mostly.) The preface to "Organization and Design" says basically this. For what it's worth, "Computer Architecture" is sitting on my shelf, and that's what I used in grad school. But based on their preface, I may buy "Organization and Design" because it may be a better reference for what I do day-to-day.
- vdm 13y agoComputer Systems: A Programmer's Perspective, Bryant & O'Halloran. http://www.amazon.com/dp/0136108040/ http://www.amazon.com/dp/0136108040/ http://csapp.cs.cmu.edu/public/pieces/preface.pdf http://csapp.cs.cmu.edu/public/pieces/preface.pdf This is a unique blend of operating systems and hardware architecture, emphasising application programming over the system implementation approach in Hennessy & Patterson.
- willvarfar 13y agoAn absolutely excellent article! For everyone interested in the topic, you might enjoy the new Mill CPU architecture talks http://ootbcomp.com/docs/ http://ootbcomp.com/docs/ - the very next talk is streamed live today (5th Feb, 16:15 PST http://ootbcomp.com/topic/instruction-execution-on-the-mill-cpu-talk-feb-5-2014/ http://ootbcomp.com/topic/instruction-execution-on-the-mill-... ) (I am a Mill forum mod; ask me anything about the Mill ;)
- fmstephe 13y agoWhat market is the mill aiming for? Servers, phones, desktop or all of the above?
- willvarfar 13y agoAll of the above! Its a family of processors, all compatible running the same binaries but each with the belt sizes, vector sizes, functional units and so on that suit it. So you'll get smaller Mills where that makes sense and absolute monsters in supercomputers, for example. When you think about the "smaller" Mill that you might have in your phone and tablet, though, its a monster compared to today's desktops! Except in the power efficiency department, that is ;)
- fmstephe 13y agoHave they indicated how they plan to get adoption?
- willvarfar 13y agoThe hackaday interview talks through some of the options on the business side: http://hackaday.com/2013/11/18/interview-new-mill-cpu-architecture-explanation-for-humans/ http://hackaday.com/2013/11/18/interview-new-mill-cpu-archit...
- Symmetry 13y agoProbably a smaller market segment than any of those, since network effects are important. Maybe render farms or networking or top end embedded or such.
- jokoon 13y agohope that makes some people want to learn about some very optimization basics.
- jwr 13y agoThis article is spectacularly good. I wish I had this available when I started doing assembly-level optimizations on x86 chips. This knowledge used to be much more fragmented and difficult to learn. VLIW could indeed be left out: you are not likely to encounter a VLIW chip, and if you do, it will come with an (excellent) compiler that will do most of the hard work for you. A good followup article would be a tutorial on how to lay out your structures/arrays in memory given your access patterns and cache architecture.
- dfox 13y agoIn my experience, when you encounter VLIW cores, there is almost no tooling for it as they tend be application specific DSP-ish cores. In that case hand optimized assembly is the way to go as there is no budget to produce optimizing compiler of anything.
- jwr 13y agoMy experience was with Texas Instruments C6000 DSPs. The compiler is excellent and if you use it well, you rarely have to resort to assembly. Even then, you normally write linear assembly, not parallel assembly, letting the assembler take care of the rest.
- pjmlp 13y agoVery nice written article. So how does ANSI/ISO C expose those details vs other languages, as many claim to?
- Symmetry 13y agoNo, and generally even the binary itself doesn't expose any of that, since you want the same program to run on your in-order atom and your OoO, multithreaded, Core iFoo.
- wtallis 13y agoThe big exceptions being SIMD and VLIW, two similar forms of explicit parallelism. They more or less require programming language support to use fully, and most older languages are purely scalar though some can be used in a manner that is fairly amenable to automatic vectorization.
- Symmetry 13y agoAt the ISA level you're right, those do require explicit instructions (or bundling of instructions). At the programming language level no. VLIWs can be programmed by normal compilers with well understood techniques (though variable memory latency makes static scheduling hard to do well in practice for most workloads). You just throw your C code into the compiler and it will spit out valid binaries and this has worked for as long as there have been VLIW machines. The techniques for auto-parallelizing 'for' loops by compilers into SIMD instructions are a more recent development, but they certainly exist. The Intel C Compiler is particularly good at that, but Clang and GCC can do this too.
- wtallis 13y agoI did mention automatic vectorization. I know it exists, but it is not perfect even in Intel's implementation. Compilers simply cannot be relied upon to always find and take advantage of opportunities for parallelism in serial code. Having language constructs that explicitly express that parallelism helps ensure that the compiler at least tries to generate parallel code, where given straight C it might just decide that the loop body is too big to bother with or that the side effects are too complicated to prove it can be parallelized.
- gtani 13y agohere's another good intro to CPU design, floating point math, linear algebra, PDE solvers etc http://www.tacc.utexas.edu/~eijkhout/Articles/EijkhoutIntroToHPC.pdf http://www.tacc.utexas.edu/~eijkhout/Articles/EijkhoutIntroT...
- kmitz 13y agoVery good overview, thanks a lot for this article