5 ms·
Intel Launches Next Gen Itanium Monster Processor
- rbanffy 16y agoWhat current OS options exist for Itanium processors? I would count HP-UX and Linux and, of course, NetBSD. Microsoft already stated 2008 R2 will be the last OS they make for ia64. That said, looks like an impressive processor.
- anonymous246 16y agoLast sentence of the article has the phrase "sufficiently intelligent compilers". :) Intentional in-joke? Wikipedia's definition: "Sufficiently Smart Compiler, any of a family of theoretically possible compilers able to perform sophisticated but unrealistic code optimizations" > Given sufficiently intelligent compilers, Itanium could begin to make economic sense in fields that couldn't previously justify the high cost of optimizing for the chip. https://duckduckgo.com/?q=%22sufficiently+smart+compiler%22 https://duckduckgo.com/?q=%22sufficiently+smart+compiler%22
- kd0amg 16y agoI would probably read it differently depending on the background of the person who wrote it (mostly based on how aware I'd expect the writer to be of the problems involved). I've seen people with a basic understanding of compilers stumble over it and not know why others in the room chuckled. In this particular case, I don't know enough about the author to say, but EPIC/VLIW architectures are kind of known for making things difficult for the compiler (meaning the joke would be very appropriate here).
- jacques_chester 16y agoThe entire bet for EPIC was that a sufficiently smart compiler would mean you could free up die space for processing transistors by ditching branch detection, prefetch logic, speculative execution, caches etc. As you point out, the SSC has yet to appear. Just look at that layout: it's dominated by cache.
- scott_s 16y agoI think even in the ideal SSC case, you'd still want as much cache as you can get. Even if the compiler can insert perfect prefetching instructions, the prefetched data has to go somewhere. And the more somewhere you have, the more aggressively you can prefetch. I think the main benefit would be much simpler instruction pipelines, which would include the points you mentioned (branch prediction, prefetch) but also all of the logic needed to keep track of dependencies in an out-of-order processor.
- jacques_chester 16y agoAbsolutely. I studied the Itanium design philosophy back in 2000 and this is exactly what they were aiming to do: drop all the complex logic devoted to keeping the pipelines full and all the units busy. True about data, though I vaguely recall EPIC had advantages there too because without needing to do branch prediction, you didn't need to speculatively fetch multiple memory addresses; meaning the same D-cache went further.
- kenjackson 16y agoThe IDC prediction chart is a thing of beauty. I'd love to have a webpage just for seeing data like this (predictions vs reality). Must be one of the best charts I've seen in a while.
- JoeAltmaier 16y agoSo, is THIS Itanium going to sell? Intel sure has a boatload of patience.
- psykotic 16y agoWhen the Itanium first launched, we got some machines from Intel so we could port our software. After excitedly unpacking one machine, we plugged it into the office power outlet and flipped the switch. Suddenly all the lights on our floor went out. Turns out you were supposed to only feed it from a data-center-strength power grid.
- kenjackson 16y agoThe nice thing though is once it was installed, you save money because you can turn off the heat in your building.
- joshu 16y agoI didn't realize Intel still made this stuff.
- Andys 16y agoYou can thank ongoing long-term enterprise and government contracts for that, to the tune of over a billion a year.
- aidenn0 16y ago1) The article says "thread level parallelism" when they mean "instruction level parallelism" 2) There is way more ILP available at run-time than compile time, and what ILP is available in both is much more tractable at run-time. An out-of-order CPU is constantly filling a buffer with instructions (or microcode), and hardware is determining the dependencies dynamically to issue them to the ALU. This is a much more tractable problem than trying to guess the control flow at compile time. A large enough prefetch buffer can overcome a really dumb compiler. 3) Requiring software to be aware of details of your hardware implementation is a really tempting idea, but it has been historically much worse than the opposite. Consider a modern x86 that is nothing like an 80386, but runs the same software, often at higher IPCs than an original 386. Now compare to MIPS which has its delay-slot for branches which often just gets filled with a NOP. Furthermore on modern MIPS cores, which have longer pipelines and branch-prediction, that slot is more-or-less useless! 4) Assuming IA64 doesn't die out, by the time compiler writers figure out how to make code that runs fast on an Itanium of today, Intel will be performing hardware gymnastics to make that code run fast on the hardware of tomorrow.
- cx01 16y agoRegarding point 2, I wonder how much of a benefit Itanium would see from JIT-compiled languages, because the JIT could then dynamically arrange instructions to maximize ILP.
- sb 16y agoWhile I certainly think this would make for interesting research, I think the runtime-complexity of VLIW algorithms (such as Monica Lam's "Software Pipelining") would definitely interfere with the upper-bounds for compilation time of JIT compilers. (But, then again you could always use a background optimization thread...)
- sb 16y agoRegarding point 2: Last time I checked (IIRC 2006-ish), there seemed to be common resentiment among scientists working in the area of programming language implementation that there is just too little ILP for successful wide-spread VLIW adoption (modulo some special use cases.) AFAI(K|R), Hennesy and Patterson's cannonical text (CA-AQA [1]) reflects this: going from 3rd to 4th edition, we find a new chapter "Limits on ILP", VLIW/EPIC elements have been moved from the main contents to the CD-ROM, too (which probably is not a good indicator, though: the 3rd edition was just too heavy to carry it around a lot ;) [1]: http://www.amazon.com/Computer-Architecture-Quantitative-Approach-4th/dp/0123704901 http://www.amazon.com/Computer-Architecture-Quantitative-App...
- jey 16y agoWhat's the main customer/application/market for Itanic, er, Itanium?
- jacques_chester 16y agoHPC work involving easily parallelised problems have been the major market for Itanium. I suspect that GPGPUs will steadily eat this space up, though.
- sfk 16y agoNot sure if it is the main market, but you can still run OpenVMS on HP Itanium servers: http://h71000.www7.hp.com/index.html?jumpid=/go/openvms http://h71000.www7.hp.com/index.html?jumpid=/go/openvms
- yuhong 16y agoYep, HP customers is the main market for Itanium nowadays, running HP-UX or OpenVMS.
- VladRussian 16y agoThe monster is back. >Itanium relies on the compiler to optimize code at run-time Thats sums it up for me :) But seriously - in case when compiler is able to do parallelization, NVidia GPU seems to be a better - cheaper, more accessible and performant - target.
- jbri 16y agoThere are serious limitations on what you can actually do on a GPU - some operations are terribly slow, and flipping back to the CPU to carry them out is also slow. They're great for embarrassingly-parallel simple operations, but once you break the ceiling on complexity you're often better off trying to vectorize it on the CPU rather than try and manage all this chatter between the CPU and GPU. Itanium seems like it fits nicely in that space - operations complex enough to be very painful on a graphics processor, but parallel enough for you to actually consider using a GPU in the first place.
- scott_s 16y agoNot all parallelism is the same. In this article, parallelism usually means instruction level parallelism (http://en.wikipedia.org/wiki/Instruction_level_parallelism http://en.wikipedia.org/wiki/Instruction_level_parallelism).GPUs are fantastic at data parallelism (http://en.wikipedia.org/wiki/Data_parallelism http://en.wikipedia.org/wiki/Data_parallelism). Being able to exploit one says nothing about the other. Itanium differs from other processor architectures is how it handles instruction level parallelism. The processor in the computer in front of you probably uses out-of-order execution (http://en.wikipedia.org/wiki/Out-of-order_execution http://en.wikipedia.org/wiki/Out-of-order_execution) to exploit ILP. This happens on the fly, as a program executes. Itanium depends on the compiler to determine where ILP is.
- VladRussian 16y ago>Being able to exploit one says nothing about the other. data parallelism is a partial case of instruction level parallelism - N instances of the same instruction run for different pieces of data. A very frequent case in high performance computing or enterprise data crunching tasks supposedly targeted by the Itanic
- jacques_chester 16y agoThat's a massive transistor budget. You could fit ~440 MIPS R10k processors on that thing.