11 ms·
Yeah this isn't quite what happened. Firstly, Intel didn't start Itanium, HP did, as a successor to their HP Precision line. I forget how they got together, b
by FullyFunctional 4y ago
Yeah this isn't quite what happened. Firstly, Intel didn't start Itanium, HP did, as a successor to their HP Precision line. I forget how they got together, but it was a collaboration between Intel and HP, but HP started it and had largely the architecture defined before Intel got involved.
Secondly, it's true that AMD hammered the nails in the coffin, but AMD wouldn't have mattered if Itanic had been faster, cheap, and on time. Itanic was a disaster partly because of overly complicated design by committee and partly because of the fundamentally flawed assumption (that you don't need dynamic scheduling, AKA OoO processing).
I have an Itanium in the garage, a monument to hubris.
UPDATE: I forgot to mention that from the outside it might seem that Intel had a singular vision, but the reality is that there were massive political battles internally and the company was largely split into IA-64 and x86 camps.
UPDATE2: Itanium was massively successful in one thing: it killed off Alpha and a few other cpus, just based on what Intel claimed.
- hinoki 4y agoI thought the biggest problem with Itanium was the fact that it was optimising the wrong thing; it maximised single thread performance by going all in on speculative execution, but it turns out that optimising joules per instruction is much more important.
- ghaff 4y agoCertainly one of the issues with Itanium was it was fighting the last war. It was trying to optimize for instruction-level parallelism when power efficiency and thread level parallelism were coming into vogue. Arguably companies like Sun overoptimized for the latter too soon but it was the direction things were going. A senior exec at Intel told me at the time the focus on frequency in the case of Netburst was driven by Microsoft being uncomfortable with highly multi-core designs--and I have no reason to doubt that was one of the drivers. There was a lot of discussion around the challenges of parallelism, especially on the desktop, at the time. It generally wasn't the problem the hand wringing suggested it would be.
- zozbot234 4y agoParallelism is still very much underused on desktop. Most desktop CPU's are used much of their time for running single-threaded JavaScript from some clunky website - no parallelism whatsoever. It's only with the latest-gen CPU's that have things like big.LITTLE that it's becoming a real game changer.
- ghaff 4y agoI think it's generally fair to say though that the applications that really need a lot of performance (e.g. multimedia) multi-thread pretty well and there are typically a lot of background tasks running that consume core cycles as well. What is probably more generally true is that a modern laptop or desktop is ridiculously overpowered for most of what we throw at it. I'm typing this on my downstairs 2015 MacBook and it's perfectly fine for the almost entirely browser-based tasks I throw at it.
- zdragnar 4y agoMy wife has only ever owned cheap Chromebooks, and has never complained that they were slow. I've used them with her streaming videos and such, and I agree- simple web browsing isn't slowing anything down on modern hardware. Even on a modern ultralight laptop, I can run two chrome profiles, three instances of vscode running different projects, docker and a few other things and the CPU never gets pegged. There's a ton of memory pressure from a memory leak somewhere that I haven't bothered tracking down yet- I suspect the SWC compiler (thanks, rust) but haven't proven it yet. All that and I'm still getting 8+ hours of battery life.
- deleted 4y ago[deleted]
- formerly_proven 4y agoPA-RISC seemed fairly neat from the somewhat limited information you can find online (there’s a 1.1 and 2.0 ISA manual on kernel.org). Where there major issues with the ISA? Killing your entire product line and starting over of course rarely worked.
- rsaxvc 4y agoThe PA-RISC 1.1 ISA encoded specific implementation details that didn't age well, like the branch delay slot and instruction address queues. And required in-order memory accesses, because there wasn't support for cache coherent IO.
- ghaff 4y agoAnd, at a higher level, Unix vendor-specific processor designs were on the way out. Designing and manufacturing a processor just for your relatively low volumes in the scheme of things Unix systems was just way too expensive. One of the other problems with Itanium was that it was supposed to be an "industry standard" 64-bit processor. But Intel and HP were never quite able to square that with a situation where HP at least saw themselves as more equal than others given their role in the design.
- pjmlp 4y agoHP-UX 11 was one of the UNIXes I worked on (1999 - 2002, 2005), and the only issue I had, at least during the first employer was their ongoing transition to 64 bits, and the C compiler we had available (aC) was a mix of K&R C with some ISO/ANSI C compliance. I don't recall any issues with the ISA, and we really liked using its early container capabilities (HP Vault).
- inkyoto 4y agoItanium was a second Intel VLIW design. The first one was the i860, which was a mixed bag of either being eye popping fast, if instruction bundles were handcrafted by a human, or being as slow as a dog if it was a compiler that emitted the code. Perhaps, there was a belief back then that compilers could be easily optimised or uplifted to generate fast and efficient code, and that did not turn out to be the case. Project management and mismanagement certainly did not help either. I wonder how a VLIW architecture would pan out today given advances in compilers in last three decades, and whether a ML assisted VLIW backend could deliver on the old dream.
- pclmulqdq 4y agoGPUs today have a more VLIW-like architecture than CPUs and almost every neural network accelerator is a VLIW chip of some kind. It's worked out really well. The big problem is SMT, since it's hard to share a VLIW core between processes, while a superscalar core shares really well.
- inkyoto 4y agoIndeed. RISC-V instruction fusion is a little bit VLIW like. I wonder how it handles the fused instruction transition across CPU cores tho (I have not looked into it).
- pclmulqdq 4y agoMost superscalar cores do macro-op fusion, including ARM and x86. They don't transition the fused op - it usually executes in one cycle (which is the point of the fusion) so you transition either before or after it.
- cwzwarich 4y ago> GPUs today have a more VLIW-like architecture than CPUs Older GPUs seemed more VLIW-like because they were descended from fixed-function rendering pipelines and essentially just exposed the control signals via instruction encodings. Over time, shader cores have become less VLIW-like, e.g. look at any reverse engineering of recent Nvidia architectures. This makes sense for the same reason you give for SMT: if you're trying to execute from multiple instruction streams on the same execution units, it makes more sense to use small individual instructions rather than large puzzle pieces.
- throwawaymaths 4y ago> fundamentally flawed assumption (that you don't need dynamic scheduling, AKA OoO processing) I'm still unconvinced this is fundamental. It certainly was flawed back then, but compiler theory has improved a LOT since then, we have polyhedral optimization, e.g. that we didn't have access to... You could probably optimize delay line technology that way.
- JonChesterfield 4y agoIf you know how long data will take to go to/from memory then you can schedule pretty well. If you don't know whether some value will hit in the L1 or the L3 cache there's wild variance on how long it'll take so you have to do something else in the meantime. On x64, that's the pipeline and speculation. On a GPU, you swap to another fibre/warp until the memory op finished. Fundamentally the hardware knows how long the memory access took, the software can only guess how long it will take. That kills effective ahead of time scheduling on most architectures.
- throwawaymaths 4y agoIirc caches mostly exist to help keep pipelines full. You could maybe imagine an architecture with deterministic memory access times from specified regions.
- cwzwarich 4y agoSuch architectures exist (e.g. for signal processing). They just aren't good for running general-purpose software.
- Kon-Peki 4y agoItanium sold, barely, and still has people using it. But, as you can imagine, not with general-purpose software. By the 2nd or 3rd generation, Intel had made some changes that really improved performance a lot. Was the final generation of Itanium the 4th generation?
- 4y ago
- Findecanor 4y agoI've heard a rumour that supposedly one of the lead designers of the IA-64 architecture had died prematurely mid-project. And that would have left the project without the man with the vision. Hence, design by committee.
- amrb 4y agoTalk about bad omens!
- Findecanor 4y agoFound where I heard it: <https://youtu.be/JS5hCjueqQ0?t=4054 https://youtu.be/JS5hCjueqQ0?t=4054>
- _a_a_a_ 4y agoFrank Dobberpuhl died in 2019. Also the Alpha was the utter opposite of design by committee. A quick skim of the alpha ISA would show you that.
- jasoneckert 4y agoI've always thought that killing off Alpha in favour of pushing Itanium was one of the worst things Intel/HP could have done. Not only was Alpha more advanced architecturally, it was actively implemented and mature. With active development by HP, it could have easily snowballed into the standard cloud hardware platform.
- ghaff 4y agoWhat would become Itanium was publicly announced about eight years before HP acquired Compaq.
- jabl 4y agoAlpha's fate, like the other proprietary RISC architectures that focused on the lucrative but in the end small workstation market, was sealed. With exponentially increasing R&D and manufacturing costs, massive industry consolidation was inevitable. That it was Itanium that delivered the coupe de grace to Alpha was but the final insult, but it would have happened anyway without Itanium. And it wasn't like the Alpha was some embodiment of perfection either. E.g. that mindbogglingly crazy memory consistency model.
- runroader 4y agoI don't disagree with the main points, but Alpha wasn't just focused on the small workstation market. Alpha for lots of us in the IT departments of SMBs was the go-to when your Exchange server couldn't handle the load anymore. DEC and by extension Alpha died as soon as Ken Olson was pushed out.
- _a_a_a_ 4y ago> that focused on the lucrative but in the end small workstation market I don't at all believe this was true. > that mindbogglingly crazy memory consistency model I guess the utterly competent designers were actually stupid eejits then? The memory consistency was AFAI could determine to reduce to the utmost the hardware guarantees and therefore hardware complexity. It was done for speed. I have the greatest respect for the Alpha design team, they designed a thing of elegance and even beauty. You could learn a lot from it - I did.
- ptx 4y ago> the fundamentally flawed assumption (that you don't need dynamic scheduling, AKA OoO processing) Could this have worked better with JIT-compiled applications, e.g. Java given a sufficiently clever JVM, where assumptions can be dynamically adjusted at runtime? (Edit: As opposed to an AOT compiler.)
- acdha 4y agoThis was the grand hope but it never panned out. It’s possible now that a sufficiently brilliant compiler could make a difference since there was nothing like LLVM at the time and GCC was far less sophisticated. One of the many acts of self-sabotage Intel committed was insisting on their hefty license fees for icc, which meant that almost all real-world comparisons were made using code compiled using GCC or MSVC, which were not as effective optimizing for Itanium. There’s no way they made enough in revenue to balance out all of those lost sales. The other point in favor of this approach now is that far more code is using high-level libraries. Back then there was still the assumption that distributing packages was hard and open source was distrusted in many organizations so you had many codebases with hand rolled or obsolete snapshots of things we’d get from a package manager now. It’s hard to imagine that wouldn’t make a difference now if Intel was smart enough to contribute optimizations upstream.
- ghaff 4y agoYes. Open source, high-level libraries, SaaS/Cloud, good dynamic translation (e.g. Rosetta), etc. make the sort of backward-compatibility that Intel/HP failed so miserably in providing much less of a big deal today. One of the driving forces behind Itanium was that, not only was developing custom microprocessors and OSs for a single company expensive, but even once you'd made that investment, ISVs were reluctant to support your low volumes for any amount of love and money.
- acdha 4y agoIt’s definitely interesting looking at ARM now. It’s helped by having consistently had much better price/performance but also the fact that things like phones meant a ton of the primitives people would need to switch server applications were already taken care of. Intel really would have been better off cutting their marketing department and hiring 50 more developers to work on open source like GCC, OpenSSL, Linux, etc.
- rocket_surgeron 4y ago>I have an Itanium in the garage, a monument to hubris. I love monuments to hubris so I too have one in my garage-- a four-node SGI Altix 350.
- killerstorm 4y agoIf you look at SPEC CPU benchmark, Itanium was not bad at all in terms of instructions-per-cycle. IIRC in fp performance it could beat Netburst Pentium4 running twice its clock speed, and even compared favorably to Core. I.e. if Intel produced Itaniums which ran at the same clock speed as Core CPUs, it would be world's fastest single-threaded number cruncher. So I don't really buy the "Itanium is bad architecture" story. It's fate was probably decided around 1999-2000. At that point Itanium was still pretty good against Pentium 3 and Pentium 4. And name "IA-64" indicates Intel didn't plan to make 64-bit Pentiums. So eventually Pentiums would fill low-end segment while the rest would be occupied by 64-bit Itaniums. AMD killed that plan by releasing AMD64 architecture. It was an obvious upgrade to x86, so it would clearly do better in the market than IA-64. So Intel decided to go for x86-64 too, and Itanium was doomed at that point. They didn't even bother making Itaniums with same clock speed as Xeons. So it's definitely possible that if AMD decided to stick to 32 bits at that time, Intel would have pushed optimized IA-64. Also AMD64 could be worse than it is. E.g. if they decided to increase only register size but keep the number of registers the same, IA-64 could still come on top.
- giantrobot 4y ago> AMD killed that plan by releasing AMD64 architecture. It was an obvious upgrade to x86, so it would clearly do better in the market than IA-64. One of the best aspects of the Opteron was it also happened to be a fantastic 32-bit CPU in addition to AMD64. This was a period where a lot of software, even FOSS wasn't 64-bit clean. There was a lot of pointer arithmetic hiding deep in libraries that were assuming pointers would always be 32-bits. The Opteron running a 32-bit OS at least as well as a 32-bit Athlon was a huge point in its favor. So your existing system running on new Opteron hardware ran fine and you could mix and match Xeon and Opterons in a fleet. Then switch over to 64-bit on the Opterons for (hopefully) better performance.
- acdha 4y agoOne thing to remember is that FP performance with well-scheduled code was by far its strongest performing area, and Intel put a lot of work into tuning their compiler for those specific tests. The problem was that it fell off heavily the less your code is like that, especially for the branchy code most business apps depend on. The other big problem was that the x86 compatibility story was worse than the earlier hopes. That meant that it not only wasn’t competitive with the current generation competition but often even the previous or worse - note losing to the original Pentium or even a 486 here: https://tweakers.net/reviews/204/8/intel-itanium-sneak-preview-benchmarks-rc5-chess-en-stream.html https://tweakers.net/reviews/204/8/intel-itanium-sneak-previ... Now, they could have improved that but statistically nobody was going to pay considerably more for lower performance in the hopes that a future update would improve matters. The Athlon and Opteron weren’t just fast, they also had flawless 32-bit support so even if your 64-bit software update never happened you could justify the purchase based on their price/performance.
- noNothing 4y agoItanium didn't kill off Alpha. Intel x86 pricing did. But the most important unheeded lesson in those days was software compatibilty. We went from the days of each computer having its own word length, instruction set, heck, even data format (remember the endian wars?) to source compatibility to binary compatibilty. We learned that for most useage, software stability that allowed taking advantage of Moores law, was seriously more valuable in most cases than gaining a bit more performance or price/performance by changing architectures. Intel kept the X86 price at a point where no bean counter would favor investing in new architectures. Fortunately AMD broke the headlock on x86.
- sliken 4y agoWell Intel's original plan was to keep x86 32 bit, forcing anyone that needed more into IA64. Fortunately AMD came out with x86-64, and when it was clear that IA64 wasn't going to be competitive, Intel brought x86-64 to their chips.
- kjs3 4y agobut AMD wouldn't have mattered if Itanic had been faster, cheap, and on time I dunno about that. My own personal opinion is that Intel has never been able to re-architect their way out of the fact that the cornerstone of their success is that they are selling x86, and their customers mostly don't care about the theoretical advantages of the bright shiny new thing. They just want to run their software just like they always have. There's a reason why IA-64 has joined iAPX432, i860, StrongARM and i960 as Intel footnotes (outside of the embedded market). When they were philosophizing about what Itanic should look like, the only thing that x86 obviously needed from a market perspective was a bigger address space. And AMD was smart enough to deliver on that, and here we are.