5 ms·
My recollection at the time was that there were good reasons to believe compilers would be workable. We were in the golden years of Java bytecode translations,
by frognumber 2y ago
My recollection at the time was that there were good reasons to believe compilers would be workable.
We were in the golden years of Java bytecode translations, real-world JITs, and Transmeta (which made a similar bet in hardware). That was the era we were starting to introduce rather complex transformations and optimizations. Computers were just starting to become powerful enough to where this kind of analysis and optimization felt possible and practical.
The compiler problem felt complex, but not unsolvable.
There were no good, obvious reasons I recall why the Itanium optimization problem couldn't be solved. The basic philosophy was make the hardware as fast as possible, even if it was hard to program, just rely on the compiler to deal with it.
Now, in hindsight, I can give all the reasons it didn't work. Key among them is that we drastically underestimated the degree to which complex software systems are hard to build and the way code messiness grows exponentially for these kinds of systems.
However, that wasn't really knowable in hindsight, though. Around the time, a lot of firms were making similar bets, and in 2001, I likely would have made the exact same bet with what was known at the time.
As a footnote, the anticipated progress in compilers is being made, but much more slowly than anticipated. NVidia reached its market cap on architectures being explored at that time. I had plenty of faculty promise similar SIMD/MIMD architectures would be increasingly important, but __dramatically__ underestimating the time it would take to get there.
- hylaride 2y ago> We were in the golden years of Java bytecode translations, real-world JITs, and Transmeta (which made a similar bet in hardware). That was the era we were starting to introduce rather complex transformations and optimizations. Computers were just starting to become powerful enough to where this kind of analysis and optimization felt possible and practical. Funnily, the static scheduling of Itanium was the opposite direction. > There were no good, obvious reasons I recall why the Itanium optimization problem couldn't be solved. The basic philosophy was make the hardware as fast as possible, even if it was hard to program, just rely on the compiler to deal with it. Oh, it technically could have been solved via compilers AND further architecture refinements (though it's a lesson in arrogance that Intel shunted this responsibility onto software developers). x86 is a complex mess of an instruction set, too. But it took decades of incremental hardware optimizations and compiler improvements to get it to scream. But there was a huge market for it, so anybody and everybody was working to improve it. You had everybody from big corporations to scrappy video game developers (especially John Carmack) figuring out ways to make it run faster, often via bugs in the architecture. > However, that wasn't really knowable in hindsight, though. Around the time, a lot of firms were making similar bets, and in 2001, I likely would have made the exact same bet with what was known at the time. I would argue that it was somewhat knowable. This wasn't the first time new CPU architectures were created. Many commercial UNIX's moved CPU architectures, some a few times. Some of their choices were similarly problematic. There were many technical and philosophical arguments on how reduced RISC processors should be. I'm quoting from memory, but I saw a great statement somewhere that fans of RISC tended to be people forced to program assembly in university, which was painful on CISC architectures. This often clouded their judgement that extra instructions can dramatically speed up computing and the complexity can be abstracted by higher level languages and compilers/JITs. The real world is a harsh mistress (this is not to say RISC CPUs didn't have performance benefits in many cases, but it was mostly where large amounts of data needed to be computed on like databases). > As a footnote, the anticipated progress in compilers is being made, but much more slowly than anticipated. NVidia reached its market cap on architectures being explored at that time. I had plenty of faculty promise similar SIMD/MIMD architectures would be increasingly important, but __dramatically__ underestimating the time it would take to get there. The big difference is that GPUs, which are a fundamentally different paradigm of computing from CPUs, provided an immediate improvement even before heavy optimizations could be done. GPU development was kickstarted by games and it took decades for it to branch out to other large markets (first crypto, then AI). But the real benefit of GPUs is that it didn't replace CPUs, but operated in parallel. You could even run quake without one, at obviously reduced performance. But intel also screwed up with GPUs... I'm a big believer in incentives. Intel's incentives were to make a CPU that nobody else could copy like x86. The complexity was probably internally seen as a benefit so it'd be harder to copy. Intel thought it could dictate to the market what it was going to get and it arrogantly saw its market domination as being stronger than it was, leaving AMD to create 64 bit extensions to x86 on its terms (hilariously most compilers call the architecture amd64). But the biggest incentive mismatch was there was no incentive to optimize for Itanium in the market. If nobody is buying it, are you going to spend time and money making your compiler or software better for it? If you're Microsoft, how much money are you going to sink into optimizing windows, .NET, VS Code, SQL Server, etc when there's no ROI? It's also a chicken-egg problem where you can't optimize when almost nobody is using it and seeing the bottlenecks. There were no John Carmacks spending hours in a debugger trying to squeeze out performance hacks. Meanwhile, ARM started to grow by being good at specific things at first (performance per watt at a relatively low price). This mattered first in portable devices, then moved to being important in dense datacenter environments which encouraged development of more raw performance. Now we're seeing it in top of the line consumer computers. Itanium was promised to be to be all of this at once - almost overnight. IMO it was doomed to failure the second that first sales graph was created: https://en.wikipedia.org/wiki/Itanium#Expectations https://en.wikipedia.org/wiki/Itanium#Expectations. If you try to make everybody happy at once, you usually end up with nobody being happy.
- frognumber 2y ago> I would argue that it was somewhat knowable. This wasn't the first time new CPU architectures were created. Many commercial UNIX's moved CPU architectures, some a few times. Some of their choices were similarly problematic. There were many technical and philosophical arguments on how reduced RISC processors should be. I think the key difference was that in the early days, one could only afford a single-pass compiler. Then, double-pass (but it was slow). Itanium was just at the time when compilers could _practically_ do deeper program analysis. There is no individual piece in making a good compiler for Itanium which I can't solve. At the time, most interesting CS problems that smart people were working on are what we called "microprogramming" in my college jargon -- optimizing individual algorithms, optimizing an assembly loop, etc. What made an Itanium compiler hard was (again in local jargon) "macroprogramming" -- making all those things work together. > I'm quoting from memory, but I saw a great statement somewhere that fans of RISC tended to be people forced to program assembly in university, which was painful on CISC architectures. This often clouded their judgement that extra instructions can dramatically speed up computing and the complexity can be abstracted by higher level languages and compilers/JITs. The real world is a harsh mistress (this is not to say RISC CPUs didn't have performance benefits in many cases, but it was mostly where large amounts of data needed to be computed on like databases). This doesn't match my history. The hypothetical difference was the other way around. CISC had complex instructions (which might take many cycles to run, like a string copy). RISC has simple instructions. Ergo, "reduced instruction set." The technical difference in early processors was RISC was pipelined and CISC was microcode (where all instructions took many cycles). The reason for CISC was largely so programs could be smaller. This made a huge difference if a computer has e.g. 8k of RAM. CISC, circa Pentium days, was hard to make fast because: - CISC variable length instructions were hard to decode, and RISC fixed-length ones were easy. This mostly disappeared as decode units became smaller perhaps circa 2005. - CISC was hard to pipeline. Again, circa 2005, it became easy to (1) avoid annoying instructions in code and treat them as very slow backwards-compatibility special cases in hardware; or (2) do a translation. For the most part, the practical distinction disappeared around then. > The big difference is that GPUs, which are a fundamentally different paradigm of computing from CPUs, provided an immediate improvement even before heavy optimizations could be done. GPU development was kickstarted by games and it took decades for it to branch out to other large markets (first crypto, then AI). True. Although that's more of a business distinction than a technical one. > But the real benefit of GPUs is that it didn't replace CPUs, but operated in parallel. You could even run quake without one, at obviously reduced performance. I'm betting 50% on convergence, as number of CPU cores grows, and GPU cores become increasingly complex. I think the M1 may be the track we eventually converge on, with diverse cores optimized for diverse tasks. > But intel also screwed up with GPUs... True. Although they seem on an okay track if they can get the software right. A770 16GB can be had for $300. NVidia 4060 16GB is $450 (and faster). For a lot of non-gaming users (ML, CAD, etc.), the A770 is ideal. > I'm a big believer in incentives. Intel's incentives were to make a CPU that nobody else could copy like x86. The complexity was probably internally seen as a benefit so it'd be harder to copy. Intel thought it could dictate to the market what it was going to get and it arrogantly saw its market domination as being stronger than it was, leaving AMD to create 64 bit extensions to x86 on its terms (hilariously most compilers call the architecture amd64). I am too -- although that's a business rather than technical argument -- but I'm not sure that's exactly what happened here; I think it's about half-right. I think Intel simply overestimated the amount of time it would stay dominant in the market, underestimated the competition, and underestimated the time Itanium would take to develop. I don't think they were intentionally making the CPU either easy or hard to copy.