4 ms·
> I would argue that it was somewhat knowable. This wasn't the first time new CPU architectures were created. Many commercial UNIX's moved CPU architectures, so
by frognumber 2y ago
> I would argue that it was somewhat knowable. This wasn't the first time new CPU architectures were created. Many commercial UNIX's moved CPU architectures, some a few times. Some of their choices were similarly problematic. There were many technical and philosophical arguments on how reduced RISC processors should be.
I think the key difference was that in the early days, one could only afford a single-pass compiler. Then, double-pass (but it was slow). Itanium was just at the time when compilers could _practically_ do deeper program analysis.
There is no individual piece in making a good compiler for Itanium which I can't solve. At the time, most interesting CS problems that smart people were working on are what we called "microprogramming" in my college jargon -- optimizing individual algorithms, optimizing an assembly loop, etc. What made an Itanium compiler hard was (again in local jargon) "macroprogramming" -- making all those things work together.
> I'm quoting from memory, but I saw a great statement somewhere that fans of RISC tended to be people forced to program assembly in university, which was painful on CISC architectures. This often clouded their judgement that extra instructions can dramatically speed up computing and the complexity can be abstracted by higher level languages and compilers/JITs. The real world is a harsh mistress (this is not to say RISC CPUs didn't have performance benefits in many cases, but it was mostly where large amounts of data needed to be computed on like databases).
This doesn't match my history.
The hypothetical difference was the other way around. CISC had complex instructions (which might take many cycles to run, like a string copy). RISC has simple instructions. Ergo, "reduced instruction set." The technical difference in early processors was RISC was pipelined and CISC was microcode (where all instructions took many cycles).
The reason for CISC was largely so programs could be smaller. This made a huge difference if a computer has e.g. 8k of RAM. CISC, circa Pentium days, was hard to make fast because:
- CISC variable length instructions were hard to decode, and RISC fixed-length ones were easy. This mostly disappeared as decode units became smaller perhaps circa 2005.
- CISC was hard to pipeline. Again, circa 2005, it became easy to (1) avoid annoying instructions in code and treat them as very slow backwards-compatibility special cases in hardware; or (2) do a translation.
For the most part, the practical distinction disappeared around then.
> The big difference is that GPUs, which are a fundamentally different paradigm of computing from CPUs, provided an immediate improvement even before heavy optimizations could be done. GPU development was kickstarted by games and it took decades for it to branch out to other large markets (first crypto, then AI).
True. Although that's more of a business distinction than a technical one.
> But the real benefit of GPUs is that it didn't replace CPUs, but operated in parallel. You could even run quake without one, at obviously reduced performance.
I'm betting 50% on convergence, as number of CPU cores grows, and GPU cores become increasingly complex. I think the M1 may be the track we eventually converge on, with diverse cores optimized for diverse tasks.
> But intel also screwed up with GPUs...
True. Although they seem on an okay track if they can get the software right. A770 16GB can be had for $300. NVidia 4060 16GB is $450 (and faster). For a lot of non-gaming users (ML, CAD, etc.), the A770 is ideal.
> I'm a big believer in incentives. Intel's incentives were to make a CPU that nobody else could copy like x86. The complexity was probably internally seen as a benefit so it'd be harder to copy. Intel thought it could dictate to the market what it was going to get and it arrogantly saw its market domination as being stronger than it was, leaving AMD to create 64 bit extensions to x86 on its terms (hilariously most compilers call the architecture amd64).
I am too -- although that's a business rather than technical argument -- but I'm not sure that's exactly what happened here; I think it's about half-right. I think Intel simply overestimated the amount of time it would stay dominant in the market, underestimated the competition, and underestimated the time Itanium would take to develop.
I don't think they were intentionally making the CPU either easy or hard to copy.
- hylaride 2y agoFirst of all, I’m enjoying this discussion. I also want to point out that I’m not necessarily advocating for either side in RISC vs CISC (I lean RISC personally), but I’m more pointing out how the market actually ended up and why it (was) so hard to replace x86. > I think the key difference was that in the early days, one could only afford a single-pass compiler. Then, double-pass (but it was slow). Itanium was just at the time when compilers could _practically_ do deeper program analysis. > There is no individual piece in making a good compiler for Itanium which I can't solve. At the time, most interesting CS problems that smart people were working on are what we called "microprogramming" in my college jargon -- optimizing individual algorithms, optimizing an assembly loop, etc. What made an Itanium compiler hard was (again in local jargon) "macroprogramming" -- making all those things work together. Both these circle back to a point I made earlier in the thread - who’s going to do that for a processor nobody is using? There is an alternate timeline where a combination of hardware and compiler/software iteration make Itanium competitive at a performance level, but intel’s non-technical decisions made that impossible at a practical level. It was never made cheap or available enough where anybody could tinker at the lower end, and at the high end they suffered from the chicken/egg “nobody bought it because x86 was faster today”. The Itanium did to very well in some supercomputer deployments where the code could handle the architecture’s good parallelism. But that was likely net-new code for a niche product. > The hypothetical difference was the other way around. CISC had complex instructions (which might take many cycles to run, like a string copy). RISC has simple instructions. Ergo, "reduced instruction set." The technical difference in early processors was RISC was pipelined and CISC was microcode (where all instructions took many cycles). Hypothetically, yes. In the real world a lot of work was done to minimize that over x86’s lifetime. In the early days RISC did do a lot of what was promised, but clever (some would say hacky in some situations) updates to x86 CPUs and compilers made these advantages less (although one could argue the x86 microarchitectures that showed up in the mid to late 1990s were more RISC like). Faster and better caching (variable length x86 instructions meant that common instructions can have a shorter encoding and take up less space in the instruction cache, vastly reducing expensive cache misses) also minimized a lot of these issues. The main issue was a lot of these “enhancements” kept the performance per watt ratios at very poor levels, which caused no end of headaches as laptops became more popular and left intel (and AMD with x86) with no competitive alternative to ARM in phones. > The reason for CISC was largely so programs could be smaller. This made a huge difference if a computer has e.g. 8k of RAM. CISC, circa Pentium days, was hard to make fast because: > - CISC variable length instructions were hard to decode, and RISC fixed-length ones were easy. This mostly disappeared as decode units became smaller perhaps circa 2005. > - CISC was hard to pipeline. Again, circa 2005, it became easy to (1) avoid annoying instructions in code and treat them as very slow backwards-compatibility special cases in hardware; or (2) do a translation. I agree with all these points, but even before the decode enhancements ~2005, lots of work was done to mitigate these. But the most obvious thing that x86 caught up with was clock speed. A central argument in favor of RISC was that it allowed for an ability to jack up the the clock speeds of processors, that could then iterate over the reduced instructions faster and provide better performance in most computing cases. This clock speed advantage (in practice) was eliminated, though not due to issues with RISC itself, but more because it was only the large volumes of x86 chips could justify the higher costs of staying near the bleeding edge of transistor manufacturing allowing smaller transistor sizes (something that ARM would eventually come to lead, though). The CISC instructions were often heavily improved upon over hardware iterations or moved over to new ones (MMX being a famous example) that were more in line with how code was used (hilariously MMX became kind of redundant soon after as GPUs took over those functions). There were also many cases where x86 could do things in fewer instructions than RISC, which helped even more at higher clock speeds as x86 caught up. Again, I’m not necessarily defending CISC/x86. But it had so much engineering heft thrown at it due to its install base that it often brute forced its way to performance and it was only when performance per watt metrics started to matter that an alternative came on the scene (and it was not something that Itanium would have been better at - ARM would probably be causing the same issues to intel in the data center had Itanium taken off as it was). This was always going to be a hindrance to any replacement. The fact there was genuine competition in x86 kept prices lower and development cycles active, too. A lot of what you state is correct, but in practice x86 chips were still faster for most workloads out there. In a perfect world all this time and effort would have been heaped on a far better CPU architecture (CISC or RISC), but alas… Even today x86 still outperforms ARM on most server chips, but our AWS loads are using their ARM chips because it’s cheaper per unite of compute, which is fine for most workloads. This may even get better as more focus is put on ARM via compilers or architecture iterations. > True. Although that's more of a business distinction than a technical one. It’s a technical one. They provided immediate technical enhancements, but didn’t mean all your current code had to be rewritten. It was optional (until video games got so sophisticated that they were required then) and you didn’t need to rewrite/recompile your OS to use it. However, GPUs are a lot more niche, so fundamental architecture changes can be done more easily, especially as most code out there is done via higher level APIs. > I'm betting 50% on convergence, as number of CPU cores grows, and GPU cores become increasingly complex. I think the M1 may be the track we eventually converge on, with diverse cores optimized for diverse tasks. I agree on this for most consumer products. SoCs have taken over the embedded and mobile market and that can continue to other markets. But we’ll see what happens if AI continues its current trajectory. There it’s the GPUs that matter more than the CPUs and we could see something different emerge. > I am too -- although that's a business rather than technical argument -- but I'm not sure that's exactly what happened here; I think it's about half-right. I think Intel simply overestimated the amount of time it would stay dominant in the market, underestimated the competition, and underestimated the time Itanium would take to develop. It’s all about business, though. If the world was purely technical, Amiga, DEC, or Sun would be on top of the world. Intel did everything you just said, but also tried to do too much. A lot of what intel does is over-engineered in a sense they do things in a more complex way that necessary (some cynical people would say on purpose to make larger margins on hardware/chipsets eg USB, which at a low level is very complex even for the original spec). > I don't think they were intentionally making the CPU either easy or hard to copy. At an engineering level, no. But higher up they made sure the way it worked with IP, etc would make this difficult. The fact that x86 clones exist at all was a legal miracle and a quirk of history (pretty much IBM demanding it for the original PC and intel was not big enough yet to say no): https://jolt.law.harvard.edu/digest/intel-and-the-x86-architecture-a-legal-perspective https://jolt.law.harvard.edu/digest/intel-and-the-x86-archit...