3 ms·
I’d argue that the larger problem is not that transistors can’t shrink forever, but that we stopped finding ways for additional transistors to usefully increase
by labcomputer 5y ago
I’d argue that the larger problem is not that transistors can’t shrink forever, but that we stopped finding ways for additional transistors to usefully increase the speed of processors.
For example, these techniques improve instructions per clock, at the cost of adding transistors:
* pipelining (late 80’s early 90’s) MIPS R3000, Intel 486, Motorola 68040
* superscalarity (early 90’s) Pentium, DEC Alpha 21064, MIPS R8000, Sun SuperSparc, HP PA-7100
* out of order execution (mid 90’s) Intel Pentium Pro and later, DEC Alpha 21264, Sun UltraSparc, HP PA-8000
* SIMD (vector) instructions (mid/late 90’s) Pentium MMX (integer) Pentium III (floating point), DEC Alpha 21264
* multithreading (early 2000’s) Pentium 4, planned chips from DEC and MIPS
We also got improved branch predictors and larger, more-associative caches.
But it feels like most of that progress stopped in the early 2000’s, and the only progress is slapping more cores on a die. I mean, if you can put 8 cores in a consumer level CPU, you have 8 times as many transistor (give or take) as you need to implement a CPU. Nobody seems to be building a higher IPC CPU with 2x transistors, even though they clearly have the transistors to do it.
- derefr 5y agoOn the contrary, we got a number of improvements for novel workloads by first inventing new things we expected computers to need to do all the time (e.g. encryption-at-rest, signature validation), and then giving the ISA special-purpose accelerator ISA ops for those same operations. > Nobody seems to be building a higher IPC CPU with 2x transistors I mean, there are designs like this, but they run into problems of cache invalidation and internal bus contention. The way to get around this is to enforce rules about how much can be shared between cores, i.e. make the NUMA cores not present a UMA abstraction to the developer but rather be truly NUMA, with each core having its own partition of main memory. But then you lose backward compatibility with... basically everything. (You could run Erlang pretty well, but that’s basically it.)
- labcomputer 5y ago> [...] number of improvements for novel workloads by [...] (e.g. encryption-at-rest, signature validation) [...] Sure, but that's not helping the general case. Only specific types of workloads. You could argue that adding lots of special-purpose hardware doesn't hurt from a transistor count (we have plenty) or power perspective (turn them off when not needed), but it can make layout tricky and reduce clock speed (which slows down everything else). > [...] cache invalidation and internal bus contention [...] NUMA Sure, but the context from spamizbad was specifically single-core performance. (I probably opened up a can of worms by mentioning hypertheading). The problem is many real-world workloads that add business value are not embarrassingly parallel problems. If it worked like that, Thinking Machines would have swept the court since the 1980's (they had NUMA like you are describing). The point is that, since about 2005-2010ish, single-thread performance has mostly stalled. Intel CPUs can issue slightly more instructions per clock. AMD has a slightly better branch predictor. But performance growth has mostly been the result of adding more cores (Except Apple's M1 has some magic). The things I previously mentioned gave big IPC gains on a diverse set of real workloads. Some innovations, like pipelining and multi-issue were responsible for 2x-4x IPC each. Pipelining, in particular, was a trick that also helped clock speeds. All those innovations happened between the late 1980's and early 2000's. So an observer during that time might have just assumed that similar innovations would keep coming. But they haven't. A Pentium III has probably around 15x IPC compared to an i386 (maybe 60x if you include SIMD), in addition to a 40x higher clock speed (some of which came from adding more transistors). How can you add transistors (say 2x or 3x) to a CPU to double performance on diverse, real-world problems that don't parallelize well? My point is, I don't think anyone knows, so it is irrelevant whether there is a physical limit to transistor shrinkage. We don't even know what to do with the transistors we have, so who cares if we can't have more?