19 ms·
The Dark Silicon Problem and What It Means for CPU Designers (2013)
- blackflame7000 8y agoWhy don't they start making CPUs 3 Dimensional like a cube with 6 "processors" each with multiple cores as its "sides" with the pins on the opposite sides of the cube wall. Seems to me more internal volume might allow for more cleaver head distribution channels
- asfasgasg 8y agoIf I understand correctly, neither heat dissipation nor existing manufacturing techniques are amenable to this approach. Also modern CPUs do have more than a dozen layers IIRC.
- woah 8y agoI think op is talking about 6 flat normal chips as the sides. This would allow for cooling stuff inside the cube. Maybe having only 5 of the sides as chips would make it even easier to have a heat sink. The center of the cube could be copper or something.
- jacquesm 8y agoYou can't cool from the inside of that cube without having some way of transporting the heat out of it. All you'd end up doing is heating that inside up to the temperature of the dies and after that there would be no more cooling effect (and this would happen in a few seconds after starting the whole thing up). You could do an 'inverse' of this by cooling the dies from the outside and having the interconnects in the space in between. This would still require a lot of cooling and there would be an issue with connecting the resulting assembly to the underlying PCB.
- blackflame7000 8y agoWhat if you used electronic cooling to chill a copper thermal conductor?
- jacquesm 8y agoElectronic cooling? Do you mean Peltier elements? They have a 'hot' and a 'cold' side, and their capacity is handily outstripped by any modern CPU. If that worked we would be using Peltier coolers to make quiet PCs today, after all, that problem is only 1/5th as complex.
- jonhendry18 8y agoI think woah means that the top would be open. So instead of a cube it'd be an open box. A heat sink (or arrangement of peltier coolers attached to a fan and heat sink, or whatever) would fit down into the box and be in contact with the five core-containing sides.
- rwmj 8y agoAnother problem is that companies (though not necessarily users) love thinner and thinner phones and laptops, and a cubic CPU wouldn't fit in such a device.
- gameswithgo 8y agoless surface area per transistor makes the heat problem worse. they are already a little bit 3d though.
- blackflame7000 8y agoWhat if the cooling came from the pin side of the chip? Then the surface area is the same
- neilmovva 8y agoThat still means multiple silicon dies, which we have known how to do for a while (see: Intel Core 2 Quad from 2006, and more recently AMD Epyc). Having more dies lets you dissipate more heat, but then it's kinda hard to build low-latency / high-bandwidth interconnects between the dies. Inter-die buses go over a PCB or interposer, which impose higher parasitic capacitance and make it difficult/expensive to run wide interfaces. That's why techniques like "dark silicon" allocation are important - it allows us to get more perf in a single die.
- wongarsu 8y agoFor heat dissipation you want the most surface area per volume (because you can only transfer heat away in the surface area). The optimal arrangement for that would be a huge, flat, one atom thick surface. Another goal we have is low latency (and high clock rate, which is depenend on low latency), which suggests putting everything in a qube or even a shere. So we compromise somewhere in the middle with a square with a few layers.
- perl4ever 8y ago"For heat dissipation you want the most surface area per volume (because you can only transfer heat away in the surface area). The optimal arrangement for that would be a huge, flat, one atom thick surface." Made me think of this: https://en.wikipedia.org/wiki/Menger_sponge https://en.wikipedia.org/wiki/Menger_sponge Say we have roughly 300 sq mm, that's about 17.32 mm square, which is 8.24e+7 silicon atoms (0.21 nm) across. Then we have a surface area of roughly 1.36e+16 atoms - the area x 2. If we make a fractal sponge down to the limit of single atoms, then that's about 16.5 cycles of removing cubes from a 17.32 mm cube. Let's ignore the difficulty of doing it half a time. According to the formulas from wikipedia, the result has a volume of about 0.7% of the original, with 2.4e28 "sides" of atoms exposed. So the third dimension gets you about 1.8 million times the surface area. I suppose this isn't nearly as good as 4e10 flat sheets with 1 atom separation between each, but you could argue it's more practical because everything is connected...
- perl4ever 8y agoRe "I suppose this isn't nearly as good as 4e10 flat sheets with 1 atom separation between each" - I guess it should be better, actually, now that I happened to notice 10+16 < 28. I got confused about whether my reference point was the basic cube or the flat sheet.
- throwwit 8y agoI think a dimpled fabrication would be best... with whatever amounts to heat sinks on both sides.
- deepnotderp 8y ago1. Heat dissipation is now ~n x m times worse where n is your transistor layer count. And where m is the increased thermal resistance* 2. Power delivery is now ~n times worse where n is your transistor layer count. * 3. Connections between chips are very slow, power hungry and expensive. Fabrication of "monolithic" 3D is temperature wise painful and usually results in crummier transistors. With that being said, innovative 3D integration methods in specific applications can help a lot. Shameless plug: we at Vathys do this for deep learning chips. * to a first order of approximation
- tormeh 8y agoThe heat can be tackled in part by pumping water through holes in the CPU. I believe it was IBM that came up with this. Can't tell if it's feasible or not.
- blackflame7000 8y agoI was thinking by either creating a temperature differential on a copper conductor to chill the cube from within or that the motherboard/walls would provide cooling from the pin side.
- deepnotderp 8y agoYes but microfluidics limits how thin each die can be.
- agumonkey 8y agoWhat about stacked heat vias ?
- deepnotderp 8y agoWhat are those? Do you mean something like thermal "dummy" vias?
- 8y ago
- deepnotderp 8y agoSuch a "side cube" configuration would probably lead to longer wires than is useful.
- wolfgke 8y ago> Why don't they start making CPUs 3 Dimensional like a cube with 6 "processors" each with multiple cores as its "sides" with the pins on the opposite sides of the cube wall. Not directly an answer to this specific question, but a (somewhat) colleague who writes his PhD thesis about 3-dimensional chip design made a popular scientific lecture about this topic. As I understood it, the central problem is that it is very hard to produce chips with multiple (lots of) layers (where you want to have interconnects inbetween). In particular producing the interconnects between the layers is really hard if they can lie "everywhere" instead of only at the border. There exist multiple ideas how this might be done (e.g. drill holes with high-precission lasers into the substrate and try to fill them with something conductive), but none of them as of today "really works" (at least if we are talking about chips with somewhat more than 4 layers and interconnects everywhere inbetween).
- deepnotderp 8y agoFun fact: the through silicon via was invented by William Shockley himself
- burnte 8y agoJust the reverse. As volume increases, the relative amount of surface area decreases.
- Phillipharryt 8y agoFrom the article "The heat generation per unit area of an integrated circuit passed the surface of a 100-watt light bulb in the mid 1990s, and now is somewhere between the inside of a nuclear reactor and the surface of a star. " I can't tell if this is hyperbole or not, it amazes me but no amount of googling is coming up with a useful answer. Is anyone able to confirm or deny it for me?
- andlier 8y agoAfter a quick napkin calculation+google it seems the sun has around 20kW of power per square centimeter. So not entirely unfeasible that a cooler star, or a nuclear reactor is closer to the typical 10-100W/cm2 of a modern cpu/gpu. Still some orders of magnitude off from our closest star. (Hope calculation is correct)
- perl4ever 8y agoLooking at the wikipedia page on red dwarf stars[1], it appears a small star (M9V) might have 8% of the sun's radius and 0.015% of its luminosity. Thus, it would have about 156 times less area, and 2.3% of the output per area of the sun. So taking your figure as given, that means it would be less than 500W / cm^2. [1]https://en.wikipedia.org/wiki/Red_dwarf https://en.wikipedia.org/wiki/Red_dwarf
- ars 8y ago> 20kW of power per square centimeter I'm getting 6300 W/cm^2 (see: https://news.ycombinator.com/item?id=17589371 https://news.ycombinator.com/item?id=17589371)
- Cybiote 8y agoThe temperature of the sun at its surface is ~5772 Kelvin. To get power per unit area, use Stefan-Boltzman: σ * T^4, σ * (5772 K)^4 ≈ 6294 W/cm^2. Dividing the sun's luminosity (power) by its surface area will also give a similar value. A 815 mm^2, 250 W GPU will be 250 W / 8.15 cm^2 ≈ 31 W / cm^2.
- 8y ago
- jarym 8y agoWe just need a clueless CIO to turn up and ask if 'Cloud' will solve the problem :D
- stephengillie 8y agoCloud is busy facing off against Sephiroth. What's Batman up to?
- melbourne_mat 8y agoI love the cynicism :-)
- jarym 8y agohehe - not everyone got it from the downvotes I got!
- bogomipz 8y agoI had a question - the author states: >"The most obvious is the instruction decoder, which is near the start of the pipeline, and is responsible (in the loosest possible terms) for passing the inputs to each of the execution units." Why would it "in the loosest possible terms"? Isn't this "precisely" the job of the decoder?
- kabdib 8y agoIIRC in many chips the decoder stage translates native instructions into micro-ops, which are RISC-like and the main food for execution units. The translation is not necessarily a simple one (one native instruction is often more than one micro-op, and it's possible to collapse multiple native instructions -- especially stuff like prefixes -- into one or more micro-ops).
- bogomipz 8y agoAh OK that make sense. This is likely what the authors means here by "loosely." Cheers.
- bogomipz 8y agoI had a couple of question about this bit of history mentioned in the article. I'm hoping someone could shed some light on this: >"You can emulate floating-point arithmetic by using integer instructions—but taking 10–100 times as long." Exactly how is/was floating point arithmetic emulated using only integers? Why is that range given an order of magnitude? Is this dependent on the precision I'm guessing?
- phkahler 8y agoWe had floating point when programming in BASIC on old 6502 8bit computers. There were software routines for doing the math. You know, multiply the mantissa, add the exponents... If someone gave you pointers to a couple 4-byte chunks of data and told you to write code to do floating point multiply on the contents using only C-char variables, what would you write? That's why it's 100 times slower than a nice modern fmul. It wasn't quite that bad, they could use the carry flag which isn't available in C.
- sehugg 8y agoYep, here's the old Woz/Roy Rankin 6502 code for log, exp, conversions, and basic math: http://www.6502.org/source/floats/wozfp1.txt http://www.6502.org/source/floats/wozfp1.txt
- bogomipz 8y agoWha a wonderful bit of history your link is. Thanks for sharing.
- bogomipz 8y agoThanks these are all good reads and explanations. I guess I just have never had to think about FPU emulation before and it kind of threw me for a loop. It's amazing what we can take for granted now I guess :) Cheers.
- brandmeyer 8y agoIf you have a copy of Knuth's Art of Computer Programming, the section on seminumerical algorithms covers software floating point. Another source is the Handbook of Floating Point Arithmetic, which covers both hardware and software methods in detail. An implementation in source code may be found in libgcc. IIRC their ARM assembly code versions are some of the fastest routines available anywhere for software floating point emulation. While they are in assembly, ARM assembly is fairly intelligible. Finally, you can probably come up with 90% or more of the correct solutions yourself if you start with the definition of a floating-point number as (-1)^s * 1.m * 2^(e - bias) and grind through the algebra with the bias as a constant. The theorems that prove the minimum number and nature of the guard bits to get correct rounding are another matter, but you can start off by computing the equivalent infinite-precision terms and then rounding.
- JudasGoat 8y agoThe author states "For every watt of power the CPU consumes, it must dissipate a watt of heat." It was always my understanding that that in electrical devices, (with the exception of heaters) the amount of heat produced was inversely proportional to the efficiency of the device. So is it really true that all energy provided to the CPU or SOC is dissapated as heat?
- Elrac 8y agoThe author's sentence is essentially a truism, because fiddling with information _as such_ and in theory doesn't use up any energy. As practically implemented in current CPU electronics, though, information needs to be communicated from point A to point B as a change in voltage. To convey that voltage change means having to move some amount of electrical charge into or out of the tiny capacitor that is a transistor's gate, via the non-perfect conductor that is the doped-silicon trace between the two points. Resistance saps part of the energy moving those electrons around and converts it to heat.
- dnautics 8y ago> fiddling with information _as such_ and in theory doesn't use up any energy. that's not true at all. Any time you destroy information (for example an and gate can destroy information) you use energy.
- Elrac 8y agoChances are you know more about this than I do, and I'd be interested to learn more. If you're willing to explain, I wonder how an AND gate destroys information. Certainly the output of a 2-input AND gate carries less information than both inputs together. But unless the gate is destroying the input signals, which it isn't, I don't see how infomation is being destroyed.
- dnautics 8y agohttps://en.wikipedia.org/wiki/Landauer%27s_principle https://en.wikipedia.org/wiki/Landauer%27s_principle