6 ms·
Why is it so important to make the chips smaller every time? Especially if they're going into servers, where, as the article says, they will almost never be se
by gavinpc 10y ago
Why is it so important to make the chips smaller every time? Especially if they're going into servers, where, as the article says, they will almost never be seen by the customers. To go from 2bn transistors to 10bn on something that's already trivially small, why not just make the thing 5 times thicker? Obviously there would be downsides in cost of materials and power consumption, but wouldn't that be offset by avoiding the need for nano-manufacturing advances in every single round? It sounds to me like there's some kind of "manifest destiny" at work here, especially the quote: “Our job is to push that point to the very last minute.” Really?
- awinograd 10y agoI think it has to do with distance electricity has to travel. If a chip is physically bigger, it takes longer to move bits inside of it. Sort of similar to why you can't have an L1 cache the same size as main memory. Just my best guess from a single processor architecture class so definitely not positive that's the answer.
- kogepathic 10y ago> I think it has to do with distance electricity has to travel. Correct. The chip has to be small enough that the clock can propagate everywhere within the chip within a single cycle, or problems will occur. > If a chip is physically bigger, it takes longer to move bits inside of it. Yes, so either you would need to delay for some cycles to ensure that the information has propagated (which will basically nullify the performance gains from cranking up your clock) or clock parts of the chip differently, but you're pretty much always limited by the slowest part of the chip (which is why every modern chip has a cache, because otherwise it would stall waiting for data).
- Dylan16807 10y ago> The chip has to be small enough that the clock can propagate everywhere within the chip within a single cycle, or problems will occur. Problems like what? These chips are already chopped up into different clock domains, and it's easy to install some PLLs so that perfectly-synchronized clock signals can blanket a chip even if it's inches across. Moving data around is also not a big deal. The Xeons in the article already have multi-nanosecond ring busses running around between cores[1]. They don't slow the chip down because the design simply lets long-distance data transfers take multiple cycles. L3 and I/O don't have to be blazingly fast in terms of latency. [1] http://images.anandtech.com/doci/8423/HaswellEP_DieConfig.png http://images.anandtech.com/doci/8423/HaswellEP_DieConfig.pn...
- kogepathic 10y ago> Problems like what? The chip won't work. > These chips are already chopped up into different clock domains, and it's easy to install some PLLs so that perfectly-synchronized clock signals can blanket a chip even if it's inches across. Sorry, I didn't explain myself well enough. Of course chips have different clock domains, but these also come at a cost. The more synchronization you need to do between domains, the less die space you have for computationally useful stuff. > L3 and I/O don't have to be blazingly fast in terms of latency. I would argue differently, the impact of latency is highly dependent on the type of computation you're doing. If you're doing something with a lot of data (say, encoding video) then you need to be moving data as quickly as possible between the processor and memory. Any additional latency in cache or I/O will cause the performance to suffer. Ideally, you want the latency of L3 and I/O to be as low as possible.
- Dylan16807 10y agoIdeally you want L3 to be fast, but it's going to be somewhat slow just by the nature of being large. And talking to ram is going to be slow even if the intra-chip pathways are infinitely fast. An extra few percent off-core latency isn't the end of the world if it lets you fit ten times as much computation on the die. L2 and L1 won't be affected. And encoding video is nearly the platonic ideal of not caring about memory latency. You could easily make memory requests ten thousand cycles before you need the results. You just need throughput.
- jaytaylor 10y agoShrinking the process also means higher yields per wafer of silicon (i.e. more chips), which decreases cost and increases profit per unit.
- zhemao 10y agoUntil recently, shrinking the process led to some obvious benefits. The chip got faster, cheaper, and more power-efficient. The last part is really the important one and it's why the big datacenter companies kept buying new chips from Intel year after year. That calculus (so-called Dennard scaling) has now broken down, which is why Intel has abandoned their tick-tock development model in favor of a model that doesn't solely rely on process shrinks in order to achieve better performance.
- Retric 10y agoOr as Intel put it: Our customers expect that they will get a 20 percent increase in performance at the same price that they paid last year. Which is less half as fast as they used to be.
- kogepathic 10y ago> Why is it so important to make the chips smaller every time? Because the amount of power leakage (heat) is proportional to the size of the transistors. So you cannot improve the efficiency of the chip without making the transistors smaller or lowering the frequency. > Obviously there would be downsides in cost of materials and power consumption, Huge costs. Data centers aren't only worried about the power consumption of chips, because for every watt that's generated as heat, they have to use >1W to cool that (because cooling systems aren't 100% efficient themselves). As you may know from other branches of science, the resistance of a conductor rises with heat, meaning that electrons running through the conductor are more likely to hit a vibrating atom and dissipate as heat. This is why so-called "super conductors" are usually super cooled. Very little atom movement = very low chance of an electron hitting a moving atom. Silicon is a semi-conductor, but the principal is the same. If the temperature of the chip rises, so does the heat generated. This is why world-record setting overclocks are done using liquid nitrogen to cool the chip. To prevent things from getting too toasty, data centers would have to reduce the density of the servers, which means they would need larger buildings and more land to house the same number of servers. > but wouldn't that be offset by avoiding the need for nano-manufacturing advances in every single round? In short, no.
- williadc 10y ago> the amount of power leakage (heat) is proportional to the size of the transistors Small correction: leakage actually increases as transistors shrink. This is why high-k dialectric and fin-fets were such important developments. They pushed back the point at which leakage power overtakes switching power as the dominant source of waste. Even with these technologies, we have to do a lot of design work to reduce leakage. I can't even guess how many power domains are on modern cpus -- certainly dozens if not hundreds. Most of those domains can be switched off to eliminate leakage in those domains altogether. I'm nearly certain you meant switching power reduces as transistors shrink, which has been true so far. Things are getting weird with these new processes, and a lot of things that we've held as fact are looking less and less reliable.
- kogepathic 10y ago
- pkaye 10y agoIf we make the chip die area bigger the probability of a defect goes up a lot. Similarly with bigger dies, there are fewer dies to a wafer of silicon. All this works out that the die cost is proportional to 4th power of area. So to make money, you want to make the smallest die possible. At the same time you want to add more features which means more transistors. Keeping the die size same means the transistors need to become smaller. Doing that means lots of new improved equipment and research. This also has a cost. So you need to ride the right balance between the two. In most of these companies, the yield rate is one of the most guarded numbers because from that you can estimate costs very easily. All this feeds into a competitive market between chip manufacturers. If you are a memory manufacturer and can find the right balance that makes the product 5% cheaper to manufacture, they can make a ton of money and gain marketshare. There are a few companies like Apple, Intel that can differentiate on brand name but the rest are commodity products that compete on cost and features.
- LeifCarrotson 10y agoThe cost is dependent on the wafer size - bigger is more expensive. Power is dependent on the capacitance of the various switching layers, so smaller features take less power to clock on and off or can go faster. Inductance is dependent on trace length, so shorter traces can be clocked faster or at lower voltages. And that (plus speed-of-light issues) means that smaller chips can both go faster and cost less! Five wafers stacked on top of each other would take 5 times as much capacity and materials to produce and cost about 5 times as much. The interconnect would be very difficult. Lastly, while you could cool the top wafer, the bottom layers would have to go through a lot of silicon to remove heat. Five times the number of layers on each wafer would also be prohibitively expensive, you don't just cut it thicker, you vapor deposit each layer under a mask and often also need to etch off parts of previous layers or ensure that higher layers are still planar. That gets difficult when you have a few steps of logic layering and some metal layers for interconnect, and would be much more expensive (and hot) with many layers. This isn't so much of a problem with eg. NAND flash chips which are low power and the address and data pins can just be shared with a couple wafer select wires to separate them, but processors are neither low power nor trivially paralleled.
- petra 10y agoBut assuming you could stack up 5 wafers - and route between them, you get to position transistors in 3d and route them in 3d - which mens wires can be shorter, transistor driving those wires can be smaller and together you could reduce a lot of the wafer area required - maybe by doing 2 layers, you can reduce total wafer area by 50% (but also get much faster low power chip because of the short wires). But true, heat is a huge problem.
- akiselev 10y agoThe size of the final die is very important to your yield from a single wafer. Imagine you have a 4in by 4in square. If your chips are 1in^2 you can fit sixteen chips but if your chip is 4in^2, you can only fit four. Now imagine you have a thick scratch going diagonally right down the middle. With the bigger chips, you might lose all four to that single error, wiping out your yield. With the smaller chips however, you'll only lose part of your wafer. Specialized processors like those made for mainframes or RAD hardened ones can be much bigger since the set up costs will vastly outnumber the fabrication cost anyway. Companies like IBM that aren't in as cost competitive a market as Intel make big chips all of the time.
- exDM69 10y agoAnother aspect of this is that chips are rectangular but wafers are round. Smaller chip means less wasted wafer area. Intel has an advantage here: they use 12 inch wafers while other fabs use 10 inch wafers. This improves their relative yield.
- InclinedPlane 10y agoAs others have pointed out, there are several factors. One is that defects in silicon tend to exist in a given areal density. Which means that for a given wafer you'll have, on average, a certain number of defects for a given process, and that in turn translates to the same number of defective chips from that wafer. The more total chips you cram on a wafer then the higher the percentage of non-defective chips you'll get, more or less. Additionally, smaller chips typically mean faster and lower power chips. Overall these are very very strong factors which push towards shrinking chip size at every opportunity.