3 ms·
Chiplets aren't so much driven by a need for heterogenous processes as much as minimizing the yield impact of defects. Ian Cutress has a good portion covering i
by conjecTech 4y ago
Chiplets aren't so much driven by a need for heterogenous processes as much as minimizing the yield impact of defects. Ian Cutress has a good portion covering it in this recent video: https://www.youtube.com/watch?v=oMcsW-myRCU https://www.youtube.com/watch?v=oMcsW-myRCU.
In short, if your process produces 20 defects per wafer, and you can fit 100 chip on a wafer, you're going to end up with ~84% yield(ie, slightly less than 20% loss). If you are able to split that same chip into 2 equal-sized pieces and make twice as many, your yield is now above 90%.
AMD has also made them composable in such a way that they have to produce and stock far fewer ICs to fulfill all of their SKUs, which is also another fantastic benefit.
- kayson 4y agoAbsolutely. Yield is obviously a big part of cost-per-die, but it's not just defect density related. When a node is first released, the process is not as well controlled so you end up with yield loss because the process doesn't match the models (in the extremes) and dies fail at various testing stages. This obviously gets better over time, but it's still a major issue for companies designing in the latest technologies. Anecdotally I'd say there's a ~3X reduction in the "sigma" of the process over 3-5yrs. Of course these issues have always existed, but wafer costs have never been so high, and yields never so low (both because of increased defect density and worse process control with new transistors).
- Vt71fcAqt7 4y agoHow do they combine the two seperate chips into one? Is there a performance/size cost? Could we have 1,000 chiplets per CPU and acheive >99% yield?
- dwaite 4y agoThey use an internal interconnect, similar to how a server which took multiple distinct CPUs would use a motherboard-level interconnect. In addition to splitting the CPU or GPU into distinct units, you can also take other functions and use different processes for them. For example, in package I/O or L2 cache don't really see the same advantages for newer processes, so you can make these using more established (and cheaper/more available) processes.
- conjecTech 4y agoYou get most of the economic value from the first few divisions. Going from 80% yield to 90% yield shaves 14% off your cost/unit. Going from 95% to 99% only saves you 4%. I'm not an expert at how they combine chips. Like I said for AMD, they also wanted their units to be composable w/ small number of chips, so they basically have a die w/ a few cores and different die w/ memory, and I believe they have a proprietary communication mesh for connecting them. I think there is some considerable signal/energy overhead to communicating between chips. The cost of masks and interconnects is probably high enough to make a high number of diverse chiplets inviable, but I wouldn't say it's impossible in the future.
- Vt71fcAqt7 4y agoThanks. >I think there is some considerable signal/energy overhead to communicating between chips. I tried looking for some info on this but couldn't find any. Do you have any source I could read on this?
- kayson 4y agoThis is a fundamental effect because of the Shannon-Hartley theorem[1], which says your communication channel's capacity (i.e. bitrate) has a log dependency on the channel's Signal to Noise Ratio. In a practical wireline communication system, you have a transmitter with noise, a lossy interconnect, and a receiver with noise. As the interconnect gets longer, which is what happens when you go from on-die communication to die-to-die communication, your loss increases. This means you have to reduce your transmitter and/or receiver noise which requires using additional power. Another way of looking at it is bit error rate [2]. I think you'll be hard-pressed to find concrete numbers since these designs are closely-guarded trade secrets. You might find some examples by searching for wireline transceiver / PHY papers on IEEE xplore, especially at ISSCC (conference) or in JSSCC (journal). [1] https://en.wikipedia.org/wiki/Shannon%E2%80%93Hartley_theorem https://en.wikipedia.org/wiki/Shannon%E2%80%93Hartley_theore... [2] https://en.wikipedia.org/wiki/Bit_error_rate https://en.wikipedia.org/wiki/Bit_error_rate
- kayson 4y agoThere are a variety of 2.5D (chip next to chip) and 3D (chip on top of chip) packaging techniques. Each comes with different tradeoffs. Here are a couple of articles with more detailed info: https://semiwiki.com/semiconductor-manufacturers/tsmc/306329-advanced-2-5d-3d-packaging-roadmap/ https://semiwiki.com/semiconductor-manufacturers/tsmc/306329... https://www.anandtech.com/show/16051/3dfabric-the-home-for-tsmc-2-5d-and-3d-stacking-roadmap https://www.anandtech.com/show/16051/3dfabric-the-home-for-t... https://www.eetimes.com/amd-tsmc-imec-show-their-chiplet-playbooks-at-isscc/ https://www.eetimes.com/amd-tsmc-imec-show-their-chiplet-pla... https://www.techpowerup.com/292256/amd-details-its-3d-v-cache-design-at-isscc https://www.techpowerup.com/292256/amd-details-its-3d-v-cach...
- easygenes 4y agoIf you're interested in more details on the cost of producing chips on the latest nodes, this is a good back of the envelope sort of breakdown for the Ryzen 7950X. [1] They use a visual die yield calculator to talk through the logic of improving economics using chiplet designs. The figures on the actual full wafer costs for TSMC 5nm are not public, but analysts say a ballpark around $17K is pretty reasonable. Using all the figures they ballpark, the manufacturing costs of the 7950X die collection, fully packaged, sounds like it averages around $70. So the cost is high as far as large scale chip manufacture goes, but this is also a part that retails for $600; there's certainly plenty of room to remain profitable at the top. Realistically, R&D costs are the biggest consumer of profit margins still. 1: https://www.youtube.com/watch?v=oMcsW-myRCU
- conjecTech 4y agoYeah, that video is fantastic. I wonder if the higher costs will place more of an emphasis on running chips at power levels that will gives better longevity. We can effectively halve operating cost and double lifetime of chips by reducing their power levels 20%. Seems plausible if we don't have technological obsolescence to force depreciation.
- easygenes 4y agoDon't chips at stock power levels with reasonable cooling have extremely high resilience for at least 5 years? Most datacenters have no more than a 4 year refresh cycle for servers based on power and space efficiency optimizations. For example, der8auer did a test [1] with chips like the 5800X with high overclock and high voltage with stress tests for over 4000 hours and came essentially to the conclusion that the chips are likely to endure for at least 5 years under rather extreme conditions. Likely much longer under normal circumstances. The number of systems he's testing isn't statistically significant like you might get by trying this sort of thing at datacenter scale, but it is fairly illustrative. 1: https://www.youtube.com/watch?v=ZAww0c2m-ks
- conjecTech 4y ago5 years is probably a fair current estimate, but that also doesn't come for free. The fans in DC servers move a ton of air, which itself further increases power use. I was thinking more on the order of 10-20 years.