3 ms·
So the next logical step would be to remove L3 cache from the compute chiplet all together? That would let AMD either save money since the chiplet is smaller or
by keanebean86 5y ago
So the next logical step would be to remove L3 cache from the compute chiplet all together? That would let AMD either save money since the chiplet is smaller or add more logic for the same die space.
This could also mean a GPU chiplet on package with compute. Each chiplet gets at least one cache layer. The next few years could be pretty crazy.
- Scene_Cast2 5y agoLatency is still unclear, and probably worse than L3.
- coreblocks 5y agoGoing over TSV will definitely incur at least one clock cycle.
- Dylan16807 5y agoIf you're on a high end chip then you already have to talk to the other chiplet for half your L3. A stacked chip should be a lot faster than that.
- ece 5y agoThe additional L3 is per chiplet, so you're still going to have to talk to other chiplets and their L3. Latency numbers would be good to see. This is definitely better than eDRAM L4 on the I/O die, and that's still something they could do, so props for that. The power cost will really determine if we see this in Laptops or not though.
- Dylan16807 5y ago> The additional L3 is per chiplet, so you're still going to have to talk to other chiplets and their L3. You will sometimes, but traffic that still overflows to the other chiplet is probably almost all traffic that would have gone to RAM before. And as a reference point, cross-chiplet L3 is about 50 nanoseconds slower than local L3. That will dwarf most things.
- ksec 5y ago1. Latency ( One layer ) is the same, as updated in the article. 2. You will still ( for now ) need L3 Cache on the Chiplet where you have layers of SRAM on top.