3 ms·
There is this section, which acknowledges the problem of replacing components: "We also need to consider system reliability. If a dielet is found to be faulty
by richdougherty 7y ago
There is this section, which acknowledges the problem of replacing components:
"We also need to consider system reliability. If a dielet is found to be faulty after bonding or fails during operation, it will be very difficult to replace."
Their proposed solution (which is not repair):
"Therefore, SoIFs, especially large ones, need to have fault tolerance built in. Fault tolerance could be implemented at the network level or at the dielet level. At the network level, interdielet routing will need to be able to bypass faulty dielets. At the dielet level, we can consider physical redundancy tricks like using multiple copper pillars for each I/O port."
Your point is still valid, just wanted to call out their their thoughts on the issue.
- nine_k 7y agoBig CPUs / GPUs already have physical redundancy, and a way to cut / rewire a limited amount of faulty parts by laser etching in the die, prior to packaging.
- monocasa 7y agoCell processors are interesting because they do quite a bit of binning on a user facing SPI slave. You shift a thousand or so bit payload into it from a support/binrg up processor that tells it which pieces are disabled, how the PLLs are configured, etc. Hacked PS3s could reenable the binned off 8th SPE for instance.
- jdnenej 7y agoThat doesn't seem to solve the issue of when a chip fails all together. While difficult, it's not impossible or unheard of for people to replace chips on a PCB like a charge controller that has fried.
- BAReF00t 7y agoIt’s not just about faults. But about customizability and upgrades! I like to choose how much RAM and storage and which ports I want with how many generic, vector, FPGA and neural cores, thank you very much. And I like to change them later, to upgrade gradually. Even buses.