3 ms·
The weights are very unlikely to be on the chip itself. That wouldn't work for SOTA models that are terabyte scale, even quantized. This is probably an accelera
by snek_case 2mo ago
The weights are very unlikely to be on the chip itself. That wouldn't work for SOTA models that are terabyte scale, even quantized. This is probably an accelerator for specific kernels in the model, but the weights are likely loaded from memory. The chip may have SRAM to store some of the weights temporarily during inference.
- foltik 2mo agoAt least in the case of Taalas the weights are physically encoded directly on the chip. It’s composed of 4-bit multiplier cells that compute all 16 possible results in parallel. The top metal wiring layer physically selects the one that corresponds to a multiplication with that cell’s constant weight, and routes it to the next layer.
- mdp2021 2mo agoAre you sure? Source? (does not seem to be https://taalas.com/the-path-to-ubiquitous-ai/ https://taalas.com/the-path-to-ubiquitous-ai/ , for example)
- foltik 2mo agoIt’s described in this patent application [0]. There’s a bit of hand waving so the HC1 might be slightly different, but the gist is the same. https://patents.justia.com/patent/20250123802 https://patents.justia.com/patent/20250123802