2 ms·
A full wafer like Cerebras is about 60x that, and N2P has about 3x the transistor density. So right now it's technically feasible to etch a 1.4 trillion paramet
by phonon 2mo ago
A full wafer like Cerebras is about 60x that, and N2P has about 3x the transistor density. So right now it's technically feasible to etch a 1.4 trillion parameter model. So roughly DeepSeek-V4-Pro class. Imagine that running a factory, for example.
- briansm 2mo agoCerebras have special techniques to work around etching errors / bad cores on their wafers. This is possible since their wafers are effectively hundreds of identical copies of redundant cores. Can't do that for a globally unique model. Etching failure in that situation would be like brain-damage in a human, all sorts of weird effects would start appearing.
- fulafel 2mo agoThere's several ways to engineer around that as the errors are detectable. There's a big literature on how to trade off speed or transistors for error correction. [1] (Is Cerebras doing something novel? CPUs and memory blocks have been doing those things for a long time too, since the error rate is otherwise too high for normal size chips as well) [1] see eg https://www.vlsimentor.com/dft/redundancy-bisr https://www.vlsimentor.com/dft/redundancy-bisr to get some basic concepts
- phonon 2mo agoA few hundred bad bits/transistors in a trillion+ parameter model would compromise its abilities not one iota...the models are inherently lossy and resistant to "brain damage"...