2 ms·
This is almost entirely dominated by the read circuitry and the data path: it’s still taking 1/6 of a second to read the whole chip, which means that the flash
by amluto 2mo ago
This is almost entirely dominated by the read circuitry and the data path: it’s still taking 1/6 of a second to read the whole chip, which means that the flash cells aren’t working hard at all. (And that pitting the full weights of a dense model on these chips while using anywhere near all the capacity is a nonstarter if you intent to stream the weights as you run inference.)
- MobiusHorizons 2mo agoI think the theory was that when each weight will be needed is predictable, so the latency can be hidden by fetching earlier (or more likely building the data in such a way that streaming it linearly brings the right weight at the right time)