3 ms·
I had to double check those figures on Sk Hynix office web site [1], and it is not a typo or wrong capital "B". It really is 3TB per second. I literally pause
by ksec 2mo ago
I had to double check those figures on Sk Hynix office web site [1], and it is not a typo or wrong capital "B".
It really is 3TB per second.
I literally paused for 5 min and thought how is this even possible.
[1] https://news.skhynix.com/en/hbf-at-fms-2026/ https://news.skhynix.com/en/hbf-at-fms-2026/
- amluto 2mo agoThis is almost entirely dominated by the read circuitry and the data path: it’s still taking 1/6 of a second to read the whole chip, which means that the flash cells aren’t working hard at all. (And that pitting the full weights of a dense model on these chips while using anywhere near all the capacity is a nonstarter if you intent to stream the weights as you run inference.)
- MobiusHorizons 2mo agoI think the theory was that when each weight will be needed is predictable, so the latency can be hidden by fetching earlier (or more likely building the data in such a way that streaming it linearly brings the right weight at the right time)
- rcxdude 2mo agoThe crazy part is the interlink: flash bandwidth scales very well with capacity, the challenge is getting that much bandwidth in and out of the chip.