4 ms·
Likely much worse write endurance than DRAM, since it’s still flash underneath. The saving grace is that model weights are mostly read-heavy, so endurance may m
by LogTrim 1mo ago
Likely much worse write endurance than DRAM, since it’s still flash underneath. The saving grace is that model weights are mostly read-heavy, so endurance may matter less than it sounds.
- rbanffy 1mo agoTrue, but every time you load a new set of model weights, you are spending writes.
- timmmmmmay 1mo agoYeah, but how often is this? Certainly there are weird hobbyist edge cases that don't do well with this, that's true of anything, but a provider is loading weights once every few months
- rbanffy 1mo agoIt all depends how many models you are serving from that flash and whether they all fit in there together. If you need to evict and load models, the flash will die a horrible death. This is one resource you should NEVER underprovision.
- rando1234 1mo agoThis is a game changer for local inference IMO, where you will basically never need to do this. On the other hand in a serverless/cloud setting it may be more problematic if you don't want to strand GPUs with locally attached HBF.
- FuckButtons 1mo agoSure, but 1 write per month is essentially zero even for flash.