3 ms·
I wish we could stream active data to RAM, directly from the NVMe drive for the 2TB K3 model. Can't wait for prism-ml/ to release a ternary 2bit model, which wo
by ALLTaken 3mo ago
I wish we could stream active data to RAM, directly from the NVMe drive for the 2TB K3 model. Can't wait for prism-ml/ to release a ternary 2bit model, which would make this a bit more tenable.
- adrian_b 3mo agoYou can, but the reported speeds for such big models were in the range from a token per a dozen seconds to slightly better than 1 or 2 tokens per second.
- halJordan 3mo agoA 2 bit model is still 700gbs you need to stream through. I know we say reading is free (vs writing). But on a 2tb drive, you're doing a full drive read for every three tokens. Thats 333,000 drive reads just to fill up the context. Well, this is at least an moe model so not that terrible. But i think the point remains
- deleted 2mo ago[deleted]
- dyl000 3mo ago1 bit would be cool!