3 ms·
The most cost-effective way is arguably to run it off of any 1TB SSD (~$55) attached to whatever computer you already have. I was able to get 1 token every 6 o
by coder543 2y ago
The most cost-effective way is arguably to run it off of any 1TB SSD (~$55) attached to whatever computer you already have.
I was able to get 1 token every 6 or 7 seconds (approximately 10 words per minute) on a 400GB quant of the model, while using an SSD that benchmarks at a measly 3GB/s or so. The bottleneck is entirely the speed of the SSD at that level, so an SSD that is twice as fast should make the model run about twice as fast.
Of course, each message you send would have approximately a 1 business day turnaround time… so it might not be the most practical.
With a RAID0 array of two PCIe 5.0 SSDs (~14GB/s each, 28GB/s total), you could potentially get things up to an almost tolerable speed. Maybe 1 to 2 tokens per second.
It’s just such an enormous model that your next best option is like $6000 of hardware, as another comment mentioned, and that is probably going to be significantly slower than the two M2 Ultra Mac Studios featured in the current post. It’s a sliding scale of cost versus performance.
This model has about half as many active parameters as Llama3-70B, since it has 37B active parameters, so it’s actually pretty easy to run computationally… but the catch is that you have to be able to access any 37B of those 671B parameters at any time, so you have to find somewhere fast to store the entire model.