4 ms·
Well yes, but not so large that it's completely prohibitive. People have been running the full models on computers going as low as $6000: https://x.com/carrigma
by roblabla 2y ago
Well yes, but not so large that it's completely prohibitive. People have been running the full models on computers going as low as $6000: https://x.com/carrigmat/status/1884244369907278106 https://x.com/carrigmat/status/1884244369907278106
Of course this is for a personal instance, you'd need a much more expensive setup to handle concurrent users. And that's to run it, not train it.
- plagiarist 2y agoSortof a letdown that after 24 32Gb RAM sticks you only get 6-8 tokens per second.
- menaerus 2y ago6k is not that bad considering that top of the line Apple laptop costs as much. However, I don't have X so unfortunately I can't read the details.
- longitudinal93 2y agoYou can read the whole thread through nitter: https://xcancel.com/carrigmat/status/1884244369907278106 https://xcancel.com/carrigmat/status/1884244369907278106
- fspeech 2y agoA better approach is to split the model with MOEs running on CPUs and MLAs running on GPU. See the ktransformers project: https://github.com/kvcache-ai/ktransformers/blob/main/doc/en/DeepseekR1_V3_tutorial.md https://github.com/kvcache-ai/ktransformers/blob/main/doc/en... This takes advantage of the sparsity of MOE and the efficient KV-cache of MLA.
- menaerus 2y agoYou perhaps forgot to mention that for their AMX optimizations to be even feasible you'd need to spend ~$10k for a single CPU, let alone the whole system which is probably ~$100k.
- phonon 2y agoGranite Rapids-W (Workstation) is coming out soon for likely much less than half that per CPU. (Xeon W-3500/2500 launched at $609 to $5889 per CPU less than a year ago and also has AMX).
- menaerus 2y agoPoint being? Workstations that are fresh on the market and which have comparable performance of the server counterparts still easily cost anywhere between $20k and $40k. At least this is according to Dell workstations last time I looked.
- phonon 2y agoSupermicro X13SWA-TF Motherboard (16 DIMM slots with Xeon W-3500)= ~$1,000 E-ATX case = ~$300 Power Supply= ~$300 Xeon W-3500 (8 channel memory) = $1339 - $5889 Memory = $300-$500 per 64GB DDR5 RDIMM Memory will be the major cost. The rest will be around $5,000. A lot less than "$100,000"!
- menaerus 2y agoI acknowledged in my last comment that the cost doesn't have to be $100k but that it would still be very high if you opted for the workstation design. You're gonna need to add one more CPU to your design, add another 8 memory channels, beefier PSU, and a new motherboard that can accommodate this. So, 8k (memory) + 10k (cpus) + the rest. As I said, not less than $20k.
- phonon 2y ago
- mechagodzilla 2y agoI have a used workstation I got for $2k (with 768GB of RAM) - using the Q4 model, I can get about 1.5 tokens/sec and use very large contexts. It's pretty awesome to be able to run it at home.
- MysticFear 2y agoWould love to know more info & specs of your workstation.
- mechagodzilla 2y agoIt's an HP Z8 G4 (dual-socket 18-core, 3 GHz Xeons, 24x32GB of DDR4-2666, and then a crappy GPU, 8TB HDD, 1TB SSD). It can accommodate 3 dual-slot GPUs, but I was mostly interested in playing with frontier models where holding all the weights in VRAM requires a ~$500k machine. It can run the full Deepseek R1, Llama3-405B, etc, usually around 1-2 tokens/sec.
- nomel 2y agoFor me, where electricity is $0.45/kWh, assuming 1kW consumption, it would be around $80 USD/million!
- CyberDildonics 2y agoI think you might have to show your math on that one.
- nomel 2y agoThey said 1.5 tokens/second. 1 mil tokens is 667k seconds is 185 hours per million tokens. 1kW * 185hr * $0.45/kWh = $80 per million tokens. Again, assuming 1kW, which may be high (or low). The cost of the physical computation is electricity cost.
- CyberDildonics 2y ago