3 ms·
i have a 512gb ram m3 ultra mac studio setup with a gas city that runs one of my companies. today was the first time ever that a local model (GLM5.3 8-bit) was
by taylorhou 1mo ago
i have a 512gb ram m3 ultra mac studio setup with a gas city that runs one of my companies. today was the first time ever that a local model (GLM5.3 8-bit) was able to match fable5 in our tests.
GLM-5.3-Flash at true 8-bit: 341 GB on disk, 328 GB resident, 288 experts across 46 layers, loads in 65 seconds.
• 18.7 tokens/s generation, 35 tokens/s prompt, on a desk, on a $0 per-token bill.
• Runs beside our whole agent city on one box with ~130 GB to spare.
• Review test: caught 6 of 6 planted P1 defects, zero false positives, same score as the frontier model we pay for.
• CRM test: 11 of 11 required records extracted, zero wrong writes, 45 minutes, first local model to clear the bar.
• Serving a 131k-token window today; the model itself supports 1,048,576. Widened to 4 concurrent slots and still have 50gb+ of excess ram.
granted my cto still isn't moving all of our inference to glm5.3 but we've identified 40%+ that is currently handled by fable that we're routing locally instead and will do concurrent requests to verify/compare responses for a while.
- aa-jv 1mo agoWhat sort of business can you run with this setup?
- Normal_gaussian 1mo ago$0 per-token bill You still have electricity and capital investment. Envelope math suggests cheap electricity is costing you something like $0.50/mtok and the opportunity cost on the capital tied up and lost in the unit purchase and resale is going to cost you something like $2/mtok at 100% utilization (so, frontier model prices or higher at real utilization), and you don't benefit from any elasticity. Hosted GLM 5.3 flash is like $0.15/mtok in $0.50/mtok out
- icedchai 1mo agoTime to completion also must be considered. If I have to wait around for hours for a prompt to complete locally and I’ll need to iterate quickly, I’m better off hosted than local. If it’s “free” and slow it may just not be worth it.
- icedchai 1mo agoThis may work for your use case, but sounds abysmally slow for any complex coding task.
- vintagedave 1mo agoSo this is something like a $10,000 machine before RAM prices rose? I see Apple is currently selling a 256GB M5 for about $10K, so buying October's 512GB one could be, what, $13-14K? A $0 per-token bill is great but this is clearly not something for normal people, just some businesses.
- bel8 1mo agoProps to your parent commenter for including context size. Because 131k context window is prohibitively small for my coding workloads so I know a 512GB Mac won't cut it.