3 ms·
So GLM 5.2/Gemini 3.6 level intelligence for $0.28/m output. And their updated Pro model coming soon.... Plus a size you can genuinely run at home: Unsloth los
by scosman 2mo ago
So GLM 5.2/Gemini 3.6 level intelligence for $0.28/m output. And their updated Pro model coming soon....
Plus a size you can genuinely run at home: Unsloth lossless Q8 at 162GB.
- segmondy 2mo agoQ8 is ~ 151gb
- scosman 2mo agothanks, corrected! I looked at Q3
- luckydata 2mo agoI would like to see your "home"
- cmrdporcupine 2mo agoTwo (linked) DGX Sparks would do it I guess. Though probably slowly (I'd guess 15-20 tok/sec for decode, but higher for prefill). So ~$8-9k USD at current RAM prices, substantially less if they ever (sigh) drop. Electricity use would actually be relatively modest. But it makes little to no sense as long as API prices are what they are. Except for maybe privacy reasons.
- ycui7 2mo agothe rational in one’s mind is similar to buying expensive supercar but no driving it daily. owning a few GPUs is a lot cheaper than supercars.
- cmrdporcupine 2mo agoI dunno. I bought the Spark in January and it has led indirectly to paid work. I don't use it for local inference so much. I use it to learn. I also use it as my daily driving Aarch64 development system. Aside it's also very cool what else can be done with unified GPU memory, once you realize you have it...
- bethekind 2mo agoLearn model deployment, batching, all the AI inference related things?
- vardalab 2mo agoIt's at least 2.5x that speed for dual sparks and prefill is good as well. Basically going on vibes it is faster seeming than what one gets by default with openAI or Anthropic.
- wolttam 2mo ago2 sparks currently run this model at 60 t/s single session, up to just over 100 t/s aggregate with concurrency of 4. Going local has as opened up a world of use-cases I never would have entertained the idea of on metered/cloud usage. Privacy is a large part of it but, I also no longer think twice about whether to send a prompt or not based on the psychology of it costing money. Cached input tokens on local inference are free, so I don’t care about running sessions up to 500k tokens and hundreds of turns (it’s rarely useful, but DSv4 remains surprisingly coherent up there)
- perdenie 2mo ago[dead]
- sourcecodeplz 2mo agoi am scared for the PRO model maybe it is Fable level