3 ms·
Large LLMs on MacBook produce tokens at an acceptable speed but the problem is reading context. Not incremental reading like when you have a chat session, becau
by visarga 5mo ago
Large LLMs on MacBook produce tokens at an acceptable speed but the problem is reading context. Not incremental reading like when you have a chat session, because they use KV cache, but large size reading, like when you paste a big file. It can take minutes.
- bel8 5mo agoAnd unless I'm mistaken, the repo is about running it with 2bit quantization. This is probably far from the raw intelligence provided by cloud providers. Still, this shines more light on local LLMs for agentic workflows.
- antirez 5mo agoIt runs both q2 and original (4 bit routed experts). At the same speed more or less. The q2 quants are not what you could expect: it works extremely well for a few reasons. For the full model you need a Mac with 256GB.
- someone13 5mo agoOut of curiosity, do you have any theories of why it works so well at such aggressive quantization levels?
- antirez 5mo agoIt's a mix of extreme sparsity but with the routed expert doing a non trivial amount of work (and it is q8), and projections and routing not being quantized as well. Also the fact it's a QAT model must have a role I guess, and I quantized routed experts out layers with Q2 instead of IQ2_XXS to retain quality.
- happyPersonR 5mo agoNot trying to give anyone homework thinking out loud : One thing I would love to see is if this dogfoods itself Like would dsv4 with q2 be able to do this task itself on this hardware ? Sidenote: I wish I had a M4-m3 … thinking about getting a ASUS ROG Flow Z13 Gaming Laptop (Model GZ302EA-XS99) uses pcie 4.0 so disk might be a little slower, but I want to see how this does on like Vulcan :)
- antirez 5mo agoDS4 can process 460 prompt tokens per second. Not stellar but not so slow. On M3 max. See the benchmarks on readme.
- brcmthrowaway 5mo agoWhy is this the case? Are there any architectures that don't rely on feeding the entire history back into the chat? Recurrent LLMs?
- habosa 5mo agoCan you ELI5 why this is so slow for local inference but so fast for using hosted models?