10 ms·
DeepSeek 4 Flash local inference engine for Metal
- m00dy 5mo ago[dead]
- maherbeg 5mo agoThis is so sick. I'm really curious to see what focused effort on optimizing a single open source model can look like over many months. Not only on the inference serving side, but also on the harness optimization side and building custom workflows to narrow the gap between things frontier models can infer and deduce and what open source models natively lack due to size, training etc.
- dakolli 5mo agoThere will always be a huge gap between frontier models and open source models (unless you're very rich). This whole industry makes no sense, everyone is ignoring the unit economics. It cost 20k a month to running Kimi 2.6 at decent tok/ps, to sell those tokens at a profit you'd need your hardware costs to be less 1k a month. Everyone who's betting their competency on the generosity of billionaires selling tokens for 1/10-1/20th of the cost, or a delusional future where capable OS models fit on consumer grade hardware are actually cooked.
- bensyverson 5mo agoIf you looked at a graph of GPU power in consumer hardware and model capability per billion parameters over time, it seems inevitable that in the next few years a "good enough" model will run on entry-level hardware. Of course there will always be larger flagship models, but if you can count on decent on-device inference, it materially changes what you can build.
- physicsguy 5mo agoIt also massively changes the value economics of the frontier models. In a lot of cases, you really don't need a general purpose intelligence model too.
- bensyverson 5mo agoExactly… as hn readers, we sometimes forget that a lot of people are using these tools to search for the best sunscreen, or rewrite an email.
- dakolli 5mo ago[flagged]
- afro88 5mo agoNo offense, this is a crazy worthless contribution to the discussion. Why?
- dakolli 5mo agoBecause everyone in these replies is in complete denial about the physical limits of memory and scaling in general. Ya'll literally living in an alternate reality where model capability increases with a decrease in size, its simply not the case. There will be small focused models that preform well on very narrow tasks, yes, but you will not have "agents" capable of "building most things" running on consumer hardware until more capable (and affordable) consumer hardware exists.
- bensyverson 5mo agoAh, you haven't realized that consumer hardware gets more capable over time
- adrian_b 5mo agoNot this year, when many vendors either offer lower memory capacities or demand higher prices for their devices.
- bensyverson 5mo agoCorrect, the progress is not perfectly linear. But do you believe technological progress has stalled forever? If so, I'd get out of tech and start selling bomb shelters.
- dakolli 5mo agoDo you really think the trend of consumer hardware is heading towards more memory and better specs? Apple's most popular product this year is an 8gb of RAM laptop.. The trend is heading in the opposite direction, less options for strong consumer hardware and towards cloud based products. This is a memory issue more than anything. Nvidia is done selling their ddr7 to gamers and people with AI girlfriends.
- otabdeveloper4 5mo ago> a delusional future where capable OS models fit on consumer grade hardware 48 gb is enough for a capable LLM. Doing that on consumer grade hardware is entirely possible. The bottleneck is CUDA and other intellectual property moats.
- liuliu 5mo agoI am not sure where this comment is from (possibly without looking at this project?). This project is running quasi-frontier model at reasonable tps (~30) with reasonable prefill performance (~500tps) with a high-end laptop. People simply project what they see from this project to what you optimistically can expect. You can argue whether the projection is too optimistic or not, but this project definitely made me a little bit optimistic on that end.
- amunozo 5mo agoMost tasks do not require frontier models, so as long as these models cover 95-99 per cent of the tasks, closed frontier models can be left for niche and specialized cases that are harder.
- maherbeg 5mo agoThere will always be a gap, but what's interesting is that because new models are constantly coming out, we as an industry never spend any time extracting the maximal value out of an existing model. What if there are techniques, and harness workflows that could be optimized for a singular model end to end? How far can that push the state of the art. An example is https://blog.can.ac/2026/02/12/the-harness-problem/ https://blog.can.ac/2026/02/12/the-harness-problem/ for just improving edits. Or if we could really steer these open source models using well structured plans, could we spend more time planning into a specific way and kick off the build over night (a la the night shift https://jamon.dev/night-shift https://jamon.dev/night-shift)
- daveguy 5mo ago> There will always be a huge gap between frontier models and open source models (unless you're very rich). They said the same thing about open source chess engines.
- shay_ker 5mo agoHow does it compare to popular local inference engines, e.g. ollama, lm studio, or handrolled llama.cpp? I saw a brief benchmark in the readme but wasn't sure if there was more.
- speu 5mo agoI've been trying deepseek-v4-flash in OpenCode (via OpenRouter) and I'm blown away. It's no Opus, obviously, but it had zero issues with any regular coding task I threw at it. v4-flash is remarkably "good enough" for what I needed. The whole evening of coding cost me $0.52 in API credits.
- jiehong 5mo agoUsing it in Kagi Assistant is stupidly slow. I get like 10 t/s. While it’s pretty fast in the official app for example. Kagi Assistant is also kind of broken when using Qwen 3.6 Plus. So, beware of using them in Kagi at the moment.
- dev_hugepages 5mo agoProbably a provider thing. Looking at https://help.kagi.com/kagi/ai/llms-privacy.html https://help.kagi.com/kagi/ai/llms-privacy.html, they're using deepinfra. Looking at https://openrouter.ai/deepseek/deepseek-v4-flash/providers https://openrouter.ai/deepseek/deepseek-v4-flash/providers tells us that the deepseek provider achieves 49tps of throughput while deepinfra 19tps.
- jiehong 5mo agoThanks for taking the time to provide this info. I appreciate it
- amunozo 5mo agoI am curious about it producing less tokens except for the max mode. I love DeepSeek V4 Flash and I use it extensively, it's so cheap I can use it all day and still not use all my 10$ OpenCode Go subscription. I use it always in max mode because of this, but now I wonder whether I should rather use high.
- unshavedyak 5mo agoWhat do you use it for? I tend to just stick to SOTA (Claude 4.7 Max thinking), and put up with the slow req/response. I'm not sure what type of work i'd trust a less thinking model, as my intuition is built around what Claude vSOTA Max can handle. Nonetheless eventually i want to build an at-home system. I imagine some smaller local model could handle metadata assignment quite well. edit: Though TIL Mac Studio doesn't offer 512GB anymore... DRAM shortage lol. Rough.
- amunozo 5mo agoI am experimenting with some game development and my thesis' beamer. I have a 20$ Codex account and I use GPT-5.5 for planning and DeepSeek for executing in OpenCode. This makes my Codex 5h tokens to last more than 10 minutes.
- actsasbuffoon 5mo agoApple just dropped the 128GB option as well.
- fgfarben 5mo agoIt is still available for the M5 Max Macbook Pro, but yes, the Mac Studio is now only offered with up to 96 GB.
- syntaxing 5mo agoHow has opencode go been for you? Worth changing over from Claude pro?
- 5mo ago
- antirez 5mo agoA random, funny, interesting and telling data point: my MacBook M3 Max while DS4 is generating tokens at full speed peaks 50W of energy usage...
- bertili 5mo agoequals 2 or 3 human brains in power usage. Amazing work!
- antirez 5mo agoTrue quantitatively, not qualitatively. DeepSeek V4 is not capable of doing what a human brain can do, of course, but for the tasks it can do, it can do it at a speed which is completely impossible for a human, so comparing the two requires some normalization for speed.
- scotty79 5mo agoI'm sure human brain, at least my present brain, is incapable of many things DeepSeek V4 can do. Qualitatively.
- minimaxir 5mo ago"Data centers for LLMs are technically more energy efficient per-user than self-hosting LLM models due to economies-of-scale" is a data point the internet isn't ready for.
- Onavo 5mo agoThere's a bunch of companies doing garage GPU datacenters now. Probably can act as a heat source during winter too if you have a heat pump.
- kristianp 5mo agoThat's an interesting idea [1], the value being that its easier to build servers into a bunch of homes that are being built than building a datacenter. Every now and then something reminds me of "Dad's Nuke", a novel by Marc Laidlaw, about a family that has a nuclear reactor in their basement. A really bizarre, memorable satire [2]. [1] https://finance.yahoo.com/sectors/technology/articles/nvidia-wants-next-house-mini-171222508.html https://finance.yahoo.com/sectors/technology/articles/nvidia... [2] https://en.wikipedia.org/wiki/Dad%27s_Nuke https://en.wikipedia.org/wiki/Dad%27s_Nuke
- happyPersonR 5mo agoSo just gonna ask a question, probably will get downvoted I know this is flash, but…. But other than this guy, did our whole society seriously never flamegraph this stuff before we started requesting nuclear reactors colocated at data centers and like more than 10% of gdp? Someone needs to answer because this isn’t even a m4 or m5… WHAT THE FUCK Sidenote: shout out antirez love my redis :)
- liuliu 5mo agoDSv4 generates much faster on NVIDIA class hardware. It is just a very efficient model.
- AlotOfReading 5mo agoThis is built atop a tower of stuff people built with profiling and performance-oriented design. That said, I've found that most corporate environments are unintentionally hostile to this kind of optimization work. It's hard to justify until the work is already done. That means you often need people with the skills, means, and motivation to do this that are outside normal corporate constraints. There aren't many of those.
- happyPersonR 5mo agoBuilding this into agentic dev workflows (subject to token/time constraints) is something I spent a lot of time doing at work. I actually am kind of proud of that hahah But you’re right I agree In the corporate world they sadly don’t take kindly to performance profiling as a first class citizen Granted I will say optimization without requirements may not be beneficial but at least profiling itself seems worthy if you have use cases. A lot of us have been working in the network packet pusher software , distributed systems , distributed storage space I’m happy to see more stuff like this :) TLDR; I’ve not seen a lot of flamegraphs of Llm end to end … idk if anyone else has?
- fgfarben 5mo agoThe world is not China.
- 5mo ago
- nazgulsenpai 5mo agoI keep seeing DS4 and in order my brain interprets it as Dark Souls 4 (sadface), DualShock 4, Deep Seek 4.
- throwaway613746 5mo ago[dead]
- sourcecodeplz 5mo agoGreat project! This is also a fine example of a vibe-coded project with purpose, as you acknowledged.
- kgeist 5mo agoHeh, I made something very similar for the Qwen3 models a while back. It only runs Qwen3, supports only some quants, loads from GGUF, and has inference optimized by Claude (in a loop). The whole thing is compact (just a couple of files) and easy to reason about. I made it for my students so they could tinker with it and learn (add different decoding strategies, add abliteration, etc.). Popular frameworks are large, complex, and harder to hack on, while educational projects usually focus on something outdated like GPT-2. Even though the project was meant to be educational, it gave me an idea I can't get out of my head: what if we started building ultra-optimized inference engines tailored to an exact GPU+model combination? GPUs are expensive and harder to get with each day. If you remove enough abstractions and code directly to the exact hardware/model, you can probably optimize things quite a lot (I hope). Maybe run an agent which tries to optimize inference in a loop (like autoresearch), empirically testing speed/quality. The only problem with this is that once a model becomes outdated, you have to do it all again from scratch.
- joshmarlow 5mo agoAnother suggestion for optimizing local inference - the Hermes team talks a lot on X about how much better results are when you use custom parsers tuned to the nuances of each model. Some models might like to use a trailing `,` in JSON output, some don't - so if your parser can handle the quirks of the specific model, then you get higher-performing functionality.
- mirsadm 5mo agoI've built something like this. One issue is that LLMs are actually terrible at writing good shaders. I've spent way too much time trying to get them not to be so awful at it.
- wahnfrieden 5mo agoJust curious if you've tried GPT 5.5 Pro?
- davidwritesbugs 5mo agoI tried getting any sota llm (GPT 5, Opus 4.6, Deepseek V4 pro, glm-5) to write a Metal 4 shader for a bottle usdz and none of them got it right. They screwed up the normals and textures , total mess. I tried it to do it in Metal 3 and still crappy.
- visarga 5mo agoLarge LLMs on MacBook produce tokens at an acceptable speed but the problem is reading context. Not incremental reading like when you have a chat session, because they use KV cache, but large size reading, like when you paste a big file. It can take minutes.
- bel8 5mo agoAnd unless I'm mistaken, the repo is about running it with 2bit quantization. This is probably far from the raw intelligence provided by cloud providers. Still, this shines more light on local LLMs for agentic workflows.
- antirez 5mo agoIt runs both q2 and original (4 bit routed experts). At the same speed more or less. The q2 quants are not what you could expect: it works extremely well for a few reasons. For the full model you need a Mac with 256GB.
- someone13 5mo agoOut of curiosity, do you have any theories of why it works so well at such aggressive quantization levels?
- antirez 5mo agoIt's a mix of extreme sparsity but with the routed expert doing a non trivial amount of work (and it is q8), and projections and routing not being quantized as well. Also the fact it's a QAT model must have a role I guess, and I quantized routed experts out layers with Q2 instead of IQ2_XXS to retain quality.
- happyPersonR 5mo agoNot trying to give anyone homework thinking out loud : One thing I would love to see is if this dogfoods itself Like would dsv4 with q2 be able to do this task itself on this hardware ? Sidenote: I wish I had a M4-m3 … thinking about getting a ASUS ROG Flow Z13 Gaming Laptop (Model GZ302EA-XS99) uses pcie 4.0 so disk might be a little slower, but I want to see how this does on like Vulcan :)
- brcmthrowaway 5mo agoHow does this compare with oMLX?
- Havoc 5mo agoWas excited until I realized DS flash is still enormous. Oh well...glad it exists anyway & happy to see antirez still doing fun stuff
- zozbot234 5mo agoIt could run viably with SSD offload on Macs with very little memory. You could even exploit batching to make the model almost compute limited even in that challenging setting, seeing as the KV cache is so extremely small (for non-humongous context). In fact, if that approach can be made to work I'd like to see a comparison between DS4 Flash and Pro on the same (Mac) hardware.
- Havoc 5mo ago>It could run viably with SSD offload on Macs with very little memory Not really. That's going to land you somewhere in the 0.2-0.5 tokens a second range Lovely as modern nvmes are they're not memory
- zozbot234 5mo agoYou can run multiple inferences in parallel on the same set of weights, that's what batching is. Given enough parallelization it can be almost entirely compute-limited, at least for small context (max ~10GB per request apparently, but that's for 1M tokens!)
- happyPersonR 5mo agoYes I think what this demonstrates that folks are missing is that now optimization for specific scenarios is quite possible.
- Havoc 5mo agoFor offline work that's fine I guess, but batched or not <1tks is largely unusable for most usage cases
- andrefelipeafos 5mo ago[flagged]
- layoric 5mo agoVery impressive. One thing that seems odd to me is that is at like 4 minutes before it starts a response for large input? I don't use mac hardware for LLMs, but that is quite surprising and would seem to be a pretty large stumbling block for practical usage. Edit: Caching story makes a lot more sense for regular usage: > Claude Code may send a large initial prompt, often around 25k tokens, before it starts doing useful work. Keep --kv-disk-dir enabled: after the first expensive prefill, the disk KV cache lets later continuations or restarted sessions reuse the saved prefix instead of processing the whole prompt again.
- antirez 5mo agoYep that happens with coding agents sending a very large system prompt. And also when later tool calling feed it large files or diffs. But with the M3 ultra the prefill speed is almost 500 t/s that is quite into the very usable zone. With M3 max you need a bit more patience but it works well and as it emits the think process if you use the pi agent you don't wait: you read non censored chain of though. I posted a video on X yesterday using it with my m3 max. It spills tokens at a decent speed.
- zozbot234 5mo agoGiven how small the KV cache for this model seems to be for small contexts, can you clarify how the engine behaves if you try to run increasingly larger batches on your prosumer hardware (RAM 128 GB)? Does it eventually become compute limited? Also, can the engine support transparent mmap use for fetching weights from disk on-demand, at least when using pure CPU? (GPU inference might be harder, since it's not clear how page faults would interact with running a shader.) If the latter test is successful, next would be testing Macs with more limited RAM, first running simple requests (would be quite slow) then larger batches (might be more worthwhile if one can partially amortize the cost of fetching weights from storage, and be bottlenecked by other factors).
- segmondy 5mo agoCurious why you went this route, don't you think you could have achieved near this performance 80%+ or more within llama.cpp?
- lhl 5mo agoI think especially with the ability for SOTA AI to optimize kernels more people should try their hand at making better inference for their specific hardware. I have an older W7900 (RDNA3) which, besides 48GB of VRAM, has some pretty decent roofline specs - 123 FP16 TFLOPS/INT8 TOPS, 864 GB/s MBW, but has had notoriously bad support both from AMD (ROCm) as well as llama.cpp. Recently I decided I'd like to turn the card into a dedicated agentic/coder endpoint and I started tuning a W8A8-INT8 model. Over the course of a few days of autolooping (about 800 iterations using a variety of frontier/SOTA models, Kimi K2.6 did surprisingly well), and I ended up with prefill +20% and decode +50% faster than the best llama.cpp numbers for Qwen3.6 MoE. I'm currently grinding MTP and DFlash optimization on it, but I've been pretty pleased with the results, and will probably try Gemma 4 next.
- throwa356262 5mo agoPlease share your knowledge and your findings I think llama.cpp could have done a much better job supporting PC. Sure, some of it us due to bad vendor support but with so many users I am surprised we don't see more optimized inference on standard PCs
- lhl 5mo agoWhen it's in a good state I'll open source it, I am keeping track of what optimizations make the most impact, stuff like this: ### Diagnosing parallelism pathologies (L1) *Grid occupancy:* - `Grid_Size / Workgroup_Size >= CU count` (W7900 = 96, Strix Halo = 40)? - < 0.3 = massively undersubscribed. Fix grid FIRST. Micro-optimization will NOT help. - 0.3-1.0 = partially utilized; depends on VGPR/LDS pressure. - 1.0-4.0 = healthy; micro-optimization can help. *Within-block distribution:* - Does the kernel do useful work across all threads, or is there an `if (threadIdx.x == 0)` gate around a serial top-k, reduction, or scan? For c=1 decode, many kernels can't grow the grid, but they can always parallelize inside the block. - `Scratch_Size > 0` from dynamically-indexed per-thread arrays is a strong secondary signal of the within-block pathology. *Router top-k (within-block fix)*: - Kernel: `qwen35_router_select_kernel` @ c=1 decode - Before: grid=1 (can't help; num_tokens=1), blockDim=512, `if (threadIdx.x == 0)` gated 2048 serial compares. Scratch=144 B from spilled per-thread arrays. - Fix: warp-shuffle parallel argmax across the whole block + `__shared__` top_vals buffer eliminating the spill. - Result: 5.7× kernel speedup, +6.6% on 4K/D4K E2E.
- deleted 5mo ago[deleted]
- micalo 5mo ago[flagged]
- dejli 5mo agoThe beaty of it, that you can clone and make it, and it just works, no python shenanigans, what a blessing for this eco system.
- octocop 5mo agoFinally someone who pays proper respect to GGML ecosystem.
- danborn26 5mo ago[flagged]
- mudkipdev 5mo agoAI bot
- shivnathtathe 5mo agoBeen working on local-first LLM observability for exactly this use case — tracing local model pipelines without sending data to cloud. Happy to share if anyone's interested.
- kristianp 5mo agoHmm, I'm unable to order more than 96GB RAM for a Mac studio, even with the M3 ultra or M4 Max. Is this au specific? However with the MacBook Pro I can specify 128GB with the M5 Mac. https://www.apple.com/au/shop/buy-mac/mac-studio https://www.apple.com/au/shop/buy-mac/mac-studio
- smcleod 5mo agoThe studio is really old now. The new one will drop at some point no doubt with more memory options. the 128GB M5 max MBP is great though
- Terretta 5mo agoAnd yet, aside from offering 512GB, that really old Studio Ultra M3 LLMs faster (especially sustained) than the new M5 Max.
- Joeri 5mo agoIt’s not just AU: https://9to5mac.com/2026/05/05/apples-most-powerful-mac-studio-loses-its-last-remaining-ram-upgrade-option/ https://9to5mac.com/2026/05/05/apples-most-powerful-mac-stud... They’ve dropped all the mac studio configs higher than 96 gb, as well as the base mac mini. They’re also rumored to be considering taking the Neo base config off the market. This seems to be how they’re dealing with supply constraints for fab capacity and RAM.
- Terretta 5mo agoDifficult to believe this memory is made of unobtanium. Maybe Apple would rather not price it at all than experience blowback for either gouging or lack of inventory.
- ZeroGravitas 5mo agoDid I miss a simple motivating benchmark or goal? I'm assuming this is faster, and/or lets you run a bigger, smarter model than just using the generic tool chain, but it doesn't spell out the level of existing improvements over that baseline or expected improvements as far as I can see? Presumably you can work it out based on the numbers given if you have the relevant comparison values.
- npgraph 5mo agoAny direct TPS comparison to Ollama?
- zozbot234 5mo agoOllama has no local support for DeepSeek V4 at present; it's only listed as a cloud model. Even llama.cpp is still waiting for support.
- JDevAlper 5mo ago[dead]
- sev_verso 5mo agoI've tried it out with Claude Code on my existing codebase and it seemed to hold its weight (despite being the 2-bit quant). Takes minutes on prompt processing, the actual edits are reasonably quick at above 20 tks. The good: It succeeded with discovering, applying edits and writing a test for a small task I gave it. The bad: It could not address a small nitpick I had. The ugly: It hallucinated a conversation about "The Duck" that I had with it simultaneously while trying to solve another problem. I can only imagine it's one of examples in the initial Claude Code prompt: --cut-- However, the user's query is "Can you track these 3 videos here?" which seems unrelated. Perhaps the user is asking if I can track the progress of three videos they are working on? Let me re-read the user's message. The user said "Source Code" and "The Agent" and "The Duck", it could be video titles. And they are asking if I can track these 3 videos. ?? That doesn't make sense in the context. Could there be two different conversations? --cut--
- tmaly 5mo agoThe intro was the best part of the README in my opinion. The rest of the README looks and feels AI generated. I am guilty of this same thing with README files.
- fgfarben 5mo agoOn both the llama.cpp based version and the custom Metal version, the model forgets how to use tools somewhere around the 50,000 token mark.