4 ms·
Kimi K3 (2.8T) at 1 token/s on a MacBook Pro, streamed from four SSDs
- Argonautlabs 26d agoAuthor here. Some context and the caveats up front. The model is Kimi K3, 2.78T parameters, ~1.45 TB of expert weights. It does not fit in memory, so the experts stream from disk: one 17.5 MB file per (layer, expert), read with pread + F_NOCACHE, 16 of 896 per layer. The machine is an M5 Max MacBook Pro with 128 GB and three Thunderbolt 5 enclosures plus the internal SSD. Expert weights are untouched at their released MXFP4 precision; the resident attention trunk is int8, which upstream labels non-weight-exact, so I don't claim bit-exactness against BF16 — I claim token-identical output against my own reference on the prompt of record, checked on every promotion. Numbers, with the unflattering ones in the same paragraph as the good ones: 1.00 tok/s steady over a 512-token completion, 1.13 over 128 tokens, and 0.96 median on the 17-token benchmark from the upstream repo's issue #15 against the 0.684 posted there. Time to first token on a 512-token prompt is about 6.3 minutes — prefill is currently read-amplified 6.2x, which is the biggest open problem in the repo and is described in the results directory. What I think is actually interesting isn't the number, it's that four of the gains came from defects in the read path that instrumentation found and I would never have guessed. The instruments are in a second repo, ARGODRIVE — a 10 ms per-device read monitor, a per-read barrier trace that records which drive served each expert and which one landed last in every pass, and a config assertion harness that refuses to record a benchmark unless the setting under test actually fired. They're deltafin-specific today. The four findings: • A constant capped the reader threads at 16 and bounded both the demand and prefetch pools with the same value. Separating them was +14%; demand queueing went from 70% of blocked time to 7.5%. • Splitting each hot expert's read across two replicas on two devices was +10% — after the same knob had measured negative six times on layouts where every expert had one home and there was nothing to split against. • The prefetch path had no balancer at all: it walked a fixed directory order and took the first hit, so on any replicated layout it dumped everything on one enclosure. Giving it least-expected-completion dispatch with in-flight counters shared with the demand path was +11% and turned every replicated layout I had previously measured as a loss into a win. • A recorded "law" that a given draft depth was worse turned out to have been measured against a drafter that no longer existed. Re-testing it was +8%. There's also a drive-count ladder in the repo — same layout, one to four drives: 57% / 78% / 92% / 100% of the four-drive decode rate. And a catalogue of about a thousand timed runs of things that did not work, with the numbers: RAM expert caches from 8 to 40 GB (-4% to -48%), striping a single copy (-7 to -25%), two drives sharing one Thunderbolt link (-11%), streaming the attention trunk from SSD (-60%), Metal's file-loading API (-19 to -22%). That catalogue is the part I expect to be most useful to other people. The engine is a fork of gavamedia/deltafin, which is MIT and did the hard part; I've told the author about all of this and the upstream-relevant fixes are going back as PRs. Two things I'd genuinely like help with: whether anyone has done expert-major prefill scheduling on an MoE (read each expert once per layer and run its kernel over all rows routed to it — it should take prefill from 6.2x amplification to about 1x), and whether the drive ladder reproduces on other hardware.
- sampullman 26d agoThis is difficult to read, maybe just link to a gist?
- woadwarrior01 26d agoThat's because it's copy pasted from a coding agent.
- anamexis 26d agoIt's difficult to read because it doesn't have line breaks.
- sampullman 26d agoIt looks at least partially hand edited to me, although it's getting pretty difficult to tell with Astra...
- Argonautlabs 26d agoThank you! Here is the short version: Kimi K3, 2.78T parameters, ~1.45 TB of MXFP4 experts streamed from four SSDs on an M5 Max / 128 GB. 1.00 tok/s steady over 512 tokens, 1.13 over 128, ~6.3 min to first token on a 512-token prompt. Output token-identical drafter on/off on a given drive layout; the int8 trunk is non-weight-exact per upstream. The useful bits: one drive gives ≈52% of four, two ≈73%, three ≈90%; and prefill is slow because of ~9 TB of reads for a 1.4 TB model — a scheduling bug with a planned fix. README with per-run logs: github.com/argonautlabsai/deltafin — a fork of gavamedia/deltafin, who built the engine.
- voidnullvalue 26d agoBut why though? Cannot possibly be useful at such slow speeds, and costs a ton to perform that badly
- Argonautlabs 26d agoNot useful for chat, agreed — and I wouldn't pretend otherwise. It's useful for the other kind of work: scheduled, unattended jobs where nobody is waiting on the cursor. My use is day/week/month end review — go through the numbers, flag what doesn't reconcile, draft the report — and there the two things that matter are that the model is good enough to trust with the judgement (K3 is, and it's the full 2.8T model, not a cut-down one) and that the data never leaves the machine.
- voiceeh 26d ago>and I wouldn't pretend otherwise. Such of a Claudism. Not criticizing, just noticing.
- ganelonhb 26d agoI think the point is that it’s running at all…
- cyanydeez 26d agoQwen3.8-Flash-Next ships with a 51B lookup table that can be read directly from ssd or memory, which greatly improves it's speed and intelligence. It can load at 4bit quant in ~60GB. These demos are maybe useless, but if open models keep progressing, there's going to be some break through that continues whittling down just how much needs to be kept in VRAM, and progressive degredation to regular system ram and to ssds. Afterall, they're not writing anything to these, so saturing all bandwidth could bring models to the masses. all without any help from Zark Muckerberg.
- Argonautlabs 26d agoIt actully does the job. Example: every morning it takes 30-40 minutes to generate reports automatically and these reports are being sent as a pdf to read to Telegram.
- willmadden 26d agoThat's next level masochism.
- netc 26d agoAnd macOSism
- willmadden 24d agoHa.
- lukeduff 26d agoReminds me of Deep Thought from Hitchhiker's Guide to the Galaxy
- schmorptron 26d agoIt's kind of insane how having this tech at this speed 5 years ago would have probably still been seen as insanely useful and revolutionary. If LLMs were more capable but dramatically slower, I wonder how it would impact how we use it? Dramatically more thought being put into prompts, much more preparation probably
- IgorPartola 26d agoMy parents learned to program on punch cards. They told me it was a day of preparing the program, an hour of running it, just to get a syntax error.
- deleted 26d ago[deleted]
- bluedino 26d agoWrite the program, punch the cards, send the cards to another building to be loaded, program runs, printout comes out in another building, somehow this takes 2-3 days
- lurker919 26d agoCoding is the new punch card slots now. My children will listen in awe about how typing and testing used to take hours or even (gasp!) days.
- redox99 26d agoAt 1t/s it's still faster than humans for a lot of tasks, basically doing overnight what could take humans half a week. Plus you can always parallelize.
- bluechair 26d agoI missed the explanation for how the SSDs are connected. Maybe a dumb question.
- Argonautlabs 26d agoSSDs are connected via Thunderbolt 5 enclosures. I have one Gen4 and three Gen5 ssds inside enclosures. You can see specs here https://github.com/argonautlabsai/deltafin/tree/main/k3-public-bench https://github.com/argonautlabsai/deltafin/tree/main/k3-publ...
- grosswait 26d agoSo if this all streamed from the internal NVME would you expect better performance?
- Argonautlabs 26d agoJust internal 2Tb Macbook M5 Max drive it came in at roughly half the four-drive speed (0.535 vs 1.038 tok/s at 128 tokens), since one fast drive still has to serve all 16 reads per layer while four drives split the load https://raw.githubusercontent.com/argonautlabsai/deltafin/main/k3-public-bench/results/charts/drive-draw.svg https://raw.githubusercontent.com/argonautlabsai/deltafin/ma...
- cousinbryce 26d agoThink the M7 will have some storage parallelization?
- _zoltan_ 26d agoWhat exact enclosure are you using?
- Argonautlabs 26d agoOWC Express 1M2 (Thunderbolt 5, single M.2 NVMe) — three of them, two on the Mac's own ports and one behind an OWC Thunderbolt 5 hub since the machine has three ports. Each enclosure tops out at about 7.1 GB/s on whole-file reads regardless of the drive inside (a 2 TB SN8100 measures the same as the 1 TB); the drive behind the hub reads 5.7 GB/s and falls with queue depth. Details in the README's hardware section
- dusted 26d agoA medium prompt in only 11 days.
- jgalt212 26d agoA medium prompt = 1 million tokens?
- RugnirViking 26d agoHow many times have you had a model start compacting already before getting back to you? Most have 1 million context window. It's happened to me occasionally
- jgalt212 26d agoI don't use agents. I like to code and AI with a REPL.
- pvab3 26d agowhen it finishes answering you already figured out the question
- meerita 26d ago3600 words in one hour.
- saejox 26d agomake it 4x40 raid-0 ssds to achieve 40 tps. or 40 macbooks with each 4 ssd. to get 40 tps.
- Argonautlabs 26d agoBandwidth doesn't multiply like that here, and we measured it rather than assumed it. A MoE layer needs 16 expert reads and can't proceed until the slowest one lands, so a layer costs the max over its reads, not the sum. Going from one drive to four (13.6 → ~33 GB/s of combined ceilings) took decode from ~52% to 100% of our number — not 4× — with Every drive already at 90–100% of its own ceiling. RAID-0 was one of the first things tried and it lost: striping makes every read touch every drive, so the slowest drive sets every barrier. What moves this is per-read latency and read scheduling, and for long prompts not re-reading each layer's experts eight times. Numbers in results/SCALING.md and results/PREFILL.md.
- lowbloodsugar 26d agoWould the 40 Mac’s work with pipelining though?
- adrian_b 26d agoNo matter how many external drives you gather, the data coming from them must be squeezed through the peripheral interfaces of the Apple SoC. So your CPU, made by Apple, Intel, AMD etc., has a number of PCIe lanes and a number of USB/Thunderbolt ports for connecting peripherals. Those have an aggregated throughput, which sets an upper limit for the amount of data that can be read per second from all the peripheral devices. In a given computer, usually not all the lanes and ports of the CPU are actually connected, so the limit may be even lower. In desktop PCs and mini-PCs, usually only 4 + 4 = 8 PCIe lanes are available for SSDs, and when there are more SSD sockets they share some of those lanes. A much higher SSD throughput could be achieved in a desktop PC by using the GPU connector with an SSD adapter for M.2 SSDs, which has 16 PCIe 5.0 lanes, with a 64 GByte/s throughput. Taking out the GPU might actually be OK for doing AI inference, because a beefy CPU like a Ryzen 9950X should be able to keep up with a reading throughput of 88 GB/s from 6 SSDs (2 on the motherboard and 4 on the add-on PCIe card), while computing inference in the INT8 or BF16 formats, so the absence of the GPU would not reduce the inference speed when it is limited by the speed of reading the weights.
- animanoir 26d ago[dead]
- hakandmr 26d ago[dead]
- ChaseRensberger 26d agonot sure ive ever seen a #1 post on HN with only 5 stars
- pjdesno 26d agoI wonder if faster SSDs would help? In particular you can still get used Optane SSDs on eBay, although they’re fairly pricey. (not the bogus m.2 ones that are slower than a halfway decent consumer NVMe)
- Argonautlabs 26d agoProbably not. Each read here is a whole 17.5 MB expert file, so the time per read is set by the drive's throughput, not its access latency — 17.5 MB at 7 GB/s is ~2.5 ms, which is what we measure at queue depth 1 on the SN8100s. What actually moves the barrier is how fast the slowest of 16 concurrent whole-file reads completes: more drives on direct ports, stable tail behaviour under load, and scheduling. Our ladder shows even that with diminishing returns (one drive ≈52% of four, three ≈90%).
- xtracto 26d agoCould RAID0 help?
- adrian_b 26d agoWhen you have heterogeneous SSDs, e.g. you mix PCIe 5.0, PCIe 4.0 and Thunderbolt interfaces, you can obtain a greater throughput by managing in software the distribution of data, than by using RAID0. RAID0 works fine only when all the interfaces have the same speed. If one SSD is twice faster than the other, in order to achieve maximal throughput, you must take care to place the data in such a way so that you will need to read twice more data from the twice faster SSD. In general, you must distribute the data so that the amounts read from each SSD are proportional with the throughputs of the SSDs. One could write a modified RAID0 device driver, which would use unequal stripes, with widths proportional with the SSD throughputs, but I am not aware of any such already existing RAID0 driver.
- zamadatix 25d agoOne approach I've seen used is: mdadm --create /dev/md0 --level=0 --raid-devices=3 /dev/nvme0n1p1 /dev/nvme0n1p2 /dev/nvme1n1p1 Which is effectively a way to get any positive integer m:n ratioed bandwidth distribution over any number of any sized drives.
- voiceeh 26d agoThat's actually pretty neat.
- dymk 26d agoThese slop readmes are painful to read.
- alex7o 26d agoI think this is cool not for kimi but for sth like glm flash
- walrus01 26d agoNow imagine the token/s rate decline after context fill at 200,000+ context.
- Argonautlabs 26d agoFair, and we didn't measure it. Decode was flat from 128 to 512 generated tokens (0.926 → 0.923 tok/s drafter-off), but that's a 6-token prompt plus the output — total context under a thousand. The current configuration admits about 4.4k tokens of context at all, and at anything like 200k the killer wouldn't be decode, it would be prefill: today it reads each layer's experts once per 64-row pass, so 200k tokens of prompt would be measured in days, not minutes, until the scheduling fix.
- NooneAtAll3 26d agowhat's the main limitation on context size? 4.4k seems... I just realized I have no sense of scale whatsoever
- Argonautlabs 26d agoMemory, not the model. The KV cache on this engine grows about 2.8 MiB per token of context, and the machine's 128 GB is already holding the 50.7 GiB resident trunk plus reserves the speculative-verify path needs (we found the hard way that squeezing those makes the verifier reject wide batches and decode falls to single-token steps). With the current reserves the engine admits ~4.4k tokens; that's a configuration ceiling you can raise by giving the cache more of the 128 GB and accepting less headroom elsewhere. K3 itself supports far longer contexts — but see the prefill caveat above: on this setup long prompts cost minutes per 512 tokens until the scheduling fix lands.
- mandeepj 26d agoYou currently can't run a 2.8T locally; there's just no way. So, it's a good start.
- theideaofcoffee 26d agoBut my local is a 8xH100, you insensitive clod!
- BoingBoomTschak 26d agoEven if that's impressive, the README is low SNR slop as usual...
- Argonautlabs 26d agoThank you. Put a five-line TL;DR at the top of the README — what it is, the number, the honest limit, the two findings, credits — with the detail below for anyone who wants it.
- snorrah 26d agoWhat was the decision to leave the LLM to write the detail section rather than author it yourself? There's a growing resentment about asking people to read LLM-produced words, especially if it's a large amount to read. Not sure if you were aware of that or not (I think there's been links to surveys / polls just recently on H.N)
- Argonautlabs 25d ago[flagged]
- nxtfari 26d agoWe’re reaching quadratic slop. Slop projects that don’t understand what they’re shipping built on top of slop projects that also don’t understand what they’re shipping. Magnificent.
- iamshs 26d agoGood start.
- amelius 26d agoThe SSDs are necessary because Apple's architecture doesn't allow RAM upgrades. Reminds me of someone who said "640KB ought to be enough for anybody".
- kulahan 26d agoGates never said that, for what it's worth.
- buzzerbetrayed 26d ago> In an April 1985 InfoWorld editorial, James Fawcette wrote that Gates had said something like: “When we set the upper limit of PC-DOS at 640K, we thought nobody would ever need that much memory.” So yes, Bill Gates denies that story. So it just depends on who you believe. But to flat out say he never said it is too confident.
- kulahan 26d agoI was mostly curious about the source and figured that would bring it out. Sorry.
- wtallis 26d agoTo be fair, nobody has upgradable memory in any system that has enough memory bandwidth and compute power to run LLMs with decent performance. It might be interesting to compare against some decade-old x86 server or workstation stuffed full of LRDIMMs to reach 1.5–2TB of RAM, but the bandwidth would be only slightly faster than a desktop today with high-end DDR5: nowhere close to GPU bandwidth. So performance would still suck. Designing for extreme expandability comes with pretty steep tradeoffs.
- LTL_FTC 26d agoTake a look at AMD’s 12-channel memory servers. The newer Epycs are up to 16-channels now, 1.6TB/s. Pretty great for inference.
- kgeist 26d agoThere's a tendency to cite only decode speeds, but in practice, an LLM generates far fewer tokens than it has to read (unless you ask general knowledge questions). So the effective performance is much slower than the decode rate suggests, because a 512-token prompt already takes 6 minutes to load
- vlovich123 26d agoGenerally infill is significantly faster than inference due to batching. Is that not the case here for some reason?
- kgeist 26d agoThey mention it here: https://github.com/argonautlabsai/deltafin/blob/main/k3-public-bench/results/PREFILL.md https://github.com/argonautlabsai/deltafin/blob/main/k3-publ... >device bytes read during the prefill window, all four drives (arm csv) 8,977 GB at 24.1 GB/s aggregate I.e. low memory bandwidth.
- solarkraft 26d agoPrefill is generally faster than generation, but not by much on older Mac processors. I get around 70-60 tps in prefill on my M1 Max for Muse Glimmer (not sure about the generation speed, probably between 15 and 30). They allegedly improved this by “up to 7x” with M5 but I’m not sure about the exact numbers here.
- dotinvictim 26d ago[dead]
- robrenaud 26d agoShould LLMs be designed to be modular, so that instead of needing access to the whole model, for a given prompt, only a small subset of the model would be used? If knolwedge was sufficiently modularized, most of it could be ignored. Maybe a hyopthetical model of 5T of indexable weights could be used with only 50 GB of GPU ram, efficiently, because it stays resident in the GPU.
- gsora 26d agoIsn't that the definition of an MoE model?
- jmolinski 26d agoNo, not really, current MoE limit the computation, not memory requirements. Router experts are not "sticky" enough to achieve what robrenaud describes - they'd have to be chosen per prompt, or at least per chunk, not per token.
- greazy 26d agoWhat is "sticky" in this context?
- robrenaud 25d agoExperts vary per token in MoE, there is maximum flexibility. Good for driving down loss, bad for locality/gpu memory/bandwidth. If expert selection were more constrained, inference systems could take advantage of it. Keeping experts cached would mean not needing to load them from disk/ram every token.