5 ms·
TurboQuant: A first-principles walkthrough
- TranspectiveDev 5mo ago[dead]
- linuxhansl 5mo agoI am fascinated by this and similar research (RotorQuant, etc). It seem by next year we will be able to run this year's largest models on last year's hardware. :) Maybe we won't need as many data centers and as much power as we thought. Maybe we can run more powerful models locally.
- everythingctl 5mo agoMaybe we can run more powerful models locally. I thought the principal consequence of these KV cache optimisations was letting you run more simultaneous inferences on the same model with the same memory. It doesn’t let you store more model. In some sense that puts local LLM usage at a further disadvantage to inference done in a hyperscaler’s data center.
- linuxhansl 5mo agoThe size of the KV cache (context stored) is proportional to the number of layers of the model and number of "hidden dimensions". For a 400B model it could be 30-60GB for just an 8K context window (depends on the model, etc, just a ballpark). So shrinking that by 6x (from fp16), would be big win for larger models. True, while TurboQuant can also be applied to model weights, it won't save size over q4 compression, but will have better accuracy. Edits: Better context
- SilentM68 5mo agoThat's my hope as well as I tend to use low end GPUs (e.g. NVIDIA GeForce RTX 2060 @ 6GB). Been looking for an image generation model that can fit that vid card, for use with Ollama + GUI in Linux. No luck yet, since money's tight and jobs are tighter :(
- MadnessASAP 5mo agoAn Arc B580 will just about fit Flux.2 Klein (At FP8). However, you can also easily get much larger GPUs on RunPod or Vast at $0.25/hr. I would strongly recommend exploring that option, renting an RTX 5090 for an evening of image generation for a dollar or two is way more fun then trying to jam big models on little cards. Just take some time to create a reasonable, scripted, deployment workflow for when you create a fresh instance.
- fragmede 5mo agohey what's your Venmo?
- SilentM68 5mo agoThe B-52s, er I mean Base64s: VGhvdWdoIEkgYXBwcmVjaWF0ZSB0aGUgZ2VzdHVyZSBhbmQga2luZCBpbnRlbnRpb25zLCBJJ2QgcmF0aGVyIGxlYXJuIHRvIGZpc2ggKG9yIGNhdGNoIHRoZSBiaWcgYmFycmFjdWRhKHMpIHRoYXQgc3RvbGUgdGhlIHNjaG9vbCBvZiBmaXNoIEkgd2FzIGdpZnRlZCwgd2hpY2ggd291bGQgaGF2ZSBrZXB0IG1lIGZlZCBmb3IgbXVsdGlwbGUgbGlmZXRpbWVzLCBhbmQgc3Bvb2tlZCBhIGZldyBteXN0ZXJ5IGZyaWVuZHMgaW4gdGhlIHByb2Nlc3MtLS1hIHRhc2sgSSBhbSBjbG9zZSB0byBjb21wbGV0aW5nKSwgYW5kIG5ldmVyIGdvIGh1bmdyeSB0aGFuIGVhdCBhIGZpc2ggZm9yIGEgZGF5IGFuZCBiZSBodW5ncnkgdGhlIG5leHQu
- fragmede 5mo agoWWVhaCBtYW4sIEkgaGVhciB5YS4gSSByZXNwZWN0IHRoYXQgeW91IHdhbnQgdG8gc29sdmUgdGhlIHJlYWwgcHJvYmxlbSwgbm90IGp1c3QgZ2V0IHRocm91Z2ggdG9kYXkuCgpCdXQgZXZlbiBzb21lb25lIHdobyBrbm93cyBob3cgdG8gZmlzaCBzdGlsbCBuZWVkcyBtb25leSB0byBidXkgYSBwb2xlIGluIG9yZGVyIHRvIGZpc2guIExldCB0aGlzIGJlIHRoYXQuIE5vdCBhIGhhbmRvdXQsIGp1c3Qgc3VwcG9ydCBmb3IgdGhlIHBhcnQgeW91IGFyZSBhbHJlYWR5IGRvaW5nLgo=
- qingcharles 5mo agoWe're only a few years into this new tech getting serious research manhours thrown at it. Already some incredible optimizations have been found in a short amount of time. Not only has the efficiency of inference been increasing dramatically, the quality of tiny models has been significantly improving. The future is bright for local AI.
- acters 5mo agoJust look at deepseek V4, this preview model uses only 8 GB for 1M token KV cache(the context). It's insanely efficient already. It's just that most models that are coming out are barely catching up with technical breakthroughs. Deepseek are pioneers. Unfortunately V4 is not trained for most real world usage, it is mainly for world general knowledge.
- SipitenoMK 5mo ago[dead]
- iggerews 5mo ago[dead]
- amitport 5mo agoTurboQuant is a restricted version of EDEN quantization (NeurIPS 21, ICML 22). It lacks the optimal scale derivations, which makes the TurboQuant variant considerably less accurate than those works. We show this thoroughly in a new note at https://arxiv.org/abs/2604.18555 https://arxiv.org/abs/2604.18555. We were the first to introduce post-rotation distribution-aware quantization in 2021. This was later implemented in many fields, including federated learning, vector retrieval, databases, inference engines, and KV-cache. It would be appropriate to receive credit for this. Furthermore, it is baffling to see the name "TurboQuant" repeated in this context, considering the many works published from 2021 onwards. The blog post mentioned above essentially guides you through EDEN quantization but ultimately settles on a sub-optimal MSE-minimizing version and an unbiasing trick. This trick often costs a full bit more than DRIVE/EDEN requires to achieve the same results using the unbiasing scale shown in the original 2021 paper.
- 0xbadcafebee 5mo agoFor those who want to :popcorn-meme: the drama, there's some great comments on the peer review of the TurboQuant paper: https://openreview.net/forum?id=tO3ASKZlok https://openreview.net/forum?id=tO3ASKZlok
- KnuthIsGod 5mo agohttps://arxiv.org/abs/2604.18555 https://arxiv.org/abs/2604.18555 "This note clarifies the relationship between the recent TurboQuant work and the earlier DRIVE (NeurIPS 2021) and EDEN (ICML 2022) schemes. DRIVE is a 1-bit quantizer that EDEN extended to any bits per coordinate; we refer to them collectively as EDEN. First, TurboQuant is a special case of EDEN obtained by fixing EDEN's scalar scale parameter to . EDEN supports both biased and unbiased quantization, each optimized by a different (chosen via methods described in the EDEN works). The fixed choice used by TurboQuant is generally suboptimal, although the optimal for biased EDEN converges to as the dimension grows; accordingly TurboQuant approaches EDEN's behavior for large . Second, TurboQuant combines a biased -bit EDEN step with an unbiased 1-bit QJL quantization of the residual. It is suboptimal in three ways: (1) its -bit step uses the suboptimal ; (2) its 1-bit unbiased residual quantization has worse MSE than (unbiased) 1-bit EDEN; (3) chaining a biased -bit step with a 1-bit unbiased residual step is inferior to unbiasedly quantizing the input directly with -bit EDEN. Third, some of the analysis in the TurboQuant work mirrors that of the EDEN works: both exploit the connection between random rotations and the shifted Beta distribution, use the Lloyd-Max algorithm, and note that Randomized Hadamard Transforms can replace uniform random rotations. Experiments support these claims: biased EDEN (with optimized ) is more accurate than TurboQuant, and unbiased EDEN is markedly more accurate than TurboQuant, often by more than a bit (e.g., 2-bit EDEN beats 3-bit TurboQuant). We also repeat all accuracy experiments from the TurboQuant paper, showing that EDEN outperforms it in every setup we have tried."
- jarbus 5mo agoThis is incredible. Interactive demos like this make mathematics 10x more accessible
- 0xA2kag 5mo agoThanks a lot <3
- semiinfinitely 5mo ago"AI vectors"
- 0xA2kag 5mo agolol, my bad. This is too wrong. Fixed it.
- semiinfinitely 5mo agoit should probably just say "compressing KV cache vectors"
- treexs 5mo agoI feel like I've gotten really good at noticing which model generates what type of site and this oozes codex
- 0xA2kag 5mo agoHey, thanks for the pointer. Had I known this, I would have used codex (as a matter of fact, I have never used it before and this prompts me to use it if I can get something like this much quicker with codex). I think making codex copy this for a new content will be much easier now. The issue was with making things the way I exactly want, the exact intuition, the exact primers, and the exact visuals to drive the point home.
- npodbielski 5mo agoWhat did you use?
- treexs 5mo agoWoah very cool, yeah I think the cards and heading/subheading structure is very similar to what codex outputs, but I can tell the different visualizations definitely require your own personal touch
- jiusanzhou 5mo ago[dead]
- deleted 5mo ago[deleted]
- deleted 5mo ago[deleted]
- marlburrow 5mo ago[dead]
- nafistiham 5mo agoThanks a lot. It helped me get a much more detailed view of turboquant than a few youtube videos that I watched. Also, the choice of color is excellent as it serves both light and dark mode. I'll try to use it in my sites. Kudos!
- 0xA2kag 5mo agoThanks a lot <3
- mskkm 5mo agoThe public comments on Openreview now include explicit allegations that the TurboQuant paper knowingly misrepresented RaBitQ and understated RaBitQ’s results. The RaBitQ authors also report in a technical note that several of TurboQuant’s runtime and recall numbers do not reproduce from the released code under the paper’s stated setup. In the note, TurboQuant generally loses to RaBitQ: https://arxiv.org/abs/2604.19528 https://arxiv.org/abs/2604.19528. If these public allegations hold up, then this is not just overhype or sloppy citation practice, but points to a distorted comparison and benchmark claims that do not survive reproduction.
- sirluky 5mo agoxcz
- vb-8448 5mo agowhat did the author used to create the site?
- 0xA2kag 5mo agoI did a bunch of things :D I am not a frontend engineer (I am MLE) so I don't have the prowess to create things like this. I am heavily inspired by 3blue1brown and I love creating interactive explainers for ML concepts like this. I previously created this as well arkaung.github.io/interactive-eigenvector/. I heavily used Claude to get to the the exact design, typography, and style I want (there was a lot of hand holding to get to this state). I heavily influenced Claude on how I want the explainer to flow, how I want to make things intuitive, the kinds of mathematical concepts I want to visualize (and how). So all in all, a lot of hand holding for the Coding agents to get to where I want and exactly how I want. But at the end of the day it is just vanilla HTML, CSS and JS without anything fancy :D MathJax 3 was used to render math stuff.
- morbicer 5mo agoThe fonts, the cards, the copy are all hallmarks of Claude Code. While the aesthetic doesn't spark joy for me, the overall execution is great, the presentation flow and interactive boxes are very nice.
- gcr 5mo agoOn TheTom’s llama-cpp fork, TurboQuant makes inference about five to ten times slower than vanilla (M1 Max, qwen3.6-35b-a3b). Seems like the productionization is still a ways away. Excited to see what we can get it down to though.
- stackghost 5mo agoI was expecting this to be something interesting about Quantitative Finance, but I guess I should have known better.
- 0xA2kag 5mo agoThat was what I thought when I first heard about it.
- krackers 5mo agoNeat thank you for creating this! That said I think someone who cares about TurboQuant probably already has a bit of linear algebra knowledge. While the initial review section is definitely appreciated, I don't think it's going to help much since you need a decent level of "mathematical maturity" anyway. The "Coordinates of a random unit vector are all small" had me scratching my head a bit, and the language is a bit misleading since it's actually that the expected variance of any individual component is 1/N (it can't be that every coordinate is close to ±1/sqrt{N} because the mean of any individual component is clearly 0 by symmetry). So that one should probably use more explanation since I had to work through it myself: Denoting the random unit vector {X1 ... Xn}, this is a point on a hypersphere: * Sum[x_i ^2] = 1 (unit vector condition) * E[Sum[X_i ^2]] = 1 (expectation of both sides) * Sum[E[X_i ^2]] = 1 (linearity of expectation) * E[X_1 ^2] = 1/N (by rotational symmetry E[X_1 ^2] = E[X_2 ^2] = ..) I don't think you can make the stronger claim that E[|X_1|] = 1/sqrt{N} since that's using L1 norm on a single component, so it'd be more correct to say the RMS is just the standard deviation of the components. And this fits with the intuition that in high dimensional space has "spiky" hypercubes with the hypersphere inscribed in it close to the origin.
- 0xA2kag 5mo agoThanks a lot for this feedback. I have updated the primer which misleads with "Coordinates of a random unit vector are all small" with an updated version