4 ms·
The efficient frontier of LLM inference
- brrrrrm 1mo agothis is a nice and concise writeup. what's striking to me is that these techniques really have not changed in /years/. sure, precision has become slightly lower, spec decoding acceptance has gotten slightly better and the complexity of parallelism is trickier with mixture of experts. but no new concepts in a very long time! the absolute most impactful improvements for inference comes at architecture design time. I firmly believe everyone who cares about impacting model efficiency should look there
- philipkiely 1mo agoI think the biggest net new recent technique is P/D disaggregation. And that spec dec is very different now especially post DSpark/DFlash. But overall yes the fundamentals of LLM performance optimization have been remarkably stable over the last few years.
- brrrrrm 1mo agoperhaps its unfair to say this in hindsight, but it's a fairly straightforward application of little's law that's been around for some time https://arxiv.org/html/2401.09670v2 https://arxiv.org/html/2401.09670v2
- Ozzie-D 1mo ago[flagged]
- regularfry 1mo agoIt's also the hardest point at which to try to work, because when you change the model architecture you need to completely retrain from scratch.
- nedo_var 1mo ago[dead]
- datadrivenangel 1mo agoThe author does not deeply mention that quality/intelligence is a third dimension here in addition to throughput and latency, and the frontier is jagged so quality and intelligence require bespoke benchmarks to evaluate tradeoffs for speed and cost.
- philipkiely 1mo agoThese are both good points that I attempted to cover, quotes: > In practice, the efficient frontier is very jagged. Rather than a smooth, continuous line between outcomes, small changes can have big impacts. These cutoff points are often unintuitive and must be discovered empirically through sweeps. > However, quantization introduces a new set of tradeoffs between quality and serving efficiency. This is a particularly jagged frontier, where a large degree of improvement to serving efficiency is possible with little-to-no reduction in model quality, especially when using microscaling floating-point number formats like MXFP4 and NVFP4. Would appreciate ideas on how to explain in greater depth
- calclavia 1mo agogood recap on the recent inference techniques!
- paidx 1mo ago[flagged]
- ttoinou 1mo agoInference techniques either move a deployment along the latency–throughput frontier or push the entire frontier out, creating more efficiency to allocate. This is a tautology. You can say that with anything. Gastronomy techniques will make a previous recipe better, or create a new recipe better than others, or a mix of both.
- Ifkaluva 1mo agoThe point is to classify them into two kinds. The kind that shifts the frontier is more powerful, since improves capabilities without incurring tradeoffs.
- philipkiely 1mo agoI also wrote this as somewhat of a defense of the techniques that don't move the frontier -- there is a lot of value in being able to pick a point on the curve.
- jing09928 1mo ago[dead]
- killerdog10 1mo ago[dead]
- jumploops 1mo ago> Speculative decoding is the process of guessing which tokens a model might generate, then validating those guesses. As a computer engineer, it’s always interesting to see optimizations applied at different levels of the stack. Speculative execution became pretty popular in the 90s, eventually used in basically every x86 design. Then in the mid-2000s the Speculator[0] paper brought that concept to distributed systems, which we’re still seeing work on[1][2]. Everything old is new again (: [0]https://www.cs.princeton.edu/courses/archive/fall07/cos518/papers/spec-execution.pdf https://www.cs.princeton.edu/courses/archive/fall07/cos518/p... [1]https://www.usenix.org/system/files/osdi25-shen-weihai.pdf https://www.usenix.org/system/files/osdi25-shen-weihai.pdf [2] https://www.microsoft.com/en-us/research/publication/distributed-speculative-execution-for-resilient-cloud-applications/ https://www.microsoft.com/en-us/research/publication/distrib...
- freakynit 1mo agoCan we expect similar issues such as spectre and meltdown that intel experienced with speculative execution.. but, in the form of prompt injection/poisoning?
- rf15 1mo agoOk, I'll bite: no, considering these are very different domains and you don't get system access by getting the wrong speculative branch for your next text token, you just get a slightly different (but probably still related enough) text.
- StevenWaterman 1mo agoSpeculative decoding is lossless because the main model checks whether it agrees with what the drafter outputted
- ranger_danger 1mo ago> you don't get system access by getting the wrong speculative branch for your next text token I think you could if the client supports tool/MCP calls.
- arjie 1mo agoYou know what I'm curious about? Whether you have brand guidelines inside the company, a Claude skillset, or the blog post author makes the charts in line with the brand colours and so on.
- philipkiely 1mo agoI draw my diagrams on notecards and send them to our designer who brings them to life. The images start out looking like this: https://philipkiely.com/images/blogs/how-to-write-a-book/design.JPG https://philipkiely.com/images/blogs/how-to-write-a-book/des...
- arjie 1mo agoMakes perfect sense. The classic way!
- fsckboy 1mo ago"the efficient frontier" is an important landmark of (investment) portfolio theory. It proves/explains/illustrates how you can combine selections from a diffuse cloud of individual investments and still land on a frontier that is better than any of your individual choices. It's the entire basis of "diversify your portfolio". The efficient frontier of LLM inference is a line, not a frontier. this is a frontier: https://upload.wikimedia.org/wikipedia/commons/e/e1/Markowitz_frontier.jpg?utm_source=en.wikipedia.org&utm_campaign=imageinfo&utm_content=thumbnail_unscaled https://upload.wikimedia.org/wikipedia/commons/e/e1/Markowit... no matter how good is something a smart person writes down, a pleb will come along and try to hang on its coattails. If you want to steal an idea for this, steal indifference curves, they'd make more sense.
- iciac 1mo agoThe authors are referring to a Pareto Frontier https://en.wikipedia.org/wiki/Pareto_front https://en.wikipedia.org/wiki/Pareto_front
- regularfry 1mo ago> The efficient frontier of LLM inference is a line, not a frontier. The efficient frontier of portfolio theory is also a line. Not sure I get your point here.
- fsckboy 1mo agonot sure i get why you don't get it. why have the word frontier if it doesn't mean something different than other words?
- regularfry 1mo agoBecause... it's an accurate description of the concept it's being applied to?
- yeasin-arafat 1mo ago[flagged]
- soricus 1mo ago[dead]
- qingcharles 1mo ago> A model is a “frontier model” if it offers the highest degree of intelligence at a given cost or size. I would define a "frontier model" as offering the highest degree of intelligence at any cost, or without regard to cost. The frontier today is clearly Fable/Mythos, with the "efficient frontier" at Opus/Sol.
- orangeboats 1mo agoYour partial quote is quite misleading. The article obviously talks about "efficient frontier", not "intelligent frontier". >In the AI industry, we borrowed the term “efficient frontier” from economists. We use it to talk about managing tradeoffs, most often the tradeoff between cost and capabilities for models. A model is a “frontier model” if it offers the highest degree of intelligence at a given cost or size.
- censor25 1mo agoNice read. I was wondering what can one do to get into inference engineering as simple theoretical knowledge is not sufficient and switching profiles is tough for someone with years of experience.
- clem_rw 1mo agoMy daily struggle is trying to make a 7B model respond in under 500ms without breaking the bank. This hits home.
- bit_rot73 1mo agoMy RTX 3090 is still laughing at my attempts to run 70B models efficiently.
- pan_lid 1mo ago[dead]
- juggle73 1mo ago[flagged]
- kgeist 1mo agoI'm currently trying to write an inference engine that combines the benefits of llama.cpp (one binary deployment, good support for heterogenous non-datacenter compute, wide quantization support) with the benefits of vLLM/SGlang (things like proper paged attention for better VRAM utilization and high concurrency). Datacenter hardware is expensive and there's shortage of it but llama.cpp is slow/unoptimized for concurrent use, while vLLM/SGLang easily crash on non-common setups (things like, if you do pipeline parallelism for RTX5090+RTX4090, they will randomly crash with RAM caching enabled or select wrong kernels because they usually assume that every rank is the same device type; they also don't support Q5-Q6). For me what's most interesting is to optimize inference for lack of good datacenter hardware and how to optimize for it best. I've been running an AI server in the office, and so far I've find these techniques most important for concurrent use on cheap hardware: pipeline parallelism (to accomodate for PCie), RAM caching (to quickly restore contexts into VRAM), speculative decoding (including domain-specific ngrams, they already can speed up code generation considerably without the overhead of a draft model), good kernels highly optimized for a specific device, support for Q5-Q6 (almost as good as Q8), FP8 contexts (more context to fit), paged attention (for better VRAM utilization), prefix caching, continuous batching (this is the default everywhere). So far the main bottlenecks have been llama.cpp's poor VRAM utilization for contexts (you either have fixed-size slots, or use unified KV cache where each request attends to attention from all other requests and then unnecessary portions of attention are masked out), and lack of decode/prefill segregation: when a request starts prefilling a long context, all decoding threads slow down to like 5 tok/sec. On the other hand, vLLM/SGLang feel superbuggy if you don't run them on some officially approved node like 8xH200
- armcat 1mo agoI think we can soon include "recursive depth" strategy that Astra is employing, which (I suspect) is using recursive internal state changes in the transformer as opposed to full forward-pass + sampling which has traditionally been the case with thinking/CoT. Similar method was used here (but different context - encoding tools inside the transformer weights for fast execution): https://www.percepta.ai/blog/can-llms-be-computers https://www.percepta.ai/blog/can-llms-be-computers
- copperwire 1mo agoQuantization and speculative decoding unlocked major savings for our smaller models. Still chasing that ideal cost/performance ratio.
- amelius 1mo agoHow do you know for sure that you stay on an efficient frontier when you change a parameter? I think this presentation says more about what knobs you can turn and in what direction the outcome will move (it may be worse than a competitor) than it says about frontiers.
- artyomsv 1mo ago[dead]
- luciana1u 1mo ago[flagged]
- entrope 1mo agoThis article describes what is usually called a Pareto frontier: the best known achievable trade-offs between two (or more) optimization goals. "Efficient frontier" in common usage seems to be specifically a Pareto frontier for financial risk versus return of an investment portfolio. Even outside of finance, points on a Pareto frontier are called Pareto optimal or Pareto efficient. A Pareto frontier is sometimes shown with more than two dimensions, although usually people will pick just two for simplicity. Within LLMs, and even inference naturally, there are many other potential parameters that one might optimize: Unsloth typically shows a Pareto frontier for size of a quantized model versus KL divergence. Others trade total concurrent tok/s against single-stream tok/s. KV cache size, context length and context coherency are other trade-offs that are closely related to inference. Total intelligence is usually a defining characteristic of a "frontier model", with cost (per token or task) as a salient trade-off. Cost is one parameter that is implicitly fixed by the "throughput versus latency" analysis: using a GB300 versus Radeon R9700 moves the curve enormously and probably changes the shape of it. Lots of threads here argue over local vs cloud inference regarding cost efficiency, often with privacy and control as competing objectives.
- ggcr 1mo ago2026 has been the year where spec-dec has matured, it has been adopted by all big OSS engines and i'm sure it's present in quite a lot of inference providers as the default I feel like P/D dissaggregation will be the next big one for providers, as prefill tends to be compute bound while decode mem bound which I guess each will have a different type of node