7 ms·
I wonder if it's related that that OpenAI has found a way to cut inference costs by half, according to The Information. https://www.theinformation.com/newslett
by postalcoder 3mo ago
I wonder if it's related that that OpenAI has found a way to cut inference costs by half, according to The Information.
https://www.theinformation.com/newsletters/ai-agenda/openai-discovers-new-way-cut-inference-costs-half https://www.theinformation.com/newsletters/ai-agenda/openai-...
- drivebyhooting 3mo agoWhat’s the technique? And did they buy it from thinking machines?
- turtleyacht 3mo agoMaybe cache similar answers from others. Surprised if this is not already being done.
- nimchimpsky 3mo ago[dead]
- wahnfrieden 3mo agoLike google search, this does not work because of how common long tail use is. What you think could be a big chunk, is more likely to be a fraction of a percent of queries. And what use is similar query caching - so you (very often! if actually cost effective, maybe half the time) get a response to a query that was different from yours. Including for when you have a lot of context input already. You’re going to get trash. If it were constrained to only very common initial prompts, and somehow the long tail did not actually dominate as it does with Google search (can't find the reference at the moment but it was a famous article some years ago), it also wouldn't account for serious enough cost savings. Long context is what is expensive. This might only work in constrained domains like customer service where there’s tolerance for generic answers and escalation paths. For technical work? For general purpose use, with secretly canned responses charged at full price?
- turtleyacht 3mo agoPlease pardon the pure speculation incoming. Yes, caching the answer doesn't seem useful. Caching the progression, the graph, may be. This is similar to making code changes with ed(1) instead of editing in vi. The transform script(s) are cached and can be played back or adjusted. Surely for some breadth of question inputs, they map more often to similar answers--but not static answers; instead, evented edits. It's nearly untenable for a human to keep private edit scripts to generate code changes. The extra steps for custom regex, essentially one-offs for a shared codebase, is inefficient. But maybe not to an LLM.
- wahnfrieden 3mo agoI don't understand how this fits LLM architecture at all
- dools 3mo agoI would be very surprised if they hadn’t sorted out some form of shared KV caching
- wahnfrieden 3mo agoI wouldn't
- ewild 3mo agopeople really dont understand how the transformer works to think this is something trivial if possible at all
- dools 3mo agoOrly? “ Automatic Prefix Caching (APC in short) caches the KV cache of existing queries, so that a new query can directly reuse the KV cache if it shares the same prefix with one of the existing queries, allowing the new query to skip the computation of the shared part.” https://docs.vllm.ai/en/latest/features/automatic_prefix_caching/ https://docs.vllm.ai/en/latest/features/automatic_prefix_cac...
- simpleintheory 3mo agoThey're already doing this under the name "fast answers" [1]: https://www.reddit.com/r/OpenAI/comments/1stsxvc/new_feature_fast_answers/ https://www.reddit.com/r/OpenAI/comments/1stsxvc/new_feature...
- layla5alive 3mo agohttps://archive.ph/NEwVz https://archive.ph/NEwVz "However, these inference optimizations, which rival Anthropic refers to as “compute multipliers,” are a big focus for all the labs. Anthropic CEO Dario Amodei has been publicly talking about the concept since at least mid-2023, when he said on a podcast that the company limits “the number of people who are aware of a given compute multiplier” because it could give other AI labs a leg up if they were to be able to replicate them. (Compute multipliers can also refer to efficiency optimizations in the model-training phase.)" Yes, on a world with finite resources where your industry is singlehandedly siphoning ALL THE RESOURCES - hoard general efficiency optimizations and treat them as trade secrets - winning is all that matters, normal people and other species and the planet be damned. Everything I hear about Dario these days makes me like him less and less. He sure did seem to speed run the 'tech leader with scruples' to 'tech villain' path! I guess all the cycles are compressing as we approach the singularity..
- razodactyl 3mo agoNot sure I know where I fall regarding your point: Yes to trade secrets, but also science and AI should be for the good of all. OpenAI seems to be trading roles back with Anthropic becoming misanthropic. I hope they both start heading in the direction of how the AI field was prior to LLMs. Collaboration and benefit for all should always be the primary motivator.
- georgemcbay 3mo ago> Collaboration and benefit for all should always be the primary motivator. Of all the things to never happen, this is never going to happen the most. That train left the station for good once hundreds of billions to trillions of dollars were involved. On the bright side, in the long run I suspect the vast majority of the value of AI will not be captured by the model making labs and the vast investments in them are going to implode, so...
- fragmede 3mo agoimplode, how?
- minimaxir 3mo agoSemi-related, has anyone noticed their GPT 5.5 usage in Codex being cut in half as of a couple days ago? I got a lot more mileage out of my session usage yesterday for the same workload.
- oxmom 3mo agoI've noticed less quota and 5.5 intelligence degrading. I didn't run the analysis like the post the other day, but I had noticed decreasing ability to complete tasks, more laziness. Switched back to 5.4 and it's much better. Maybe they're getting ready to launch 5.6? https://github.com/openai/codex/issues/30364 https://github.com/openai/codex/issues/30364
- neon_diogenes 3mo agoYes, I’ve noticed the same thing recently. Yesterday I burned through four 5-hour usage-limit resets in roughly four hours of wall-clock time. Same workflow, but it feels like I’m getting about half the mileage now, maybe even less. Guess that means Sol is dropping soon..
- torginus 3mo agoI wonder if AI labs are actively manipulating the narrative (and thus investor sentiment) by airing problems, and then solving them weeks to months later. I wouldn't be surprised if they have a lot of stuff figured out that is not included in the current version, just to make a steady product cycle with years of tangible improvements from one version to another (this is a common practice in the industry). For example, if inference isn't too expensive, but they figure out how to cut costs, then price goes down. After all, why pay OpenAI when a smaller datacenter can give you similar models? But, if they make a huge issue about how inference is too expensive, they engineer a crisis of their own creation - then, once they deploy the solution (which they might already have), then they're back on top.