3 ms·
Autoregressive models can't just resume so they have to re-parse the entire prompt again for each execution. By caching them they resume from where it left off
by burtonator 2y ago
Autoregressive models can't just resume so they have to re-parse the entire prompt again for each execution.
By caching them they resume from where it left off from before thereby completely bypassing all that computation.
For large contexts this could save a ton of compute!
I think this feature and structured outputs are some of the biggest inventions in LLMs this year.
- minimaxir 2y agoPrompt caching has been a thing for LLMs since GPT-2 (e.g. transformers's `use_past=True`), it's more of a surprise that it took this long for the main LLM providers to provide a good implementation.
- brylie 2y agoI’m building an app with OpenAI, using structured outputs. Does OpenAI also support prompt caching?
- minimaxir 2y agoNot currently.
- cma 2y agoI'm sure internally they use it for the system prompt at least, probably since launch. And maybe for common initial user queries that exactly match.
- Onavo 2y agoThey are certainly not passing the savings on to the users.
- minimaxir 2y agoYet. I suspect OpenAI will release a similar offering soon. (hooray, free market competition!)
- HeatrayEnjoyer 2y agoThat $100 billion data center has to get paid for somehow.