5 ms·
Seems like their competition (Google, Anthropic) have both shipped prompt caching, which is a lot more developer-friendly and gets a lot of the same benefits as
by gamegoblin 2y ago
Seems like their competition (Google, Anthropic) have both shipped prompt caching, which is a lot more developer-friendly and gets a lot of the same benefits as fine-tuning.
You just cache a prompt with a ton of examples that you would have otherwise fine-tuned on. You can trivially update that prompt whenever, no asynchronous fine-tuning job needed.
Wonder if OpenAI will stick with fine-tuning or go towards prompt caching (or both). Fine-tuning has uses, but prompt caching gets you 99% of the benefits for 1% the effort.
- eightysixfour 2y agoThey, in my experience, result in different outcomes. Fine-tuning is better at shaping the response, for example getting a correct JSON format, or making the responses shorters and concise, or use emojis, or whatever. Long prompts are better for in-context learning, fine-tuning doesn't seem to impart knowledge well, it just increases or decreases the likelihoods of existing knowledge coming out.
- maeil 2y agoGenuinely curious as you have experience with fine-tuning, which I don't yet. Given how simple it has become to get the correct response shape, do you feel there's still a point to it? If you need JSON and want to be very sure you can just use function calling, if you only need simple boundaries, use Claude-like XML, if it's something very complicated you can give a few shots. From my understanding, fine-tuning allows for reducing model size, meaning less latency and cost. That seems like it would be the biggest advantage, no?
- eightysixfour 2y ago> Given how simple it has become to get the correct response shape, do you feel there's still a point to it? I am personally working on systems which don't do mrequire uch structured data output; my work is more around structure, style, etc. Fine-tuning is more effective and more consistent for me in those contexts. One thing to remember is that the big models, like 4o and Claude, both are fine-tuned for chat interactions. This tuning can often make them dumber, more verbose, etc. because that's what the human feedback testing likes. You can look at some of the rating tests on chatbot arena as an example where two give the same answer, but people have voted more favorably for the one that delivers it with "more personality." There are SLMs that are better tuned for specific use cases and outperform the big models as a result of this, and the examples OpenAI showed in the article make it clear the big models can benefit from this as well. > From my understanding, fine-tuning allows for reducing model size, meaning less latency and cost. That seems like it would be the biggest advantage, no? This is partially correct, fine-tuning can be a step of this process. The full pipeline looks more like: 1. Create a prompt which gets a large model to output (mostly) correct responses. 2. Build a dataset of those inputs/outputs, probably with human or LLM curation/judging in the loop since some will still be wrong. 3. Fine tune a small model on those inputs/outputs. Now you have a smaller model which behaves more like the large model you were able to prompt engineer into instruction following.
- gamegoblin 2y agoI agree that fine-tuning isn't good at imparting knowledge, though in my experience, K-shot prompting is very close to as good as fine-tuning at getting output formats (for recent models, at least) So except for cases where you are really trying to perfectly match a particular writing style and want to fine-tune on loads of some author's text, I think prompt caching mostly dominates fine-tuning for most real-world workloads. I'd still fine-tune if I wanted my model to have a really particular "voice" though.
- WanderPanda 2y agoI would have expected fine-tuning to be good at imparting knowledge. Pre-training is often done for a single epoch only and models soak up the knowledge like crazy without multiple passes so why would fine-tuning be any different?
- eightysixfour 2y agoBecause the learning rate vs. pre-training is completely different. This is not accurate, but my mental model is that the LLM's initial training establishes the "space of concepts and ideas" while tuning (like RLHF and fine-tuning) changes how it expresses those concepts and ideas. It works well for me in deciding my approach.
- leobg 2y agoThe problem is that it’s almost impossible to teach knowledge to an LLM without teaching a specific form of expression at the same time. When you have a question/answer tuple in your training data, you are also teaching the model that every other way of answering the question is wrong. So while the LLM would probably be capable of generating maybe 100 answers to the question that would be equally useful (just using different phrasing, different choice of words, etc.) come on, you are forcing it to update its parameters to suppress all of these, except the one specific form that you selected. So you’re not really adding knowledge. Instead, you’re chiseling knowledge away.
- eightysixfour 2y agoI'm not sure this is entirely true, but I guess we could test it by generating a large enough dataset around a specific concept, and see if we can add that concept to a model. Or change one that exists to something else entirely. For example, create an "idea", generate thousands of Q&A pairs about that idea from different angles, and generate conversations about that idea, then train the model on it. This is essentially the Phi process, but with a single concept instead of "everything." My guess is that fine-tuning cannot add the concept to the model without suffering from catastrophic loss everywhere else. However if we added that same data to the pre-training dataset and retrain the model, the model would express the idea correctly.
- esafak 2y agoDoes their prompt caching work with semantically equivalent queries, or do they match lexically?
- simonw 2y agoYou have to mark exactly which parts of the prompt you want to cache. The only vendor I've seen with "automatic" prompt caching so far is DeepSeek https://platform.deepseek.com/api-docs/news/news0802/ https://platform.deepseek.com/api-docs/news/news0802/
- deleted 2y ago[deleted]
- KTibow 2y agoOpenAI has a mysterious reference to caching in their guide to latency. > Maximize shared prompt prefix, by putting dynamic portions (e.g. RAG results, history, etc) later in the prompt. This makes your request more KV cache-friendly (which most LLM providers use) and means fewer input tokens are processed on each request. source: https://platform.openai.com/docs/guides/latency-optimization https://platform.openai.com/docs/guides/latency-optimization
- okl 2y agoSounds very similar to how e.g. docker/layers work - the later you put the more dynamic stuff, the higher the change that previous layers are reused. Meaning the cache entries are likely chained, i.e., each one depends on the preceding entry.
- htrp 2y agoHow do you know it's not prompt caching in the backend?
- simonw 2y agoOpenAI haven't documented if they do this, but it should be reasonably easy to determine via experiments against their API. Claude's prompt caching gives a very material performance improvement - they claim up to a 4x performance boost in https://www.anthropic.com/news/prompt-caching https://www.anthropic.com/news/prompt-caching - so if OpenAI have similar techniques, even undocumented, they should become visible through sending the same prompt a bunch of times and measuring the latency.
- NavinF 2y agoOpenAI docs say they also improve latency with caching: https://news.ycombinator.com/item?id=41302358 https://news.ycombinator.com/item?id=41302358
- nprateem 2y agoIMO they are definitely doing something. I've had countless occasions where the first few times I run a prompt I get excellent output, then several tries later it just returns the same dumb-sounding AI shite.
- impossiblefork 2y agoIt's also pretty much required if you want to do anything complicated with prompt networks or agent-type stuff. Without caching it's too expensive.