3 ms·
curious why dont they bake the system prompt in the model itself ? Why do we pay for these tokens on every API call ? These are just free $ for them, unnecessa
by pulkitsh1234 2mo ago
curious why dont they bake the system prompt in the model itself ? Why do we pay for these tokens on every API call ?
These are just free $ for them, unnecessary bloating the context
- epolanski 2mo agoWhy would it be a good idea? That would make the model quite inflexible. A system prompt is about guiding the behavior for the rest of the conversation. If I'm writing an agent for financial analysis I don't want the crap that belongs to a chat-based one, or a code-oriented one.
- solarkraft 2mo agoBaking them into the model and having them apply this strongly is hard and resource intensive, as far as I am aware. Having them in context is super easy and cheap. It is trivial to change and is 100% cacheable.
- simonw 2mo agoThese system prompts don't affect the API, they are for the Claude consumer chat products. We aren't charged extra for them. They're also prefix cached, so the cost to Anthropic and performance hit is greatly reduced.
- Dfol 2mo agoSo the people using the Claude consumer chat products pay for them via usage... That's not any better. It's actually worse.
- simonw 2mo agoI don't understand. System prompts are part of the software that customers pay to access. Complaining about that is a bit like complaining that your Netflix subscription includes paying to execute the compiled code that Netflix wrote that serves you video streams from their servers. Actually there is a difference: If Anthropic deleted a large chunk of that system prompt I guess you might get like a 1% increase in how much Opus 5 you can use via their chat allowance for your paid subscription. Is that really something worth being frustrated by?
- TZubiri 2mo agoCached. they are the first part of the input and it contains no user dependent variables, so the model is in a known state that it can reuse across all users, it does not need to recompute all that inference
- cubefox 2mo agoUnless they are using a linear architecture, the compute cost still scales O(n²) for n tokens, and nemory cost scales O(n).
- TZubiri 2mo ago>the compute cost still scales O(n²) for n tokens, That is never the cost, it's a common misconception. Cost scales linearly per tokens. Unless you are sending one token at a time and avoiding using the same machine or cache. Just look at api charges, they are charged by token, not by token squared.
- cubefox 2mo agoWhich seems to contradict the usual consensus that purely linear architectures are not sufficiently capable and unsuited for frontier models.
- jannyfer 2mo agoThey have a {{currentDateTime}} in the prompt which is interesting for something that gets prefix cached. Hopefully they are handling that properly :)
- bob1029 2mo agoIt's probably a lower resolution timestamp (seconds/minutes omitted).
- JimDabell 2mo agoYou don’t want to do that for anything you want to be able to vary, but they do something similar with a “soul document” for things they always want to apply. https://news.ycombinator.com/item?id=46125184 https://news.ycombinator.com/item?id=46125184
- supriyo-biswas 2mo agoIn this token-mania frenzy that has taken hold of the industry, I guess solutions like "soul document" and "system prompts" will continue for a while, and once the industry matures a bit we'll go back to things like LoRA[1] and control vectors[2][3]. The other explanation may be that these AI labs may be expecting more government scrutiny, and "here's a document" would probably go better than "here's some vector representation of our values" when talking to politicians. [1] https://arxiv.org/abs/2106.09685 https://arxiv.org/abs/2106.09685 [2] https://vgel.me/posts/representation-engineering/ https://vgel.me/posts/representation-engineering/ [3] https://transformer-circuits.pub/2024/scaling-monosemanticity/index.html https://transformer-circuits.pub/2024/scaling-monosemanticit...
- monkpit 2mo agoIs there a reason a document could not be converted to vectors via embedding, and you’d have both? EDIT: I see, the control vectors operate more directly upon the model, in a way embedding vectors don’t quite have access to.
- energy123 2mo agoIf it's a fine tuning step at the end, why is the need for it to vary a problem? Can't you run the fine tuning, test for regression, and deploy the weights in a day? I think the more likely reason is it doesn't work as well as in context learning. Otherwise they would prefer to avoid polluting context and degrading performance.
- TZubiri 2mo agoFine tuning isn't the same and doesn't have the same effect as selecting input tokens. Does there exist a model X that behaves exactly as a model Y with context Z? Maybe, but it's not trivial to achieve and might possibly be convoluted and more expensive.
- Marha01 2mo ago> curious why dont they bake the system prompt in the model itself ? Probably because if they did, they would need to retrain the model everytime they want to change the system prompt.
- tgsovlerkhgsel 2mo agoFully baking them in would make it expensive to update them. Caching kind of "bakes them in" (as in, removes part of the cost) while keeping it flexible.
- amelius 2mo agoFlexibility.
- EMM_386 2mo agoWhen you call via the API and want it to roleplay as a pirate or fix broken YAML in the coding harness - it doesn't need to know about the sports scores lookup tool or the recipe creation tool. There are different use cases for the same underlying model. They also can't tell Opus it might be a Fable handoff when Fable didn't exist when Opus was created. They need to be able to change them.
- deleted 2mo ago[deleted]