4 ms·
"we often make changes intended to improve the efficiency and throughput of our models.." https://status.anthropic.com/incidents/h26lykctfnsz https://status.a
by irthomasthomas 1y ago
"we often make changes intended to improve the efficiency and throughput of our models.."
https://status.anthropic.com/incidents/h26lykctfnsz https://status.anthropic.com/incidents/h26lykctfnsz
I thought Anthropic said they never mess with their models like this? Now they do it often?
- mccoyb 1y agoI read this as changes to quantization and batching techniques. The latter shouldn’t affect logits, the former definitely will …
- jjani 1y agoThey already have a track record of messing with internal system prompts (including those that affect the API) which obviously directly change outputs given the same prompts. So in effect, they've already been messing with the models for a long time. It's well known among founders who run services based on their products that this happened, everyone who does long output saw the same. It happened around November last year. If you had a set of evals running that expected an output of e.g. 6k tokens in length on 3.5 Sonnet, overnight it suddenly started cutting off at <2k, ending the message with something like "(Would you like me to continue?)". This is on raw API calls. Never seen or heard of (from people running services at scale, not just rumours) this kind of API behaviour change for a the same model from OpenAI and Google. Gemini 2.5 Pro did materially change at time of prod release despite them claiming they had simply "promoted the final preview endpoint to GA", but in that case you can give them the benefit of it being technically a new endpoint. Still lying, but less severe.
- simonw 1y agoCan you expand on "messing with internal system prompts" - this is the first I have heard of that.
- jjani 1y agoWe talked about this a few weeks ago so it can't be the first time you're hearing about this :) [1] You hadn't heard about it before because it really only affected API customers running services including calls that required output of 2.5k+ (rough estimate) tokens in a single message. Which is pretty much just the small subset of AI founders/developers that are in the long-form content space. And then out of those, the ones using Sonnet 3.5 at the time for these tasks, which is an even smaller number. Back then it wasn't as favoured yet as it is now, especially for content. It's also expensive for high-output tasks, so only relevant to high-margin services, mostly B2B. Most of us in that small group aren't posting online about it - we urgently work to work around it, as we did back then. Still, as I showed you, some others did post about it as well. The only explanations were either internal system prompt changes, or updating the actual model. Since the only sharply different evals were those expecting 2.5k+ token outputs with all short ones remaining the same, and the consistency of the change was effectively 100%, it's unlikely to have been a stealth model update, though not impossible. [1] https://news.ycombinator.com/item?id=44844311 https://news.ycombinator.com/item?id=44844311
- simonw 1y agoI didn't see that as being about changing system prompts, I thought it was about changing model weights.
- simonw 1y agoAnthropic have frequently claimed that they do not change the model weights without bumping the version number. I think that is compatible with making "changes intended to improve the efficiency and throughput of our models" - i.e. optimizing their inference stack, but only if they do so in a way that doesn't affect model output quality. Clearly they've not managed to do that recently, but they are at least treating these problems as bugs and rolling out fixes for them.
- irthomasthomas 1y agoDo you have any links to statements they have made about that? I can only find this from last year. >We've read and heard that you'd appreciate more transparency as to when changes, if any, are made. We've also heard feedback that some users are finding Claude's responses are less helpful than usual. Our initial investigation does not show any widespread issues. *We'd also like to confirm that we've made no changes to the 3.5 Sonnet model or inference pipeline.* https://www.reddit.com/r/ClaudeAI/comments/1f1shun/new_section_on_our_docs_for_system_prompt_changes/ https://www.reddit.com/r/ClaudeAI/comments/1f1shun/new_secti... That statement aged poorly. The recent incident report admits they "often" ship optimizations that affect "efficiency and throughput." Whether those tweaks touch weights, tensor-parallel layout, sampling or just CUDA kernels is academic to paying users: when downstream quality drops, we eat the support tickets and brand damage. We don't need philosophical nuance about what counts as a "model change." We need a change log: timestamped, versioned, and machine-readable that covers any modification that can shift outputs: weights, inference config, system prompt, temperature, top-P, KV-cache size, rollout percentage, the lot. If your internal evals caught nothing but users did, give us the diff and let us run our own regression tests. History proves inference changes can drastically alter outputs. When gpt-oss launched, providers using identical weights delivered wildly different qualities due to inference configurations. We need transparency about all changes whether model weights or infrastructure. Anthropic's eval suite clearly missed this real-world regression. Proactive change notifications would let us run our own evals to prevent failures. Without this, we're forced to reactively troubleshoot. An unacceptable risk for production systems.