3 ms·
Do you have any links to statements they have made about that? I can only find this from last year. >We've read and heard that you'd appreciate more transpar
by irthomasthomas 1y ago
Do you have any links to statements they have made about that? I can only find this from last year.
>We've read and heard that you'd appreciate more transparency as to when changes, if any, are made. We've also heard feedback that some users are finding Claude's responses are less helpful than usual. Our initial investigation does not show any widespread issues. *We'd also like to confirm that we've made no changes to the 3.5 Sonnet model or inference pipeline.*
https://www.reddit.com/r/ClaudeAI/comments/1f1shun/new_section_on_our_docs_for_system_prompt_changes/ https://www.reddit.com/r/ClaudeAI/comments/1f1shun/new_secti...
That statement aged poorly. The recent incident report admits they "often" ship optimizations that affect "efficiency and throughput." Whether those tweaks touch weights, tensor-parallel layout, sampling or just CUDA kernels is academic to paying users: when downstream quality drops, we eat the support tickets and brand damage.
We don't need philosophical nuance about what counts as a "model change." We need a change log: timestamped, versioned, and machine-readable that covers any modification that can shift outputs: weights, inference config, system prompt, temperature, top-P, KV-cache size, rollout percentage, the lot. If your internal evals caught nothing but users did, give us the diff and let us run our own regression tests.
History proves inference changes can drastically alter outputs. When gpt-oss launched, providers using identical weights delivered wildly different qualities due to inference configurations.
We need transparency about all changes whether model weights or infrastructure. Anthropic's eval suite clearly missed this real-world regression. Proactive change notifications would let us run our own evals to prevent failures. Without this, we're forced to reactively troubleshoot. An unacceptable risk for production systems.