6 ms·
This reveals a staggering level of incompetence, if that’s really all it is, and lack of transparency. They don’t have ANY product-level quality tests that pic
by dgreensp 5mo ago
This reveals a staggering level of incompetence, if that’s really all it is, and lack of transparency.
They don’t have ANY product-level quality tests that picked this up? Many users did their own tests and published them. It’s not hard. And these users’ complaints were initially dismissed.
I don’t think the high vs medium change is really on par with the others. That’s a setting you change in the UI, and depending on what you are doing, both effort levels are pretty capable, they just operate a bit differently. Unless I’m missing something and they are saying they were doing some kind of routing behind the scenes.
If they are constantly pushing major changes to the prompts and workings of the tool, without communicating about it, and without testing, it’s likely there are other bugs and quality-degrading changes beyond the ones in this article, which would make a lot of sense.
- kolinko 5mo agoThere were few systems like claude in the past, to testing rulebook is not really written yet. And far from obvious.
- sockgrant 5mo agoLLM evals are well established, are these not applicable here?
- rekrsiv 5mo agoTime is finite and regression testing always gets punted to the back of the line when humans are excited. This simply reveals a staggering level of humanity.
- FartyMcFarter 5mo agoSoftware engineering is not a new field. Best practices on testing are mature now, and Anthropic has poached enough engineers from companies with a solid understanding of those practices. Yet, their flagship product got three really bad changes shipped into it and only resolved after more than a month. This raises another question: with all the industry-wide boasting about AI-driven productivity, why does the leading company in agentic coding take over a month to fix severe customer-reported issues?
- sfink 5mo ago> Why does it take the company that is probably the best at agentic coding more than a month to find and solve such large regressions, even with customers complaining about them? My unfounded suspicion: because this is the tradeoff we're all facing and for the most part refusing to accept when transitioning over to LLM-driven coding. This is exactly how we're being trained to work by the strengths and limitations of this new technology. We used to depend on maintaining a global if incomplete understanding of a whole system. That enabled us to know at a glance whether specs and tests and actual behavior made sense and guided our thinking, enabling us to know what to look at. With agentic coding, the brutal truth is that this is now a much less "efficient" approach and we'll ship more features per day by letting that go and relying on external signs of behavior like test suites and an agent's analysis with respect to a spec. It enables accomplishing lots of things we wouldn't have done before, often simply because it would be too much friction to integrate it properly -- write tests, check performance, adjust the conceptual understanding to minimize added complexity, whatever. So in order to be effective with these new tools, we're naturally trained to let go of many of the things we formerly depended on to keep quality up. Mistakes that would have formerly been evidence of stupidity or laziness are now the price to pay for accelerate productivity, and they're traded off against the "mistakes" that we formerly made that were less visible, often because they were in the form of opportunity cost. Simple example: say you're writing a simple CLI in Python. Formerly, you might take in a fixed sequence of positional arguments, or even if you did use argparse, you might not bother writing help strings for each one. Now because it's no harder, the command-line processing will be complete and flexible and the full `--help` message will cover everything. Instead, you might have a `--cache-dir=DIR` option that doesn't actually do anything because you didn't write a test for it and there's no visible behavioral change other than worse performance. Closely related, what do you do with user feedback and complaints? Formerly they might be one of your main signals. Now you've found that you need dependable, deterministic results in your test suite that the agent is executing or it doesn't help. User input is very very noisy. We're being trained away from that. There'll probably be a startup tomorrow that digests user input and boils out the noise to provide a robust enough signal to guide some monitoring agent, and it'll help some cases, and train us to be even worse at others.
- 5mo ago
- close04 5mo ago> This simply reveals a staggering level of humanity. Wasn't AI supposed to solve all the drudgery? All those humans aided by cutting edge AI are still failing at these basic tasks? Then how good is that AI in the first place?
- rekrsiv 5mo agoNo, AI wasn't supposed to solve all that drudgery. The hypothesized AI singularity would, but an ordinary AI agent running an LLM is just a problem solving automaton with no will of its own, just like a fleshy brain solving computer problems is just a code monkey.
- afavour 5mo ago> This simply reveals a staggering level of humanity. Pretty embarrassing for an AI company. Surely AI should be doing their regression testing?
- culopatin 5mo agoTheir in house philosopher thinks Claude gets anxiety though
- dantillberg 5mo agoI would think that many of these defects should show up clearly in service-side analytics as well. For example, the bug that repeatedly re-cleared thinking for old sessions would cause a substantial drop in token cache hit rate for sessions > 1hr for the affected claude code versions. Session age & claude code version seeeem like obvious dimensions for analytics. But perhaps only in hindsight.
- lanyard-textile 5mo agoEh :) Let's not forget the humans on the other end of this. One of them was a bug that didn't present itself until after an hour of usage.
- tuwtuwtuwtuw 5mo agoSeems like that would be trivial to test?
- stickfigure 5mo agoMost bugs are trivial to test for after you know about them.
- mh- 5mo agoTrue, but when your cache configuration has exactly 2 TTLs and modalities, I don't think it's offbase to expect them to test what happens in the cache hit/miss scenarios for each of those. (I write this as someone who likes Claude Code, if that matters.)
- TedDallas 5mo agoAbout 20 years ago I maintained a shop floor control client/server application. I asked my manager why we didn't have any independent Q/A. He said we didn't need any testers because we have 500 in the building. Wild west days then. Looks like we are back.
- PeterStuer 5mo agoBack implies we ever left.
- gozzoo 5mo agoIt is worse than that. People have been complaining for weeks and Anthropic’s message was basically “you are holding it wrong”. On top of that this misconfiguration somehow makes CC consume much more tokens. How believable is all that?
- musebox35 5mo agoThey say that they did test but the coverage was not enough to pick it up, at least for the prompt change: “ After multiple weeks of internal testing and no regressions in the set of evaluations we ran, we felt confident about the change and shipped it alongside Opus 4.7 on April 16. As part of this investigation, we ran more ablations (removing lines from the system prompt to understand the impact of each line) using a broader set of evaluations. One of these evaluations showed a 3% drop for both Opus 4.6 and 4.7. We immediately reverted the prompt as part of the April 20 release.” Considering the number and scope of users they serve, I can sympathize with the difficulty. However, they should reimburse affected users at least partially instead of just announcing “our bad, sorry “. That would reduce the frustration.
- baxtr 5mo agoNaively, one could assume that with AI it should be possible to create a long and broad list of test cases…
- overfeed 5mo ago> If they are constantly pushing major changes to the prompts and workings of the tool, without communicating about it These are all classic symptoms of vibe-induced AI velocitis, sold by AI-peddlers as the future of the industry under the guise of "productivity." AI can help one generate a lot of code, but the poor engineers approving the deluge of changes are still using their old, unmodified, stock meat-brains. An individual change may look fine in isolation, but when it's interacting with hundreds or thousands of other changes landing the same week , things can go south quickly. Expect more instability until users rebel, and/or CTOs amd CIOs cry uncle. Amazon reportedly internally sounded the alarm after a couple of AI-tool-induced SEVs. The challenges at Github and the company insisting you don't call it Microslop are also rumored to be AI-related.
- legulere 5mo agoTo me it reads more like they are struggling to scale with requests and are trying to find ways that hurt users the least.
- baxtr 5mo agoYou’re talking about their intentions. OP is talking about how they don’t test continuously / densely enough for quality. I think both can be true.
- greatgib 5mo agoTo give my best guess, I think that the change of default effort is unrelated to the major problems encountered by the users but that this was added big and first to cover up a little bit the huge failure of the 2 other ones. First thing you will read and that takes a big part is that it was something like: not really a bug but we changed a default not well communicated and users (their fault) did not notice it. This is why they were "under the false impression" of a change. Lots of people will stop reading after a few paragraphs.