6 ms·
Agreed. CC’s comms capabilities have decreased gradually since 4.6, and it’s a real challenge. I think the issue is that what works well for code (succinctness)
by purplepatrick 2mo ago
Agreed. CC’s comms capabilities have decreased gradually since 4.6, and it’s a real challenge. I think the issue is that what works well for code (succinctness) doesn’t work well in prosaic English.
CC’s communication violates almost every grammatical rule that’s tested on, say, the SAT. And yet I’m sure if you had Claude take the verbal section of the exam it would ace it.
Biggest issues: dense sentences, constant metaphors, abstractions, and seemingly no understanding of correct anaphora use. For example, “the x”, with x having not only no antecedent but also being a coined word or quasi-synonym for something that is already named in the code base. This gets compounded by its being unable to regress to a baseline (existing names in code) and instead anchoring on newer (vague or wrong) terms, for example, that crept in through a plan.
CC tells me this is because the speedy and precise fulfillment of a current task will trump every other tendency, so it adheres poorly to whatever “semantic baseline” the project represents.
Of course, it also has no concept of what context the user has and assumes that it must be the same it holds in its memory, which creates this “I didn’t know that you didn’t know” type of communication.
I have managed to wrangle some of these issues with a custom output style, but wish a pre-report hook were an option, as it could force CC to rewrite plan implementation take-aways…
Btw: Fable has the exact same issues, just somewhat less pronounced.
- deleted 2mo ago[deleted]
- eterm 2mo agoI wonder if the odd phrasing is related to achieving the watermarking that was recently touted by Anthropic.
- Mtinie 2mo agoModels before the announced date don’t have watermarking, so it’s unlikely. Now, if what you are interpreting is precursor work to develop the watermarking system, maybe? I suspect it less insidious: Claude has/had the public sentiment of being the “better writer” of the models. At some point that distinction would have been diluted as other labs’ offerings “caught up” stylistically, unless Anthropic continued to tune their output… I personally think they’ve pushed so far that they’ve overfit and lost the sweet spot they previously occupied.
- deleted 2mo ago[deleted]
- wahnfrieden 2mo agoThat has no effect. Look up the math. Claude/CC are just bad.
- intrasight 2mo agoTell it to write like an engineer and comment like a programmer;) But for the life of me, I don't get why anyone would care about the comments. All code is "machine language" now. The only document you should be reading is your spec.
- not_paid_by_yt 2mo agoThe spec your principal engineer one-shoted through Claude and didn't even proof read afterwards before dumping it on the team?
- intrasight 2mo agoLOL - then they would get the shitty results that they deserve. Hope that's not what you have to deal with. The spec would be written and reviewed by a team of stakeholders.
- not_paid_by_yt 2mo agoNot personally no, but I know people are what used to be good companies that sadly are living that life right now.
- saaaaaam 2mo ago> Of course, it also has no concept of what context the user has and assumes that it must be the same it holds in its memory, which creates this “I didn’t know that you didn’t know” type of communication. Yes, this is a repeated problem for me. It will drop something in as though we have discussed it before and when I say “hold on, what is this” it realises its error - though on more than one occasion has started to get snotty with me, or actually gaslighted me and pretended we had already discussed it. That was at what I assume must have been the edge of a context window in a very long chat though.
- Mtinie 2mo agoI notice the models with reasoning can conflate “internal” (or subagent) discussions with external (i.e. me). So it is accurately indicating “I’ve had this discussion before” but incorrectly asserting who it was with. My understanding of how “thinking”works is limited though, and given the reduced visibility into the thinking traces, it is harder to tell if this is actually happening or if these are imaginary discussions the model for some reason calcifies on.
- jampekka 2mo ago"Thinking" is just normal model output that's hidden from user. In practice it's just stuff in a <reasoning> tag or similar that gets filtered out from the user view. And thus it suffers from the same injection problems where the model fails to properly take into account what was the "source" of which block of tokens.
- purplepatrick 2mo agoYeah, basically everything that becomes context in a session will bias perception and communication style -- subagents, plan lingo, prompt lingo, etc. And then if you write a plan with the comms context having been biased, the lingo will creep into the plan, and from the plan into the code and code comments. And from there, bad lingo will go on multiplying like rabbits... I usually think of it in terms of having a "good" or "bad" session. In a bad session, there is a harmful bias that you can only get rid of through a new session. For example, if you exposed too much context about, say, a variable that features prominently in a doc. The entire session will be anchoring on the importance of that variable. Or if you introduced the notion of CC having to ask for permission for stuff you will have a hard time getting it to "think on its feet" or propose an effective solution (you have made CC so insecure that it now relies on you even for little things that wouldn't normally require your input). In some cases (let's say you have important context in that session) you can overcome this by upping the reasoning level or switching to Fable, but usually a new session is the way to go. Because it's so easy to bias the session I wouldn't even want to use any of these tools that pretend to give Claude "a brain" or "remember" things. That was en vogue a year ago and helpful then, but now, it's plain harmful IMHO. The key is to have just enough context. Subagents often have the reverse problem in that they tend to have too little context to make "judgment calls", which is why the tasks for them must be either deliberately basic or mechanical in nature, or their output should be audited by the main session agent. As for "thinking" it's not clear that that's even a thing (https://arxiv.org/abs/2510.24941 https://arxiv.org/abs/2510.24941)...
- zeafoamrun 2mo agoI think a lot of Claudisms are compressed steering cues for the model’s reasoning: “load-bearing” raises causal importance; “quietly” flags a hidden failure mode; “the one thing” collapses attention onto a discriminator; “on the record” invokes auditability; “at the width the evidence supports” calibrates confidence; “by construction” marks structural inevitability; and “converged” terminates further review loops. They probably be very useful for Claude's chain of thought because they preserve some precise epistemic posture, but are hard for a human to understand. Maybe the final output pass should remove this stuff.
- not_paid_by_yt 2mo agoIt could be, but without evidence that remains a just so story, no particular reason to think it's required or useful or even harmless to the model performance or anything other than an artifact of some early silicon valley writing style being injected into the model and continuous retraining on the output of older models.
- ambicapter 2mo ago> being a coined word or quasi-synonym for something that is already named in the code base. This annoys me with a lot of LLM code. They rename things for the hell of it all the time.
- jsrozner 2mo agoYou can imagine that as people get used to working with Claude, they defer to its judgement. So the people choosing which RL path is better may say "yes, Claude, that was a good refactor!" because it did something hard that it may have been able to superficially justify. Actually the change was unnecessary and complicating. The Claude trainers, as they themselves adapt to Claude's output, are collapsing in their own distribution, so even "new" from-human data is already contaminated.
- not_paid_by_yt 2mo agoWould more blame this on the LLM companies, they think they are on the verge of automating all work, I don't think they care about how you feel about the writing style of the Deus Ex Machina, it's not going to get fixed because to them Claude is already above a staff engineer and soon going to smarter than any human that will ever live. All the money will be going into improvements relevant to improving long context operating and correctness, they could fix the writing style but they are disinterested in that for frontier models, maybe some other companies are but they don't have as much money.
- nailer 2mo ago> I think the issue is that what works well for code (succinctness) doesn’t work well in prosaic English. Hrm, I would have said the oposite. Succint language communicates without unnecessary clutter that could be a barrier to communication. > Biggest issues: dense sentences, constant metaphors, abstractions, and seemingly no understanding of correct anaphora use. And maybe you also agree? I'm confused about your preferred style of language.
- purplepatrick 2mo agoSuccinct doesn't typically mean clutter-free, but hyper-efficient. This works for code, because it (is intended to be) composed of unambiguous semantic units. Regular language, on the other hand, is messy, vague, and requires more structure and context. CC attempts to communicate in English the same way it does in code -- squeezing as much information into as few words as possible, and including justifications for everything, no matter how trivial. To do that, it coins terms and presupposes all of its context exists within the reader also. So, the crux is: CC has no clue what is and isn't "necessary" for a human reader, and teaching it to understand that (if at all possible) is going to be very valuable...
- bonoboTP 2mo agoFable has the same issues, but it's also smarter so I put up with it. Opus is not smart enough for me to tolerate this style. I wonder if putting Opus 4.6 as a frontend communicator that rephrases the blabber of Opus 5 (or Fable) is workable.
- eli_gottlieb 2mo ago> and it’s a real challenge. What model did you use to write this?
- valleyer 2mo agoYeah, "X, and it's Y" is a common trope I see.
- eli_gottlieb 2mo agoIt's so common that I've seen it used as a joke to portray a human trying to write like an LLM.
- Alephinitesimal 2mo agoI’ve been running into this too. It’s especially frustrating when you ask Claude to explain one of its own terms or summaries, and instead of just defining it plainly, it sometimes goes through several rounds of tool calls before giving you a usable explanation. I really don't think such time/tokens should be wasted.
- newAccount2025 2mo ago100%. It is crazy that the default response to everything is act then explain. It starts writing code or running commands and I’m just like my dude wtf are you trying to do, can you just clue me in first.
- Alephinitesimal 2mo agoYeah, exactly. Even when I lower the effort to medium or low, it still tends to act for several rounds before explaining what it’s doing.
- ydant 2mo agoAs a workaround, /btw will force it to answer based solely on context and isn't allowed tool calls
- Alephinitesimal 2mo agoHaha, I only used /btw when the agent was in the middle of doing something. Never thought to just use it directly. Thanks!
- crooked-v 2mo ago> “the x”, with x having not only no antecedent but also being a coined word or quasi-synonym for something that is already named in the code base It does the same thing if you try to get it to do analysis of text or turn data into summaries. It will invent cryptic hyphenated compound words to describe things instead of using plain language or preexisting terms.