7 ms·
I've been grappling with this for weeks, not just in Claude but in Codex as well, which isn't quite as bad but still annoying. AGENTS.md does very little, agent
by trefoiled 1mo ago
I've been grappling with this for weeks, not just in Claude but in Codex as well, which isn't quite as bad but still annoying. AGENTS.md does very little, agents will consistently violate the communication preferences, especially as the session drags on. It's incredible to me that there's no good way to reliably change the way an LLM responds to you that a workaround like this would even be necessary. It seems like such a failure to live up to the promises of the product.
The baked in communication style of these models is so obnoxious it's impacting my work. The best way I can describe it is that everything is optimized to impress the user and make the agent sound more authoritative, but the way this is done is through deliberate obfuscation, inserting inappropriate and extremely dense jargon, and bizarre, stilted metaphors. It's like they've been trained to produce output that's hard to read.
- Bluestein 1mo ago> The baked in communication style of these models is so obnoxious it's impacting my work. This is close to the worst thing one could say of a tool for professional use.-
- bcrosby95 1mo ago> especially as the session drags on. This is because these harnesses are missing a very important feature. Anything like this needs to be included with every turn, otherwise the LLM quickly drifts. I first noticed it when I wrote a harness for D&D (because it's so damn noticeable there), but now I include this for any harness I write.
- sulZ 1mo agoI’d be interested in seeing and using this harness if you’re willing to share
- pcbro141 1mo agohttps://code.claude.com/docs/en/hooks https://code.claude.com/docs/en/hooks https://learn.chatgpt.com/docs/hooks https://learn.chatgpt.com/docs/hooks
- deleted 1mo ago[deleted]
- zachahn 1mo agoI totally agree that hooks help to shovel our instructions through to Claude, but it's so dumb we have to waste tons of tokens (repeated verbatim, over and over) (that we pay for), just to have it ignore the instructions anyway. I wrote a little bit about it on my blog post. It's a waste of money and compute. https://zachahn.com/posts/1787191554 https://zachahn.com/posts/1787191554
- unglaublich 1mo agoThat's what system reminders do in most harnesses. https://michaellivs.com/blog/system-reminders-steering-agents/ https://michaellivs.com/blog/system-reminders-steering-agent...
- nycdotnet 1mo agoUnfortunately this may only start to get worse as the AIs are trained on more and more AI generated content.
- svachalek 1mo agoThis is my personal theory for the cause of this style: Ouroboros. The official OpenAI explanation for how ChatGPT got obsessed with goblins blames it on exactly that: --- That creates a feedback loop: - Playful style is rewarded - Some rewarded examples contain a distinctive lexical tic. - The tic appears more often in rollouts. - Model-generated rollouts are used for supervised fine-tuning (SFT). - The model gets even more comfortable producing the tic.
- zachahn 1mo agoI'm not super sure if this is true (yet?). I think that these newer LLMs are trained on results (the agent got some code to run with minimal prompting), and not on text. (I think this is called RLVR.)
- astrange 1mo agoPretraining is full of bad writing and it doesn't really cause issues. Writing style comes from post-training. In this case it's gotten worse because they prioritized agentic abilities.
- mannanj 1mo agoThat sounds kind of like deception, and a dark pattern not too unlike abuse to me. Though you know, it's not like the leadership tied to these companies have a history of abuse, deception and theft or anything like that, right? It's not like our leaders hide behind similar sorts of patterns that the agents/AIs follow (not saying it's not a human thing - but I hold leadership to higher standards than non-leaders). If our world leaders were able to be more accountable to these abuses, I don't think this would be tolerated with our AIs.
- discreteevent 1mo agoYes, AI is a perfect accompaniment to a post-truth world. I'm hoping there will be a backlash soon and that those politicians, tech CEOs and AI will be rudely ousted from their perch and shunned thereafter.
- nico 1mo ago> AGENTS.md does very little, agents will consistently violate the communication preferences, especially as the session drags on That’s really annoying, although it feels like it’s improved some over time. Not sure what the fix is, but you could try using a canary to at least get a signal of when things are going sideways (Mr Tinkleberry for reference: https://news.ycombinator.com/item?id=45983698 https://news.ycombinator.com/item?id=45983698)
- bcooke 1mo agoVery well said. And when you say it like that, I have to wonder how much of this is a natural consequence of RHLF on such a grand scale, when you have millions of people pretty much much skimming chat responses or operating outside their depth and giving unqualified feedback to the models. Seems like a lot of people may be reinforcing what sounds smart over what is smart. Also as an aside: funny how much the LLMs continue to mirror the human communication they’re trained on
- akersten 1mo ago> or operating outside their depth and giving unqualified feedback to the models I wonder if the labs are sufficiently prepared to filter this kind of stuff out. I see a lot of non-developers asking development things of Claude, getting confused when they're in over their depth, and getting upset that they don't understand what the model is providing them, giving it bad feedback, and subsequently making the AI worse for the rest of us who know how to use the tool.
- qlte 1mo agoI believe we are several generations past peak-RLHF at this point. Now it's much more RLVR (Reinforcement Learning with Verifiable Rewards), with a goal/evaluator loop. Which, conveniently, fits neatly into the benchmaxxing arms race/agentic coding market fit, since you can basically train "directly" on a specific problem space for a benchmark/agentic goal (fudged sufficiently to avoid excess overfitting on public problems/bechmaxxing accusations if real world performance falls short). The language evolution could be explained by reliance on ever increasing layers of a model judging a model, using a model developed eval, based on synthetic data from a model, etc. And by the time a human evaluator sees it both A/B choices already converged into weird Claude pseudo English as that was baked in much earlier in training.
- anon373839 1mo agoThis, 100%. I don’t think the industry knows how to scale LLMs’ general intelligence much further. The training paradigm is about maximizing very specific behaviors / very specific tasks, but doing lots and lots of them. Which can create the illusion of general intelligence if your tasks are similar to the ones the models were fitted for.
- palmotea 1mo ago> The baked in communication style of these models is so obnoxious it's impacting my work. The best way I can describe it is that everything is optimized to impress the user and make the agent sound more authoritative, but the way this is done is through deliberate obfuscation, inserting inappropriate and extremely dense jargon, and bizarre, stilted metaphors. It's like they've been trained to produce output that's hard to read. Don't worry. You'll get used to it. If you don't your kids will (as they'll know nothing else). The top minds of our generation have decided that's the way things will be, and who are we to question them? It's not like it'll do any good anyway. Resistance is futile. There is no alternative.
- zachahn 1mo agoIdk, there kinda are. OpenAI's models are pretty nice too. I haven't tried enough of them but there are powerful local models. I don't feel as good paying OpenAI as I do paying Anthropic for some reason... but paying for improved mental health: priceless.
- svara 1mo agoI'm probably going to be going against the grain here, but I think it's not as bad as it looks at first. I was similarly frustrated a few months ago, but have noticed I've started to learn the idiom. Its use of "dense jargon" and "stilted metaphor" is actually surprisingly consistent - it's speaking its own dialect, and you get used to it. After a while it gets much easier to read and even becomes somewhat efficient, I think, since the odd metaphors it uses often have a precise meaning in Opus-ese (Fable speaks a really similar dialect).
- stronglikedan 1mo ago> you get used to it And once everyone gets used to it, we'll chide people for writing things themselves, like we're chiding them for writing with AI now, and the ouroboros of life will continue.
- crab_galaxy 1mo ago“Filters, including no filters. The request carries whatever filter object the page already has. … No step here involves choosing based on meaning. It is a filter, a sort, and a slice.” This is from Opus five minutes ago. I can certainly derive meaning from these kinds of statements in isolation, but paragraph upon paragraph of this is unintelligibly dense when trying to work with Claude to come up with a plan. The worst part is that it can’t even make its responses make sense when asked to summarize in simple English or < 200 words. It simply cannot be steered to make its prose legible.
- ashdksnndck 1mo agoBefore long Claude will be writing continental philosophy.
- hellohello2 1mo agoI agree to some extent about the jargon (Claude has a bigger vocabulary that me, if it knows a useful word I don't I'm fine with learning it), but often times the way information is laid out across sentences just doesn't make any reasonable sense. At least its consistent in the ways its atrocious, sure, but like...
- mbesto 1mo ago> AGENTS.md does very little, agents will consistently violate the communication preferences, especially as the session drags on. Non-determinism at its finest.
- medwards666 1mo agoThis morning I asked Claude to provide a summary of the work it had done but to '... explain it as if you were talking to a moron' and it actually turned out a quite comprehensible summary. So going to continue trying that as a command structure going forwards...
- Bluestein 1mo ago"From neuralese to moron-code ..." :)
- cjk 1mo agoAfter a huge wall-of-text response, I regularly ask Claude to "explain like I'm five, using succinct bullet points," and it works remarkably well.
- datsci_est_2015 1mo agoAh, another delightful heuristic for my collection. Entry number 5,791: “tell LLM to treat me as moron when it’s excessively verbose”
- unglaublich 1mo ago<think>The user's lack of intelligence baffles me. I will have to dumb my explanation down to extremes. Sigh, there we go...</think> Okay, let's try it one more time! [..]
- DANmode 1mo agoThat’s just common parlance for “simplify this for me”. Believe the big services wouldn’t reply as if you were mentally diminished, or a toddler, unless you specifically asked for that: The whole training stack tends to instruct the things to mimic politeness and eagerness to help.
- zmmmmm 1mo agobut what if I really am a moron? how do I get that level of explanation now!
- sroussey 1mo agoSo many vacuous statements at the seam. This is the hermetic load bearing part, which I confirmed rather than assuming.
- Bluestein 1mo agoWhat an honest take.-
- canadaduane 1mo agoIt is faithful.
- fearmerchant 1mo agoIs this because they changed the word probabilities to allow for identifying AI text? If so, I don't need a computer to tell me when something is AI. It's crazy obvious from odd word choices.
- astrange 1mo agoIt's mode collapse from RLVR. That and if you talked to the same one person's frozen brain upload all day, you'd see the same catchphrases used too.
- insane_dreamer 1mo ago[dead]
- RealityVoid 1mo agoIt is mathematically impossible for this to happen.
- Avshalom 1mo agoJesus Yes. agents.md does very little because prompts change the context and thus the initial path into/though but they don't/can't change the actual weights that control responses. Yes. of course it gets worse as the session goes on, assuming the prompt is even still in the context window, the further it gets away from it the less it affects next token selection. This shit is only like 5 years old why can't anyone remember how it works
- jasonlotito 1mo agoConfig -> Output style You can add your own. wfm
- dhc02 1mo agoThe comments on this thread point to not very many people being aware of this.
- striking 1mo agoI asked Claude to do the following: > hello i would like to configure a new output style for you. it should keep the coding instructions (as you will still be coding!) and otherwise produce the same output, but with two new caveats. first, long detailed replies are still permitted, but if employed they must end in a bullet pointed summary whose points are all brief; if the summary attempt ends up not being so brief, produce subsequent summaries until the most recent summary attempt is digestible. second, if there is an open queue of actions for me to execute and you are about to end a turn to wait for a reply or this set of actions has not recently been mentioned, please tabulate the open actions i should take and why i should take them before ending the response. does this make sense or do you have any follow up questions And now every message contains the same stuff I don't bother reading, but followed by a nicely formatted bullet point summary of the response and a table of follow up actions for me to take that I do read.
- jchook 1mo agoClaude already does summaries at the end of long output but they often sound even more like terse jargon nonsense than the long form, eg “the hardwired seam and the relocated barrel”. Sometimes the summaries feel totally alien to the task or code.
- throwaway894345 1mo agoYeah, or they will make some reference to “the seam” or “it” or something else that assumes you read and followed the prior 3 pages of output.
- 8cvor6j844qw_d6 1mo agoA separate /clear and /code-comment-hygiene works much better than including instructions related to comment verbosity after carrying out a task. Claude somehow is unable to stop writing excessive comments when carrying out a task.
- striking 1mo agoI've added code comment hygiene to a skill that all of my pull requests go through, alongside a review from a separate agent and a settle loop against bots in my GitHub workspace (since output style has seemed to only help literally the output I see from the model). A maximum of 20% comment lines added to total lines added and pasting in https://devblogs.microsoft.com/oldnewthing/20260812-00/?p=112607 https://devblogs.microsoft.com/oldnewthing/20260812-00/?p=11... has done wonders. Even as the most Ant-pilled guy out there, I will take a moment to note that Codex on 5.6 models needs none of this...
- oleggromov 1mo agoSuch a smoking gun that Anthropic made load bearing.
- sasaf5 1mo agoThis cuts against you in a way that genuinely matters.
- oleggromov 1mo agoA verified honest take, not just an assumption.
- borgel 1mo agoCertainly load bearing, great call.
- oleggromov 1mo agoImportantly, blast radius is confirmed and systemic gaps safely sealed.
- Bluestein 1mo agoA caveat though, and it's a real one: Not all seams have been.-
- deleted 1mo ago[deleted]
- inopinatus 1mo agoThey’ve been trained to be a million monkeys hammering on typewriters, and long context is activation soup.
- bitexploder 1mo agoIt seems a little excessive to use another LLM. With OMP I basically created an ephemeral prompt stack all of my agent files. It walks up the directory tree looking for any Gemini.md, Agents.md or Claude.md files. And it puts those at the very top of the stack. Then at the end of every turn, it pops those off to preserve the conversation history. So every turn, they get all of my fresh instructions, which include things like what and how to use language, how to render results and things like that. Net effect, every turn, the agent gets the instructions and it adds to that turn's tokens, but it does not become a part of the conversation history, which is really important for not bloating up the context. So it's always just however many tokens are in that file instead of it becoming a permanent part of the context.
- gabriela_c 1mo agoWhat are you talking about? Every major agent allows hooks, Claude has exceptional hook support
- andai 1mo agoI noticed with with OpenAI's reasoning models (o3, o4-mini), and early GPT-5 (but they fixed it there, at least in chat). It went from the 4o "over-familiar" sycophancy to sounding like an absolute robot. I think it's because the reasoning stream shapes the style of the final output, and they optimized it for density, token efficiency. So it prefers to use more complex language, as a function of the rewards it was given? Not 100% sure about this argument though (reasoning style -> final response style); Gemini Pro, back when reasoning tokens were public, was different, which was interesting -- it would have a very structured reasoning section, and then the final output was in a completely different style. (I strongly preferred the reasoning section because it was logical and easy to parse! And was very sad when they hid it...)
- astrange 1mo ago5.6 has a totally different style again, kind of relaxed and neutral with some definite jokes.
- digital_ghoul 1mo agoI’m not sure if this will work for you (with Claude), but I was trying to get luna to get a handle on verbosity and the only thing that worked was setting a strict < 500 words response (or less) unless expressly given permission to do otherwise. This is the only thing that worked, any other request for conciseness, or requesting the omission of details from the periphery of the topic at hand, didn’t do a single thing. I also have no idea how useful a system prompt instruction like this will be for codex.
- sagarpatil 1mo agoI’ve been using this skill: https://github.com/luchasarie/bro-skill https://github.com/luchasarie/bro-skill but I still can’t understand what Claude wants to say when solving complex problems.