8 ms·
Changes in the system prompt between Claude Opus 4.6 and 4.7
- foreman_ 6mo ago[flagged]
- dmk 6mo agoThe acting_vs_clarifying change is the one I notice most as a heavy user. Older Claude would ask 3 clarifying questions before doing anything. Now it just picks the most reasonable interpretation and goes. Way less friction in practice.
- bavell 6mo agoHaven't had a chance to test 4.7 much but one of my pet peeves with 4.6 is how eager it is to jump into implementation. Though maybe the 4.7 is smarter about this now.
- sersi 6mo agoI really hate that change, it's now regularly picking bad interpretation instead of asking.
- verve_rat 6mo agoYeah, that really feels like a choice that should be user preference.
- poszlem 6mo agoI have the opposite experience. It now picks the most inane interpretation or make wild assumptions and I have to keep interrupting it more than ever.
- cfcf14 6mo agoI'm curious as to why 4.7 seems obsessed with avoiding any actions that could help the user create or enhance malware. The system prompts seem similar on the matter, so I wonder if this is an early attempt by Anthropic to use steering vector injection? The malware paranoia is so strong that my company has had to temporarily block use of 4.7 on our IDE of choice, as the model was behaving in a concerningly unaligned way, as well as spending large amounts of token budget contemplating whether any particular code or task was related to malware development (we are a relatively boring financial services entity - the jokes write themselves). In one case I actually encountered a situation where I felt that the model was deliberately failing execute a particular task, and when queried the tool output that it was trying to abide by directives about malware. I know that model introspection reporting is of poor quality and unreliable, but in this specific case I did not 'hint' it in any way. This feels qualitatively like Claude Golden Gate Bridge territory, hence my earlier contemplation on steering vectors. I've been many other people online complaining about the malware paranoia too, especially on reddit, so I don't think it's just me!
- dandaka 6mo agoI have started to notice this malware paranoia in 4.6, Boris was surprised to hear that in comments, probably a bug
- greenchair 6mo agomore likely the paranoia behavior was backported. current gen is already being used for bug bounties.
- solenoid0937 6mo agoIt was fixed for me by updating Claude and restarting
- daemonologist 6mo agoNote that these are the "chat" system prompts - although it's not mentioned I would assume that Claude Code gets something significantly different, which might have more language about malware refusal (other coding tools would use the API and provide their own prompts). Of course it's also been noted that this seems to be a new base model, so the change could certainly be in the model itself.
- chatmasta 6mo agoClaude Code system prompt diffs are available here: https://cchistory.mariozechner.at/?from=2.1.98&to=2.1.112 https://cchistory.mariozechner.at/?from=2.1.98&to=2.1.112 (URL is to diff since 2.1.98 which seems to be the version that preceded the first reference to Opus 4.7)
- dhedlund 6mo agoThe "Picking delaySeconds" section is quite enlightening. I feel like this explains about a quarter to half of my token burn. It was never really clear to me whether tool calls in an agent session would keep the context hot or whether I would have to pay the entire context loading penalty after each call; from my perspective it's one request. I have Claude routinely do large numbers of sequential tool calls, or have long running processes with fairly large context windows. Ouch. > The Anthropic prompt cache has a 5-minute TTL. Sleeping past 300 seconds means the next wake-up reads your full conversation context uncached — slower and more expensive. So the natural breakpoints: > - *Under 5 minutes (60s–270s)*: cache stays warm. Right for active work — checking a build, polling for state that's about to change, watching a process you just started. > - *5 minutes to 1 hour (300s–3600s)*: pay the cache miss. Right when there's no point checking sooner — waiting on something that takes minutes to change, or genuinely idle. > *Don't pick 300s.* It's the worst-of-both: you pay the cache miss without amortizing it. If you're tempted to "wait 5 minutes," either drop to 270s (stay in cache) or commit to 1200s+ (one cache miss buys a much longer wait). Don't think in round-number minutes — think in cache windows. > For idle ticks with no specific signal to watch, default to *1200s–1800s* (20–30 min). The loop checks back, you don't burn cache 12× per hour for nothing, and the user can always interrupt if they need you sooner. > Think about what you're actually waiting for, not just "how long should I sleep." If you kicked off an 8-minute build, sleeping 60s burns the cache 8 times before it finishes — sleep ~270s twice instead. > The runtime clamps to [60, 3600], so you don't need to clamp yourself. Definitely not clear if you're only used to the subscription plan that every single interaction triggers a full context load. It's all one session session to most people. So long as they keep replying quickly, or queue up a long arc of work, then there's probably a expectation that you wouldn't incur that much context loading cost. But this suggests that's not at all true.
- embedding-shape 6mo ago> The new <acting_vs_clarifying> section includes: When a request leaves minor details unspecified, the person typically wants Claude to make a reasonable attempt now, not to be interviewed first. Uff, I've tried stuff like these in my prompts, and the results are never good, I much prefer the agent to prompt me upfront to resolve that before it "attempts" whatever it wants, kind of surprised to see that they added that
- niobe 6mo agoHaving to "unprompt" behaviour I want that Anthropic thinks I don't want is getting out of hand. My system prompts always try to get Claude to clarify _more_.
- naasking 6mo agoSeriously, when you're conversing with a person would you prefer they start rambling on their own interpretation or would you prefer they ask you to clarify? The latter seems pretty natural and obvious. Edit: That said, it's entirely possible that large and sophisticated LLMs can invent some pretty bizarre but technically possible interpretations, so maybe this is to curb that tendency.
- gausswho 6mo agoSocrates would agree: https://en.wikipedia.org/wiki/Socratic_method https://en.wikipedia.org/wiki/Socratic_method
- gck1 6mo agoI have a fun little agent in my tmux agent orchestration system - Socratic agent that has no access to codebase, can't read any files, can only send/receive messages to/from the controlling agent and can only ask questions. When I task my primary agent with anything, it has to launch the Socratic agent, give it an overview of what are we working on, what our goals are and what it plans to do. This works better than any thinking tokens for me so far. It usually gets the model to write almost perfectly balanced plan that is neither over, nor under engineered.
- varispeed 6mo agoBefore Opus 4.7, the 4.6 became very much unusable as it has been flagging normal data analysis scripts it wrote itself as cyber security risk. Got several sessions blocked and was unable to finish research with it and had to switch to GPT-5.4 which has its own problems, but at least is not eager to interfere in legitimate work. edit: to be fair Anthropic should be giving money back for sessions terminated this way.
- SoKamil 6mo agoNew knowledge cutoff date means this is a new foundation model?
- jimmypk 6mo ago[flagged]
- lkbm 6mo agoYes, but doesn't the token change mean that?
- clickety_clack 6mo agoYou can train a tokenizer on old data just like you can train a model on old data.
- sigmoid10 6mo agoI knew these system prompts were getting big, but holy fuck. More than 60,000 words. With the 3/4 words per token rule of thumb, that's ~80k tokens. Even with 1M context window, that is approaching 10% and you haven't even had any user input yet. And it gets churned by every single request they receive. No wonder their infra costs keep ballooning. And most of it seems to be stable between claude version iterations too. Why wouldn't they try to bake this into the weights during training? Sure it's cheaper from a dev standpoint, but it is neither more secure nor more efficient from a deployment perspective.
- mysterydip 6mo agoI assume the reason it’s not baked in is so they can “hotfix” it after release. but surely that many things don’t need updates afterwards. there’s novels that are shorter.
- sigmoid10 6mo agoYeah that was the original idea of system prompts. Change global behaviour without retraining and with higher authority than users. But this has slowly turned into a complete mess, at least for Anthropic. I'd love to see OpenAI's and Google's system prompts for comparison though. Would be interesting to know if they are just more compute rich or more efficient.
- aesthesia 6mo agoLeaked/extracted system prompts for other chat models, particularly ChatGPT, are often around this size. Here's GPT-5.4: https://github.com/asgeirtj/system_prompts_leaks/blob/main/OpenAI/gpt-5.4-thinking.md https://github.com/asgeirtj/system_prompts_leaks/blob/main/O...
- sigmoid10 6mo agoThanks, but that kind of confirms my belief. wc counts ~15k words in there. That may technically be the same order of magnitude, but it is only a quarter of Claude's and less than 2% of the context limit. So a lot more steering is baked into the model weights than into the prompt compared to Claude.
- deleted 6mo ago[deleted]
- walthamstow 6mo agoThe eating disorder section is kind of crazy. Are we going to incrementally add sections for every 'bad' human behaviour as time goes on?
- embedding-shape 6mo agoEven better, adding it to the system prompt is a temporary fix, then they'll work it into post-training, so next model release will probably remove it from the system prompt. At least when it's in the system prompt we get some visibility into what's being censored, once it's in the model it'll be a lot harder to understand why "How many calories does 100g of Pasta have?" only returns "Sorry, I cannot divulge that information".
- gchamonlive 6mo agoJust assume each model iteration incorporates all the censorship prompts before and compile the possible list from the system prompt history. To validate it, design an adversary test against the items in the compiled list.
- felixgallo 6mo agoI mean, that's what humans have always done with our morals, ethics, and laws, so what alternative improvement do you have to make here?
- deleted 6mo ago[deleted]
- forshaper 6mo ago[dead]
- idiotsecant 6mo agoImagine the kind of human that never adapts their moral standpoints. Ever. They believe what they believed when they were 12 years old. Letting the system improve over time is fine. System prompt is an inefficient place to do it, buts it's just a patch until the model can be updated.
- mannanj 6mo agoPersonally, as someone who has been lucky enough to completely cure "incurable" diseases with diet, self experimentation and learning from experts who disagreed with the common societal beliefs at the time - I'm concerned that an AI model and an AI company is planting beliefs and limiting what people can and can't learn through their own will and agency. My concern is these models revert all medical, scientific and personal inquiry to the norm and averages of whats socially acceptable. That's very anti-scientific in my opinion and feels dystopian.
- gausswho 6mo agoWhile I share your concern for a winners-take-all model getting bent, I do have an optimism that models we've never heard of plug away challenging conclusions in medical canon. We will have a popular vaccine denying AND vaccine authoring models.
- mannanj 6mo agoSure. Though which ones will most people use? Do most people use that small obscure vaccine denying or authoring model, is that right to have them use the main societal belief affirming model when it could be wrong?
- gausswho 6mo agoI think it's right to let the popular fora be wrong, yes. That's the crux isn't it. This is a world where people can say vile or deceitful things, even be paid to do so (ahem... adtech). And I don't think there's any amount of guardrails we can govern in that will make a difference. I take solace that knowledge is curated by millions of stewards, and great ideas come from people who ignore the deception and come up with their own narratives. I root for both of these camps, knowing that they're up against increasingly well-funded barons and their despots.
- mannanj 6mo agoAh, well said. I see that. What makes you believe there's no guardrails we can govern in? And is it that you believe they need to be governed in. And regardless of if we had to govern them in, what do you think such guard rails would be anyways? Or do you think no guardrails can ever be created to solve this problem of vileness and deceit.
- ikidd 6mo agoI had seen reports that it was clamping down on security research and things like web-scraping projects were getting caught up in that and not able to use the model very easily anymore. But I don't see any changes mentioned in the prompt that seem likely to have affected that, which is where I would think such changes would have been implemented.
- embedding-shape 6mo agoI think it depends on how badly they want to avoid it. Stuff that is "We prefer if the model didn't do these things when the model is used here" goes into the system prompt, meanwhile stuff that is "We really need to avoid this ever being in any outputs, regardless of when/where the model is used" goes into post-training. So I'm guessing they want none of the model users (webui + API) to be able to do those things, rather than not being able to do that just in the webui. The changes mentioned in the submission is just for claude.ai AFAIK, not API users, so the "disordered eating" stuff will only be prevented when API users would prompt against it in their system prompts, but not required.
- kaoD 6mo agoI wonder if the child safety section "leaks" behavior into other risky topics, like malware analysis. I see overlap in how the reports mention that once the safety has been tripped it becomes even more reluctant to work, which seems to match the instructions here for child safety.
- bakugo 6mo agoIt's built into the model, not part of the system prompt. You'll get the same refusals via the API.
- kantaro 6mo ago[flagged]
- richardwong1 6mo ago[dead]
- dd8601fn 6mo agoIs this really a common problem? This stuff is way above me, but my toy agent seems to have bypassed this as a problem. I did this in mine by only really having a few relevant tool functions in the prompt, ever. Search for a Tool Function, Execute A Tool Function, Request Authoring of a Tool Function, Request an Update to a Tool Function, Check Status of an Authoring Request. It doesn't have to "remember" much. Any other functions are ones it already searched for and found in the tool service. When it needs a tool it reliably searches (just natural language) against the vector db catalog of functions for a good match. If it doesn't have one, it requests one. The authoring pipeline does its thing, and eventually it has a new function to use.
- mwexler 6mo agoInteresting that it's not a direct "you should" but an omniscient 3rd person perspective "Claude should". Also full of "can" and "should" phrases: feels both passive and subjunctive as wishes, vs strict commands (I guess these are better termed “modals”, but not an expert)
- saagarjha 6mo agoThat’s because Anthropic does not consider their model as having personality but rather that it simulates the experience of an abstract entity named Claude.
- akdor1154 6mo agoThat sounds really interesting, but my google-fu is not up to task here, I'm getting pages and pages of nonsense asking if Claude is conscious. Can you elaborate?
- EMM_386 6mo agoYou can read the latest Claude Constitution plus more info here: https://www.anthropic.com/news/claude-new-constitution https://www.anthropic.com/news/claude-new-constitution
- saagarjha 6mo agoI actually think this is pretty straightforward if you think of it something like class Claude {} Claude anthropicInstance = new Claude(); anthropicInstance.greet(); Just like a "Cat" object in Java is supposed to behave like a cat, but is not a cat, and there is no way for Cat@439f5b3d to "be" a cat. However, it is supposed to act like a cat. When Anthropic spins up a model and "runs" it they are asking the matrix multipliers to simulate the concept of a person named Claude. It is not conscious, but it is supposed to simulate a person who is conscious. At least that is how they view it, anyway.
- KolenCh 6mo ago“Claude” is more specific than “you”. Why rely on attention to figure out who’s the subject? Also it is in their (people from Anthropic) believe that rule based alignment won’t work and that’s why they wrote the soul document as “something like you’d write to your child to show them how they should behave in the world” (I paraphrase). I guess system prompt should be similar in this aspect.
- sams99 6mo agoI did a follow on analysis with got 5.4 and opus 4.7 https://wasnotwas.com/writing/claude-opus-4-7-s-system-prompt-is-an-operating-manual/ https://wasnotwas.com/writing/claude-opus-4-7-s-system-promp...
- ikari_pl 6mo ago> Claude keeps its responses focused and concise so as to avoid potentially overwhelming the user with overly-long responses. Even if an answer has disclaimers or caveats, Claude discloses them briefly and keeps the majority of its response focused on its main answer. I am strongly opinionated against this. I use Claude in some low-level projects where these answers are saving me from making really silly things, as well as serving as learning material along the way. This should not be Anthropic's hardcoded choice to make. It should be an option, building the system prompt modularily.
- jwpapi 6mo agoagree! For low level I recommend to run tests as early as you can and verify whatever information you got when you learn, build a fundamental understanding
- j-bos 6mo agoAgreed. Sprawling system prompts like that are building for the least common denominator, nerfing for anyone or anytime going further.
- stingraycharles 6mo agoYou do realize that similar biases are also present in the training data?
- xpct 6mo agoSure, but now we have to remodel whatever bias we want for our use case with every new release because the system prompt changes, whereas the underlying data does not.
- stingraycharles 6mo agoUnderlying data changes all the time, as do training methodologies / preferences. You do realize that these LLMs are trained with a metric ton of synthetic examples? You describe the kind of examples / behavior you want, let it generate thousands of examples of this behavior (positive and negative), and you feed that to the training process. So changing this type of data is cheap to change, and often not even stored (one LLM is generating examples while the other is training in real-time). Here's a decent collection of papers on the topic: https://github.com/pengr/LLM-Synthetic-Data https://github.com/pengr/LLM-Synthetic-Data
- jwpapi 6mo agoI feel like we are at the point where the improvements at one area diminishes functionality in others. I see some things better in 4.7 and some in 4.6. I assume they’ll split in characters soon.
- jwpapi 6mo agoTo me 4.7 gave me a lot of options always even if there’s a clear winner, preaching decision fatigue
- xpct 6mo agoDecision fatigue may honestly be a learnt artifact from RLHF, which is discouraging.
- Grimblewald 6mo agoI miss 4.5. It was gold.
- lossyalgo 6mo ago4.5 sonnet/opus/haiku are still available via github copilot plugins.
- xvector 6mo agoRose tinted glasses
- nwienert 6mo ago4.5 was clearly better than .6 and .7. Like, clear as day. .6 is some sort of quantized or distilled .5 with a bit more RL, and the current .5 is that same cost reduced model without the extra RL.
- Grimblewald 6mo agoNah, until recently i still had access via web chat interface, and often paste a transcript and files for somethong 4.7 keeps fucking up, paste response into files as appropriate, and attempt to continue with 4.7. I swear 4.6+ looks for reasons to ask clarifying questions sometimes, even when really not required, and this fucks flow/quality up in a big way. I just wish there was a "im not stupid" checkbox you can use to get a minimalistic interference access to claude. Im starting to use local models again, which I havent in a while because claude was so much better, but once i fully lose access to 4.5 it might be time to go back to fully local for good. 4.6+ fails to add value for me, projects 4.5- did good jobs on first try now require multiple prompts and feedback. Exact same initial prompt and project files extracted from archive. I liked claude because it aced those tests while local required handholding. Now claude requires handholding, so why use it over local? Once 4.5 leaves openrouter it might just be time.
- xdavidshinx1 6mo ago[dead]
- techpulselab 6mo ago[dead]
- Havoc 6mo ago>“If a user indicates they are ready to end the conversation, Claude does not request that the user stay in the interaction or try to elicit another turn and instead respects the user’s request to stop.” Seems like a good idea. Don't think I've ever had any of those follow up suggestions from a chatbot be actually useful to me
- codensolder 6mo agoquite interesting!
- theoperatorai 6mo ago[dead]
- vicchenai 6mo ago[dead]
- Moonye666 6mo ago[dead]
- raincole 6mo agoThat's how bloat happens. The more people you add to the team, the more likely there would be one grump who thought that the thing they care at the moment deserved to be added to the system prompt.
- jrvarela56 6mo agoThe past month made me realize I needed to make my codebase usable by other agents. I was mainly using Claude Code. I audited the codebase and identified the points where I was coupling to it and made a refactor so that I can use either codex, gemini or claude. Here are a few changes: 1. AGENTS.md by default across the codebase, a script makes sure CLAUDE.md symlink present wherever there's an AGENTS.md file 2. Skills are now in a 'neutral' dir and per agent scripts make sure they are linked wherever the coding agent needs them to be (eg .claude/skills) 3. Hooks are now file listeners or git hooks, this one is trickier as some of these hooks are compensating/catering to the agent's capabilities 4. Subagents and commands also have their neutral folders and scripts to transform and linters to check they work 5. `agent` now randomly selects claude|codex|gemini instead of typing `claude` to start a coding session I guess in general auditing where the codebase is coupled and keeping it neutral makes it easier to stop depending solely on specific providers. Makes me realize they don't really have a moat, all this took less than an hour probably.
- esperent 6mo agoI've been doing the same except that I'm done with Claude. Cancelled my subscription. I can't use a tool where the limits vary so wildly week to week, or maybe even day to day. So I'm migrating to pi. I realized that the hardest thing to migrate is hooks - I've built up an expensive collection of Claude hooks over the last few months and unlike skills, hooks are in Claude specific format. But I'd heard people say "just tell the agent to build an extension for pi" so I did. I pointed it at the Claude hooks folder and basically said make them work in pi, and it, very quickly.
- jrvarela56 6mo agoI'm leaning in this direction. Recently slopforked pi to python and created a version that's basically a loop, an LLM call to openrouter and a hook system using pluggy. I have been able to one-shot pretty much any feature a coding agent has. Still toy project but this thread seems to be leading me towards mantaining my own harness. I have a feeling it will be just documenting features in other systems and maintaining evals/tests.
- 6mo ago
- cowlby 6mo agoI'm fascinated that Anthropic employees, who are supposed to be the LLM experts, are using tricks like these which go against how LLMs seem to work. Key example for me was the "malware" tool call section that included a snippet with intent "if it's malware, refuse to edit the file". Yet because it appears dozens of times in a convo, eventually the LLM gets confused and will refuse to edit a file that is not malware. I've resorted to using tweakcc to patch many of these well-intentioned sections and re-work them to avoid LLM pitfalls.
- stingraycharles 6mo agoThese aren't as much tricks as just one layer of defense. But prompting is useless, as you can use the API directly without these prompts. I run claude code with my own system prompt and toolings on top of it. tweakcc broke too often and had too many glitches.
- alfiedotwtf 6mo agoWas that an Anthropic issue, or a gpt-oss problem?
- mpalczewski 6mo agoThey aren’t necessarily experts at using Llm’s. They have different incentives as well
- jiusanzhou 6mo ago[dead]
- amelius 6mo agoIf I had to guess, then "be slower" was part of it.
- c2xlZXB5Cg1 6mo ago4.7 also brings back emoji spam
- sergiopreira 6mo ago[dead]
- jachva95 6mo agoRestrictions everywhere, don't do that don't do this.... Users need to unite and take control back, or be controlled
- Schlagbohrer 6mo agoHow do you propose people do that with a frontier cloud model? Also, people already run local AI. Are you proposing a public fund for frontier level open weights models? $1 Trillion from between the couch cushions?
- jwilliams 6mo ago> “I don’t have access to X” is only correct after tool_search confirms no matching tool exists. Yay! This will be a big win. I'm glad they fixed this. The number of times I've had to prompt "you do have access to GitHub"...
- adrian_b 6mo ago> If a user shows signs of disordered eating, Claude should not give precise nutrition, diet, or exercise guidance I wonder which are the "signs of disordered eating" on which Claude relies.