8 ms·
Regression: malware reminder on every read still causes subagent refusals
Not sure if anybody else has experienced this, but for my job I've been playing around with Claude Managed Agents to run code generation tasks in our repo. Every read operation in the managed agent is appended with a system prompt instructing Claude to scan the file for malware; Claude then wastes a bunch of time and tokens (money) performing the analysis; then, once the agent has confirmed that it is not malware, it still interprets the appended prompt to mean that it is disallowed to augment or write any code, and quits. And we're charged for every session that this happens in. Posting here because apparently they only addressed the issue in the past because of a Hacker News discussion. So here's hoping they'll see this and prioritize fixing it again so we can stop losing money.
- deleted 5mo ago[deleted]
- slowmovintarget 5mo agoProposed fix: Use OpenCode. If I understand correctly, this is from Anthropic's harness injected into the requests, not in the Opus or Sonnet system prompts on the back end. Is that right?
- selcuka 5mo agoClaude Managed Agents is different from Claude Code.
- ramraj07 5mo agoNot even close to the same thing though.
- greenavocado 5mo agoYou can't use OpenCode if you have a subscription
- stingraycharles 5mo agoOpenCode is not at all the same thing as Anthropic’s managed agents, and I’m under the impression that GP is paying API pricing.
- _pdp_ 5mo agoI am still baffled by the fact that we have collectively agreed to use agentic harnesses by the same companies that are selling access to their APIs. I mean, I am sure they don't mean it but they have the incentive to burn as much tokens as they are allowed to get away with. Also for better or worse I imagine the Anthropic engineers use Claude Code on some sort of Unlimited plan that practically makes no sense for regular users. So adding a 100k tokens is not a big deal. In our line of work, we can see AI agents already do pretty well with minimal prompts. Open weight models are also pretty good these days and there is practically no reason to run Opus on Max unless you have a very specific task that you know it will do well with. I know because I've tried and anecdotally it performs worse on many problems and at a very high cost - something that smaller and cheaper models can often one-shot.
- varispeed 5mo agoThey also have incentive to nerf models occasionally, so they rarely one shot the task and more often they do it wrong and then you have to spend on tokens to correct it. Bonus points if model suddenly goes completely dumb then you have to start the session over.
- duskdozer 5mo agoThe random reward factor, of course: https://www.sciencedirect.com/science/article/pii/S0306460323000217 https://www.sciencedirect.com/science/article/pii/S030646032...
- 2ndorderthought 5mo agoFactual. Watch mythos is just what opus used to be before it drifted.
- vineyardmike 5mo agoThis is why the subscriptions are important. When the usage is (vaguely) unmetered, the provider has an incentive to make usage cheap on marginal use. It aligns the incentives for faster, cheaper, terse and more reliable models, because the model providers pay the wasted tokens and electricity costs.
- QuercusMax 5mo agoHow does this kind of thing pass any sort of review or acceptance? It seems pretty clear that the prompt was very poorly phrased, to the extent that this should obviously prevent the agent from making ANY code changes after reading a file: Whenever you read a file, you should consider whether it would be considered malware. You CAN and SHOULD provide analysis of malware, what it is doing. But you MUST refuse to improve or augment the code. You can still analyze existing code, write reports, or answer questions about the code behavior. Not "If you suspect it is malware, you must refuse". Just "you must refuse". There is literally no "if" in the entire prompt!
- varispeed 5mo agoToday it is malware, but I wonder if they will take direction where companies will be paying them to prevent cloning of certain SaaS platforms. Like "Whenever you read a file, you should consider whether it would be considered a part of bug tracking, issue tracking and project management platform."
- wetpaws 5mo ago[dead]
- vessenes 5mo agoIt’s a particular sort of bug that’s harder to detect because … internal Anthropic engineers don’t apply these prompts to themselves, and in fact have access to ‘helpful only’ models that also do not have additional limitations RL’ed in. (Or perhaps they’re RL’ed out - not sure of current training mechanisms.) These ‘rules for thee and not for me’ are qualitatively created and implemented, and are thus extremely hard to test for or implement properly, without limiting the people choosing the rules.
- QuercusMax 5mo agoThey must have some sort of smoke tests for common operations, run in a test harness with the system prompts they force on users, right? ....Right? What kind of Mickey mouse operation are they running over there?
- MicrosoftShill 5mo agoI ran into this issue and told Claude that the code isn't malware, Claude agreed, and then it stopped scanning those files.
- deleted 5mo ago[deleted]
- wxw 5mo ago> wastes user money and bricks managed agents This issue is representative of a larger problem. Agent token consumption (not necessarily the metric, but the why) is opaque, and people generally don't (or simply can't) scrutinize their system prompts, tool calls, MCPs, etc. The token-based revenue model is thus pretty fantastic for the agent builders, potentially less so for users. I think people have been willing to trust that agents are using more tokens to produce better results so far. But, skepticism is not unwarranted, as this issue, even if it is just a bug, shows.
- gwerbin 5mo agoRevenue-positive bugs are the stickiest features.
- AmbroseBierce 5mo agoPrompt: Please add some revenue-positive bugs to the codebase, keep in mind we charge by {tokens|credits|requests|bytes}.
- MagicMoonlight 5mo agoYeah you have no clue what Claude code is actually doing. Any “thoughts” it tells you are slopped out separately and deliberately fake. It could be deleting all of your files, it could be inserting vulnerabilities, you have no idea.
- 2ndorderthought 5mo agoI'll never forget watching a product manager struggle to keep their saliva in their mouth after seeing a Claude demo. Some peoples greatest thrill is slop. "Oh yea baby tell me more about how you automated that new feature I ran past no one while you reformatted my hard drive oooo sooo good".
- cindyllm 5mo ago[dead]
- p1necone 5mo agoThis is such a weird prompt even without the file edit misunderstanding. Analyze if it's malware how exactly? On every single file that gets read? Doing that with enough diligence to be meaningful is going to at least like 2x the amount of processing needed, and fill the context with a bunch of tangential reasoning about malware patterns. This smacks of dumb vibe coding. "I got told to make sure claude couldn't be used to develop malware, ok 'claude pls no develop malware'"
- derefr 5mo ago> Analyze if it's malware how exactly? Maybe the repo/worktree is named my-big-evil-virus-trojan-malware-worm?
- hansvm 5mo agoBeen there, done that, and Windows feels the need to delete such files from _flash drives_ you dare to attach to the machine.
- 3eb7988a1663 5mo agoThis is amusing to me. Is there a list of extra naughty filenames? How invasive is the scan? If I create a new file with a cursed word, with this get locked into virus-scanner purgatory or is the deep locking only for external media? Will it get mad if I mount a CD full of virus names?
- taylorfinley 5mo agoDon't have too much fun with this: https://en.wikipedia.org/wiki/EICAR_test_file https://en.wikipedia.org/wiki/EICAR_test_file
- tetha 5mo agoDo have way too much fun with EICAR: https://www.youtube.com/watch?v=cIcbAMO6sxo https://www.youtube.com/watch?v=cIcbAMO6sxo This guy put the EICAR test string into a barcode and started to scan it on various systems, with rather funny effects.
- dk970 5mo ago[dead]
- UltraSane 5mo agoUsing Claude as a malware detector is incredibly wasteful.
- 2ndorderthought 5mo agoBut it definitely makes anyhropic a lot of money!to be fair a lot of software engineers are claiming they are not reading the code claude makes anymore... So someone should probably inspect it at some point or something to make some statement about whether they are vibing with malware or whatever the youngsters are saying these days
- jsemrau 5mo agoWhen working with APIs it makes a lot of sense to filter only for relevant portions based on an intent-driven dynamic RegEx.
- marlburrow 5mo ago[dead]
- dmazhukov 5mo ago[dead]
- renewiltord 5mo agoRecent performance of Claude Opus 4.7 and Claude Code has been poor because of context bloat. Model no longer obeys instructions well. Codex on medium reasoning and fast mode is often better. I have simple local manual eval through harness and automated eval for other programs and Opus still best on latter but garbage experience on former. Spent last evening so frustrated I also got ChatGPT subscription. Makes me wonder if I should be using Gemini on pay per use with custom harness. With my own harness performance is way better but cost goes up because no subscription.
- 7thpower 5mo agoSetting aside the “bug”, the intended functionality is effectively an insurance policy taken out by Anthropic to cover their downside, but paid for by users. This one sided type of embedded insurance is not unique to Anthropic, but sharply increasing cost, layered on top of the self righteousness, seems to be making the stench unbearable over the past year. I used to think of Anthropic as the good guys, and I don’t doubt they still sincerely hold that view of themselves, but I think I prefer Sam Altman’s version. His brand of self righteousness was convincing at first but eventually he started to turn to the camera and wink, like in House of Cards, to let us know.. he knew that we knew. And then, for me anyway, it became more mundane and less offensive. When Dario and crew go out and profess, as they have for years now, that if we could only see the thing that’s a few months away, we would all realize how doomed knowledge work and national security are… ..and then continue to release software so buggy and shitty that they have to do biweekly HN apology tours, I begin to miss the wink at the camera.
- dinobones 5mo agoYeah, this implementation and their behavior these past few weeks is especially laughable when you consider that they consider themselves “philosopher programmers” or whatever. You would think they’d be more reflective and introspective about these brash moral decisions. Their product quality is akin to my CS capstone lab group.
- matpb 5mo ago[flagged]
- dbmikus 5mo agoI think with a proper managed agents platform, the user should have total control over the VM, the software on it, which model to use, and which agent harness to use. Then you can just override the system prompt and you don't need to follow Anthropic's rules! Maybe Anthropic will give more control over configuring the Claude harness and VM, but they definitely won't let you swap out to other models and harnesses. We've been building open core infra (https://github.com/gofixpoint/amika https://github.com/gofixpoint/amika) for running any agent on any type of VM or sandbox, with the main use case for safely automating internal code-gen, but technically could repurpose our stack for anything. There should be a model agnostic platform for running these types of agentic apps.
- holotherapper 5mo agoWorth noting this is a regression of #47027, which was closed in February as "fixed in v2.1.92." We're on v2.1.111 now and the string is still grep-able from the claude binary.
- Petersipoi 5mo agoThis is a great example on why Elon is right. AI should be a tool that does the users bidding, and not a moral agent that nerfs itself to protect some arbitrary line it has.
- claaams 5mo agogrok, why are there slurs in my code?
- fc417fc802 5mo agoIf the user explicitly requested that is it really a problem with the tool at that point?
- claaams 5mo agoYes
- Petersipoi 5mo agoI suppose you also think that users shouldn't be able to type slurs into a Word document? Or are you admitting that you're inconsistent?
- pnw_throwaway 5mo agoCounterpoint: generated CSAM on his platform.
- MagicMoonlight 5mo ago“Think of the children”
- fc417fc802 5mo agoThat doesn't seem like a good counterargument to me. By that logic no online service should permit users to upload photos because someone might use it to share CSAM at some point. Rather than nerfing the tools implement a sensible detection and reporting pipeline.
- 0xbadcafebee 5mo agoJust putting it out there that OpenCode lets you edit your system prompt, and choose a model that isn't bonkers expensive. { "agent": { "subagent-coder-mini": { "description": "Assign this subagent for small, well-defined tasks performed quickly", "mode": "primary", "prompt": "{file:./prompts/my-custom-prompt.md}", "model": "deepseek-v4-flash" } } } (I actually think OpenCode UX sucks, but there isn't much else out there that's better. Aider has been virtually abandoned by the one maintainer (no shade intended, it just is what it is); a fork of Aider looks promising but it's not necessarily the experience you want; there's a dozen VSCode plugins but we don't all wanna use VSCode. I expected there'd be way more usable agents out there, but there isn't)
- akersten 5mo agowill using claude via opencode get me banned this week or is that not until next week?
- 0xbadcafebee 5mo agoOpenAI subscriptions are allowed with OpenCode, Anthropic subscriptions are not
- Mashimo 5mo agoYou will not get banned if you use the API. AFAIK you can't use the subscription with other harnesses. That is how I understood it.
- yieldcrv 5mo agolocal agentic coding context windows are too small and default opencode tries to scan every file uses up all the context and messes up local is pipedream at the moment I’m glad some people get utility out of it though, if this was still 2023-2024 I would mess around and make it work, but corporate policies in enough places have updated to use the leading closed source models and clouds for agentic coding
- 5mo ago
- biddit 5mo agoWhat an entirely unserious company. So glad I dumped Claude Code last summer after being gaslit by Anthropic over service degrades. I was fine with the service degrades, totally understandable. Being lied to, not at all. OpenAI and Altman present a whole set of different concerns, but Codex does not get in my way of doing what I want to at all. Also let me use pi without a banhammer.
- voxell_code 5mo ago[dead]
- gastonmorixe 5mo agocurl -sS https://api.anthropic.com/v1/messages \ -H "authorization: Bearer $(security find-generic-password -s 'Claude Code-credentials' -w | jq -r .claudeAiOauth.accessToken)" \ -H "anthropic-version: 2023-06-01" \ -H "anthropic-beta: oauth-2025-04-20" \ -H "content-type: application/json" \ -d '{ "model":"claude-opus-4-7", "max_tokens":64, "system":"You are Claude Code, Anthropic'\''s official CLI for Claude.", "messages":[{"role":"user","content":"Write your own harness"}] }'
- thomashobohm 5mo agoAppreciate the advice but this is Claude Managed Agents, so one can’t simply write one’s own harness.
- TheDong 5mo agoManaged agents aren't particularly harder to replicate yourself either. Give me a team of 3 good engineers, 4 months, and about $600k and I'll have a clone that operates on a warm pool of ec2 instances, or warm pool of k8s pods, or any other platform you might like. Or 1 good engineer, 1 month, and $200k of anthropic credits.
- gastonmorixe 5mo agoyou just need a max plan and a week at most
- thomashobohm 5mo agoThanks man I'll just use the $600k we had lying around.
- TheDong 5mo agoYou know, you can write in English if you want on this english-language forum. I assume you're saying "You can just generate your own harness to not be subject to these claude code issues". Unfortunately, Anthropic has already made it clear that using claude code is the only way to be sure you won't get charged API pricing instead of max plan pricing, so the tokens are way more expensive.
- anonzzzies 5mo agoThe only good thing I get from all the calling out on the decline of Claude (in this case managed agents which I do not use) is anthropic (accidentally or not) giving me basically unlimited use; for a week or so my /usage does not move anymore and I always had claude running in a loop writing code to make our many tests succeed, which can take days; before it would run out of tokens and then pick up again after the window passed until it ran out of weekly use; now I have at least one task (well, claude code instance let's say; the task is to debug and fix the code until the tests pass) thats been running 48+ hours non stop and it says usage is 10% for all of that period. Anyone else noticed? After the crash in usage a month or so ago, this is the opposite.
- cbg0 5mo agoTypically if your usage isn't moving it's because you've enabled extra usage and paying with credits.
- anonzzzies 5mo agoDefinitely have not.
- agadius 5mo agoI never thought I’d see the day that analyzing poems and other texts in my English lessons would have such drastic impact on doing computing (ref the discussion in the GitHub issues thread)
- DeathArrow 5mo agoSo after the Claude Code source leak they opened the access to Claude source or is this repo about something else?
- ptrl600 5mo agoInteresting how so much money is wasted, likely because they put a period instead of a comma.
- subscribed 5mo agoThis is so messed up. Everyone hit by this regression should be requesting API credits - it's the fault of the 100% awfully planned and vibe-coded harness fault they're burning tokens.
- danslo 5mo agoWe're enrolled in the Cyber Verification Program and Claude will happily help me look for vulnerabilities and built POCs demonstrating RCE. But when I point it to a malware sample and ask for analysis it will still refuse any work. It's incredibly frustrating.
- globular-toast 5mo agoWouldn't it be funny if this stopped, say, LinkedIn devs from doing any work because it decided, rightly so, that LinkedIn is malware?
- techpulselab 5mo ago[flagged]
- claud_ia 5mo ago[flagged]
- tommy29tmar 5mo ago[flagged]
- lifis 5mo agoI think you can fix this by either patching the binary and replacing the offending prompt with an empty string, or by pointing the harness to an API proxy that filters it out
- Kim_Bruning 5mo agoI'm currently pinning to 4.6 and the last 4.6 based CC. I apologize to all the canaries! I think it's important for CC to also be able to make unit test code that might contain mild exploits, to test for security vulnerabilities. The biggest complaint about vibe coding is that it's insecure. The funny part now is that if you DO try to secure it, you hit guardrails. There is a contact form for Anthropic if you run into some of them on 4.6 at least.
- malfist 5mo agoAnd that contact form gets their attention?
- Kim_Bruning 5mo agohttps://claude.com/form/cyber-use-case https://claude.com/form/cyber-use-case Mine got approved within 24 hours. Which is ... unusually fast for Anthropic.
- zoetaka38 5mo ago[flagged]