6 ms·
Thanks for the feedback. To make it actionable, would you mind running /bug the next time you see it and posting the feedback id here? That way we can debug and
by bcherny 6mo ago
Thanks for the feedback. To make it actionable, would you mind running /bug the next time you see it and posting the feedback id here? That way we can debug and see if there's an issue, or if it's within variance.
- freedomben 6mo agoHow much of the code/context gets attached in the /bug report?
- bcherny 6mo agoWhen you submit a /bug we get a way to see the contents of the conversation. We don't see anything else in your codebase.
- murkt 6mo agoWas there a change in Claude Code system prompt at that time that nudges Claude into simplistic thinking? Here is a gist that tries to patch the system prompt to make Claude behave better https://gist.github.com/roman01la/483d1db15043018096ac3babf5688881 https://gist.github.com/roman01la/483d1db15043018096ac3babf5... I haven’t personally tried it yet. I do certainly battle Claude quite a lot with “no I don’t want quick-n-easy wrong solution just because it’s two lines of code, I want best solution in the long run”. If the system prompt indeed prefers laziness in 5:1 ratio, that explains a lot. I will submit /bug in a few next conversations, when it occurs next.
- enigmo 6mo ago[dead]
- dev_l1x_be 6mo agoHoly sweet LLM, this gist is crazy. Why did they do this to themselves? I am going to try this at home, it might actually fix Claude.
- murkt 6mo agoRemember Sonnet 3.5 and 3.7? They were happy to throw abstraction on top of abstraction on top of abstraction. Still a lot of people have “do not over-engineer, do not design for the future” and similar stuff in their CLAUDE.md files. So I think the system prompt just pushes it way too hard to “simple” direction. At least for some people. I was doing a small change in one of my projects today, and I was quite happy with “keep it stupid and hacky” approach there. And in the other project I am like “NO! WORK A LOT! DO YOUR BEST! BE HAPPY TO WORK HARD!” So it depends.
- pbowyer 6mo agoLet us know if it does, because we all want it to work :)
- Avamander 6mo agoThat Gist does explain quite a few flaws Claude has. I wonder if MEMORY.md is sufficient to counteract the prompt without patching.
- liamsfr 6mo agoAnd if memory.md can’t and you need something quick and dirty for flat memory management, I wrote a plugin just for this. https://github.com/NominexHQ/pmm-plugin https://github.com/NominexHQ/pmm-plugin
- withinboredom 6mo agoIs there not a setting to change the system prompt itself? I vaguely remember seeing it in the docs.
- matheusmoreira 6mo agoThere is!! https://code.claude.com/docs/en/cli-reference#system-prompt-flags https://code.claude.com/docs/en/cli-reference#system-prompt-... --append-system-prompt --append-system-prompt-file --system-prompt --system-prompt-file Can this script be made to work without patching the executable?
- withinboredom 6mo agoMight be worth extracting the system prompt and then patching it. TBH, that's what I was expecting when I saw the gist.
- matheusmoreira 6mo agoThis might be more complex than I imagined. It seems Claude Code dynamically customizes the system prompt. They also update the system prompt with every version so outright replacing it will cause us to miss out on updates. Patching is probably the best solution. https://github.com/Piebald-AI/claude-code-system-prompts https://github.com/Piebald-AI/claude-code-system-prompts https://github.com/Piebald-AI/tweakcc https://github.com/Piebald-AI/tweakcc
- withinboredom 6mo agoInteresting. So literally triggering any of these changes probably invalidates the cache as well…
- andersa 6mo agoI didn't know we could change the base system prompt of Claude Code. Just tried, and indeed it works. This changes everything! Thank you for posting this!
- naasking 6mo agoVery interesting. I run Claude Code in VS Code, and unfortunately there doesn't seem to be an equivalent to "cli.js", it's all bundled into the "claude.exe" I've found under the VS code extensions folder (confirmed via hex editor that the prompts are in there). Edit: tried patching with revised strings of equivalent length informed by this gist, now we'll see how it goes!
- matheusmoreira 6mo agoI adapted these patches into settings for the tweakcc tool. https://github.com/Piebald-AI/tweakcc https://github.com/Piebald-AI/tweakcc Pushed it to my dotfiles repository: https://github.com/matheusmoreira/.files/tree/master/~/.tweakcc/system-prompts https://github.com/matheusmoreira/.files/tree/master/~/.twea... The tweaks can be applied with npx tweakcc --apply
- andoando 6mo agoIsnt the codebase in the context window?
- frog437 6mo agodepending on how large your codebase is, hopefully not. At this point use something like the IX plugin to ingest codebase and track context, rather than from the LLM itself.
- frog437 6mo agoThis is crazy.. tokensSaved = naiveTokens - actualTokens - naiveTokens = 19.4M — what ix estimates it would have cost to answer your queries without graph intelligence (i.e., dumping full files/directories into context) - actualTokens = 4.7M — what ix's targeted, graph-aware responses actually used - tokensSaved = 14.7M — the difference
- andoando 6mo agoI mean whatever part of the code that is read by the AI has to be in the content window at some point or another nSprewd throughout your sessions Id think even with a huge codebase, 90% of it is going to be there
- koverstreet 6mo agoI'll have a look. The CoT switch you mentioned will help, I'll take a look at that too, but my suspicion is that this isn't a CoT issue - it's a model preference issue. Comparing Opus vs. Qwen 27b on similar problems, Opus is sharper and more effective at implementation - but will flat out ignore issues and insist "everything is fine" that Qwen is able to spot and demonstrate solid understanding of. Opus understands the issues perfectly well, it just avoids them. This correlates with what I've observed about the underlying personalities (and you guys put out a paper the other day that shows you guys are starting to understand it in these terms - functionally modeling feelings in models). On the whole Opus is very stable personality wise and an effective thinker, I want to complement you guys on that, and it definitely contrasts with behaviors I've seen from OpenAI. But when I do see Opus miss things that it should get, it seems to be a combination of avoidant tendencies and too much of a push to "just get it done and move into the next task" from RHLF.
- jchanimal 6mo agoOne of the thing is we’ve seen at vibes.diy is that if you have a list of jobs and you have agents with specialized profiles and ask them to pick the best job for themselves that can change some of the behavior you described at the end of your post for the better.
- necrotic_comp 6mo agoOpus definitely pushes me to ignore problems. I've had to tell it multiple times to be thorough, and we tend to go back and forth a few times every time that happens. :)
- pimeys 6mo ago"I see the tests failing, but none of our changes caused this breakage so I will push my changes and ask the user to inform their team on failing tests."
- JamesSwift 6mo agoa9284923-141a-434a-bfbb-52de7329861d d48d5a68-82cd-4988-b95c-c8c034003cd0 5c236e02-16ea-42b1-b935-3a6a768e3655 22e09356-08ce-4b2c-a8fd-596d818b1e8a 4cb894f7-c3ed-4b8d-86c6-0242200ea333 Amusingly (not really), this is me trying to get sessions to resume to then get feedback ids and it being an absolute chore to get it to give me the commands to resume these conversations but it keeps messing things up: cf764035-0a1d-4c3f-811d-d70e5b1feeef
- bcherny 6mo agoThanks for the feedback IDs — read all 5 transcripts. On the model behavior: your sessions were sending effort=high on every request (confirmed in telemetry), so this isn't the effort default. The data points at adaptive thinking under-allocating reasoning on certain turns — the specific turns where it fabricated (stripe API version, git SHA suffix, apt package list) had zero reasoning emitted, while the turns with deep reasoning were correct. we're investigating with the model team. interim workaround: CLAUDE_CODE_DISABLE_ADAPTIVE_THINKING=1 forces a fixed reasoning budget instead of letting the model decide per-turn.
- diavelguru 6mo agoLove this. Responding to users. Detail info investigating. Action being taken (at least it seems so).
- onoesworkacct 6mo agoThis kind of thing is harder for regular end-users to understand following the change removing reasoning details.
- matheusmoreira 6mo agoI just asked Claude to plan out and implement syntactic improvements for my static site generator. I used plan mode with Opus 4.6 max effort. After over half an hour of thinking, it produced a very ad-hoc implementation with needless limitations instead of properly refactoring and rearchitecting things. I had to specifically prompt it in order to get it to do better. This executed at around 3 AM UTC, as far away from peak hours as it gets. b9cd0319-0cc7-4548-bd8a-3219ede3393a > You're right to push back. Let me be honest about both questions. > The @() implementation is ad-hoc > The current implementation manually emits synthetic tokens — tag, start-attributes, attribute, end-attributes, text, end-interpolation — in sequence. > This works, but it duplicates what the child lexer already does for #[...], creating two divergent code paths for the same conceptual operation (inline element emission). It also means @() link text can't contain nested inline elements, while #[a(...) text with #[em emphasis]] can. I just feel like I can't trust it anymore.
- koverstreet 6mo agoThat's pretty much been my day - today was genuinely bad, and I've been putting up with a lot of this lately. Now on Qwen3.5-27b, and it may not be quite as sharp as Opus was two months ago, but we're getting work done again.
- matheusmoreira 6mo agoLiterally two weeks ago it was outputting excellent results while working with me on my programming language. I reviewed every line and tried to understand everything it did. It was good. I slowly started trusting it. Now I don't want to let it touch my project again. It's extremely depressing because this is my hobby and I was having such a blast coding with Claude. I even started trying to use it to pivot to professional work. Now I'm not sure anymore. People who depend on this to make a living must be very angry indeed.
- jacquesm 6mo agoI can see how that works: this is like building a dependency, a habit if you wish. I think the tighter you couple your workflow to these tools the more dependent you will become and the greater the let-down if and when they fail. And they will always fail, it just depends on how long you work with them and how complex the stuff is you are doing, sooner or later you will run into the limitations of the tooling. One way out of this is to always keep yourself in the loop. Never let the work product of the AI outpace your level of understanding because the moment you let that happen you're like one of those cartoon characters walking on air while gravity hasn't reasserted itself just yet.