7 ms·
Codex has always been better at following agents.md and prompts more, but I would say in the last 3 months both Claude Code got worse (freestyling like we see h
by inerte 7mo ago
Codex has always been better at following agents.md and prompts more, but I would say in the last 3 months both Claude Code got worse (freestyling like we see here) and Codex got EVEN more strict.
80% of the time I ask Claude Code a question, it kinda assumes I am asking because I disagree with something it said, then acts on a supposition. I've resorted to append things like "THIS IS JUST A QUESTION. DO NOT EDIT CODE. DO NOT RUN COMMANDS". Which is ridiculous.
Codex, on the other hand, will follow something I said pages and pages ago, and because it has a much larger context window (at least with the setup I have here at work), it's just better at following orders.
With this project I am doing, because I want to be more strict (it's a new programming language), Codex has been the perfect tool. I am mostly using Claude Code when I don't care so much about the end result, or it's a very, very small or very, very new project.
- parhamn 7mo agoI added an "Ask" button my agent UI (openade.ai) specifically because of this!
- hrimfaxi 7mo ago> Codex, on the other hand, will follow something I said pages and pages ago, and because it has a much larger context window (at least with the setup I have here at work), it's just better at following orders. Can you speak more to that setup?
- inerte 7mo agoClaude Code goes through some internal systems that other tools (Cline / Codex / and I think Cursor) do not. Also we have different models for each. I don't know in practice what happens, but I found that Codex compacts conversations way less often. It might as well be somehow less tokens are used/added, then raw context window size. Sorry if I implied we have more context than whatever others have :)
- rsanheim 7mo agoCodex does something sorta magical where it auto compacts, partially maybe, when it has the chance. I don’t know how it works, and there is little UI indication for it.
- torben-friis 7mo ago>I've resorted to append things like "THIS IS JUST A QUESTION. DO NOT EDIT CODE. DO NOT RUN COMMANDS". Which is ridiculous. Funny to read that, because for me it's not even new behavior. I have developed a tendency to add something like "(genuinely asking, do not take as a criticism)". I'm from a more confrontational culture, so I just assumed this was just corporate American tone framing criticism softly, and me compensating for it.
- mikepurvis 7mo agoI've been using chat and copilot for many months but finally gave claude code a go, and I've been interested how it does seem to have a bit more of an attitude to it. Like copilot is just endlessly patient for every little nitpick and whim you have, but I feel like Claude is constantly like "okay I'm committing and pushing now.... oh, oh wait, you're blocking me. What is it you want this time bro?"
- nineteen999 7mo ago"Don't act, just a question" works for me.
- d1sxeyes 7mo agoTry /btw
- nineteen999 7mo agoThat's not a thing in Claude ... so no.
- andyferris 7mo agoIt's new
- closewith 7mo agoIt is in Claude Code, specifically for this use case.
- darkoob12 7mo agoThis is not Claude Code. And my experience is the opposite. For me Codex is not working at all to the point that it's not better than asking the chat bot in the browser.
- thomasfromcdnjs 7mo agoA lot of people dunking but as this comment says, it is not claude code. (just opus 4.6)
- pprotas 7mo agoThis comment is right, this screenshot is not Claude Code. It’s Opencode.
- stavros 7mo agoI've added an instruction: "do not implement anything unless the user approves the plan using the exact word 'approved'". This has fixed all of this, it waits until I explicitly approve.
- AnotherGoodName 7mo agoThere’s an extension to this problem which I haven’t got past. More generally I’d like the agent to stop and ask questions when it encounters ambiguity that it can’t reasonably resolve itself. If someone can get agents doing this well it’d be a massive improvement (and also solve the above).
- stavros 7mo agoHm, with my "plan everything before writing code, plus review at the end" workflow, this hasn't been a problem. A few times when a reviewer has surfaced a concern, the agent asks me, but in 99% of cases, all ambiguity is resolved explicitly up front.
- skeeter2020 7mo agowhat gung-ho, talented-but-naive junior developer has ever done that?
- eproxus 7mo agoIn planning I sometimes add ”ask me questions as we go to iron out details and ambiguities.” Works quite well.
- vitaflo 7mo agoThis. Just asking it to ask you questions before proceeding has saved me so much time from it making assumptions I don’t want. It’s the single most important part of almost all my prompts.
- xeckr 7mo ago"NOT approved!" "The user said the exact word 'approved'. Implementing plan."
- lubujackson 7mo agoI feel like people are sleeping on Cursor, no idea why more devs don't talk about it. It has a great "Ask" mode, the debugging mode has recently gotten more powerful, and it's plan mode has started to look more like Claude Code's plans, when I test them head to head.
- hansonkd 7mo agoI love to build a plan, then cycle to another frontier model to iterate on it.
- ponyous 7mo agoIn the coworking I am in people are hitting limits on 60$ plan all the time. They are thinking about which models to use to be efficient, context to include etc… I’m on claude code $100 plan and never worry about any of that stuff and I think I am using it much more than they use cursor. Also, I prefer CC since I am terminal native.
- adwn 7mo agoTell them to use the Composer 1.5 model. It's really good, better than Sonnet, and has much higher usage limits. I use it for almost all of my daily work, don't have to worry about hitting the limit of my 60$ plan, and only occasionally switch to Opus 4.6 for planning a particularly complex task.
- bushido 7mo agoCursor implemented something a while back where it started acting like how ChatGPT does when it's in its auto mode. Essentially, choosing when it was going to use what model/reasoning effort on its own regardless of my preferences. Basically moved to dumber models while writing code in between things, producing some really bad results for me. Anecdotal, but the reason I will never talk about Cursor is because I will never use it again. I have barred the use of Cursor at my company, It just does some random stuff at times, which is more egregious than I see from Codex or Claude. ps. I know many other people who feel the same way about Cursor and other who love it. I'm just speaking for myself, though. ps2. I hope they've fixed this behavior, but they lost my trust. And they're likely never winning it back.
- cmrdporcupine 7mo agoI'm back on Claude Code this month after a month on Codex and it's a serious downgrade. Opus 4.6 is a jackass. It's got Dunning-Kruger and hallucinates all over the place. I had forgotten about the experience (as in the Gist above) of jamming on the escape key "no no no I never said to do that." But also I don't remember 4.5 being this bad. But GPT 5.3 and 5.4 is a far more precise and diligent coding experience.
- sroussey 7mo agoUse cli or extension or the app?
- AlotOfReading 7mo agoI've had some luck taming prompt introspection by spawning a critic agent that looks at the plan produced by the first agent and vetos it if the plan doesn't match the user's intentions. LLMs are much better at identifying rule violations in a bit of external text than regulating their own output. Same reason why they generate unnecessary comments no matter how many times you tell them not to.
- miohtama 7mo agoHow does one integrate critic agent to a Codex/Claude?
- bentcorner 7mo agoI just say something like "spawn an agent to review your plan" or something to that effect. "Red/green TDD" is apparently the nomenclature: https://simonwillison.net/guides/agentic-engineering-patterns/red-green-tdd/ https://simonwillison.net/guides/agentic-engineering-pattern... I've also found it to be better to ask the LLM to come up with several ideas and then spawn additional agents to evaluate each approach individually. I think the general problem is that context cuts both ways, and the LLM has no idea what is "important". It's easier to make sure your context doesn't contain pink elephants than it is to tell it to forget about the pink elephants.
- collinmanderson 7mo ago> "Red/green TDD" is apparently the nomenclature From your link: > what "red/green" means: the red phase watches the tests fail, then the green phase confirms that they now pass. > Every good model understands "red/green TDD" as a shorthand for the much longer "use test driven development, write the tests first, confirm that the tests fail before you implement the change that gets them to pass".
- AlotOfReading 7mo agoYou can just say spawn an agent as the sibling says. I didn't find that reliable enough, so I have a slightly more complicated setup. First agent has no permissions except spawning agents and reading from a single directory. It spawns the planner to generate the plan, then either feeds it to the critic and either spawns executors or re-runs the planner with critic feedback. The planner can read and write. The critic agent can only read the input and outputs accept/reject with reason. This is still sometimes flaky because of the infrastructure around it and ideally you'd replace the first agent with real code, but it's an improvement despite the cost.
- clarus 7mo agoThe solution for this might be to add a ME.md in addition to AGENT.md so that it can learn and write down our character, to know if a question is implicitly a command for example.
- casey2 7mo agoFor the last 12 months labs have been 1. check-pointing 2. train til model collapse 3. revert to the checkpoint from 3 months ago 4. People have gotten used to the shitty new model Antropic said they "don't do any programming by hand" the last 2 years. Antropic's API has 2 nines
- chrysoprace 7mo agoMaybe I should give Codex a go, because sometimes I just want to ask a question (Claude) and not have it scan my entire working directory and chew up 55k tokens.
- thomaslord 7mo agoThis is extra rough because Codex defaults to letting the model be MUCH more autonomous than Claude Code. The first time I tried it out, it ended up running a test suite without permission which wiped out some data I was using for local testing during development. I still haven't been able to find a straight answer on how to get Codex to prompt for everything like Claude Code does - asking Codex gets me answers that don't actually work.
- 0xbadcafebee 7mo agoThis is mostly dependent on the agent because the agent sets the system prompt. All coding agents include in the system prompt the instruction to write code, so the model will, unless you tell it not to. But to what extent they do this depends on that specific agent's system prompt, your initial prompt, the conversation context, agent files, etc. If you were just chatting with the same model (not in an agent), it doesn't write code by default, because it's not in the system prompt.
- hun3 7mo agoDoes appending "/genq" work? Or use the /btw command to ask only questions
- wartywhoa23 7mo agoI guess appending the actual correct handwritten brainthought code is the solution here.
- deleted 7mo ago[deleted]
- hun3 7mo agoWell, tell that to OP, not me.
- niobe 7mo agoBut that's one of the first things you fix in your CLAUDE.md: - "Only do what is asked." - "Understand when being asked for information versus being asked to execute a task."
- bdangubic 7mo agoThis - per extensive experiments - works about as well as when I tell my wife to calm down
- smackeyacky 7mo agoAsking might work better than telling
- bdangubic 7mo agoHow do you do that???? Say the words but in the form of a question? I feel like that will go a lot worse than just telling (but nicely). I have a daughter too so I am genuinely willing to try anything
- smackeyacky 7mo agoPlease and thank you and make sure you’re addressing the behaviour and not the person.
- user3939382 7mo agoClaude Code is perfectly happy to toggle between chat and work but if you’re simply clear about which you want. Capital letters aren’t necessary.
- onion2k 7mo agoCodex, on the other hand, will follow something I said pages and pages ago, and because it has a much larger context window (at least with the setup I have here at work), it's just better at following orders. This is important, but as a warning. At least in theory your agent will follow everything that it has in context, but LLMs rely on 'context compacting' when things get close to the limit. This means an LLM can and will drop your explicit instructions not to do things, and then happily do them because they're not in the context any more. You need to repeat important instructions.
- tomtomistaken 7mo agoFor Claude writing "let's discuss" at the end of the prompt seems to do it
- 112233 7mo agoI tried using codex, and it is great (meaning - boring) when it works. My problem is it does not work. Let me explain codex> Next I can make X if you agree. me> ok codex> I will make X now me> Please go on codex> Great, I am starting to work on X now me> sure, please do codex> working on X, will report on completion me> yo good? please do X! ... and so on. Sometimes one round, sometimes four, plus it stops after every few lines to "report progress" and needs another nudge or five. :(
- dwedge 7mo agoFirst time I used Claude I asked it to look at the current repo and just tell me where the database connection string was defined. It added 100 lines of code. I asked it to undo that and it deleted 1000 lines and 2 files
- exceptione 7mo agoWould `git reset --hard` have worked to in your case? I guess you want to have each babystep in a git commit, in the end you could do a `git rebase -i` if needed.
- aidos 7mo agoOne annoying thing about that flow is that when you change the world outside the model it breaks its assumptions and it loses its way faster (in my experience).
- dwedge 7mo agoWithout git I would have been screwed. AI doesn't commit anything, I do when I'm satisfied
- bagacrap 7mo agoAh, so you have not yet been forced to tell it DO NOT AMEND THE LAST COMMIT
- 7mo ago
- dr_dshiv 7mo ago“Don’t code yet” is a longstanding part of the rapport
- xboxnolifes 7mo agoI just start my prompts with "conceptually, ..." and thats usually enough to stop claude from going down the coding path.
- tempestn 7mo agoWhat about adding something like, "When asked a question, just answer it without assuming any implied criticism or instructions. Questions are just questions." to claude.md?
- lwhi 7mo agoI've found codex will find another way to do what it wants, if I deny it access to a command request.
- iainmck29 7mo agoI find this thread surprising honestly. Claude Code is my daily driver and I consider myself a real power user. If you have your commands/agents/skills set up correctly you should never be running into these issues
- sumeno 7mo agoAhh, "you're holding it wrong" The classics never go out of style
- jasonlotito 7mo agoI mean, in this case, we aren't even holding Claude Code. So weird to complain about something that isn't even in the original post.
- malfist 7mo agoYour experience is not universal.
- deleted 7mo ago[deleted]
- bartread 7mo agoAre you finding this happens even in “Plan Mode”?
- duxup 7mo agoYour experience with Claude is surprising to me. At least for me when using Claude in VSCode (extension) there’s clearly defined “plan mode” and “ask before edits” and “edit automatically”. I’ve never had it disregard those modes.