9 ms·
Too late, personally after how bad 4.6 was the past week I was pushed to codex, which seems to mostly work at the same level from day to day. Just last night I
by buildbot 6mo ago
Too late, personally after how bad 4.6 was the past week I was pushed to codex, which seems to mostly work at the same level from day to day. Just last night I was trying to get 4.6 to lookup how to do some simple tensor parallel work, and the agent used 0 web fetches and just hallucinated 17K very wrong tokens. Then the main agent decided to pretend to implement tp, and just copied the entire model to each node...
- alvis 6mo agoI don't have much quality drop from 4.6. But I also notice that I use codex more often these days than claude code
- buildbot 6mo agoIt's been shockingly bad for me - for another example when asked to make a new python script building off an existing one; for some cursed reason the model choose to .read() the py files, use 100 of lines of regex to try to patch the changes in, and exec'd everything at the end...
- kivle 6mo agoHate that about Claude Code. I have been adding permissions for it to do everything that makes sense to add when it comes to editing files, but way too often it will generate 20-30 line bash snippets using sed to do the edits instead, and then the whole permission system breaks down. It means I have to babysit it all the time to make sure no random permission prompts pop up.
- fluidcruft 6mo agoI generally think codex is doing well until I come in with my Opus sweep to clean it up. Claude just codes closer to the way my brain works. codex is great at finding numerical stability issues though and increasingly I like that it waits for an explicit push to start working. But talking to Claude Code the way I learned to talk to codex seems to work also so I think a lot of it is just learning curve (for me).
- cmrdporcupine 6mo agoYep, I'll wait for the GPT answer to this. If we're lucky OpenAI will release a new GPT 5.5 or whatever model in the next few days, just like the last round. I have been getting better results out of codex on and off for months. It's more "careful" and systematic in its thinking. It makes less "excuses" and leaves less race conditions and slop around. And the actual codex CLI tool is better written, less buggy and faster. And I can use the membership in things like opencode etc without drama. For March I decided to give Claude Code / Opus a chance again. But there's just too much variance there. And then they started to play games with limits, and then OpenAI rolled out a $100 plan to compete with Anthropic's. I'm glad to see the competition but I think Anthropic has pissed in the well too much. I do think they sent me something about a free month and maybe I will use that to try this model out though.
- davely 6mo agoI’ve been on the Claude Code train for a while but decided to try Codex last week after they announced the $100 USD Pro plan. I’ve been pretty happy with it! One thing I immediately like more than Claude is that Codex seems much more transparent about what it’s thinking and what it wants to do next. I find it much easier to interrupt or jump in the middle if things are going to wrong direction. Claude Code has been slowly turning into this mysterious black box, wiping out terminal context any time it compacts a conversation (which I think is their hacky way of dealing with terminal flickering issues — which is still happening, 14 months later), going out of the way to hide thought output, and then of course the whole performance issues thing. Excited to try 4.7 out, but man, Codex (as a harness at least) is a stark contrast to Claude Code.
- cmrdporcupine 6mo agoDo this -- take your coworker's PRs that they've clearly written in Claude Code, and have Codex/GPT 5.4 review them. Or have Codex review your own Claude Code work. It then becomes clear just how "sloppy" CC is. I wouldn't mind having Opus around in my back pocket to yeet out whole net new greenfield features. But I can't trust it to produce well-engineered things to my standards. Not that anybody should trust an LLM to that level, but there's matters of degree here.
- muzani 6mo agoFor me, making it high effort just fixed all the quality problems, and even cut down on token use somehow
- vunderba 6mo agoThis. They kind of snuck this into the release notes: switching the default effort level to Medium. High is significantly slower, but that’s somewhat mitigated by the fact that you don’t have to constantly act like a helicopter parent for it.
- muzani 6mo agoYup, they recommend a minimum of high for coding now, and cranked the default up to extra high.
- aurareturn 6mo agoFunny because many people here were so confident that OpenAI is going to collapse because of how much compute they pre-ordered. But now it seems like it's a major strategic advantage. They're 2x'ing usage limits on Codex plans to steal CC customers and it seems to be working. I'm seeing a lot of goodwill for Codex and a ton of bad PR for CC. It seems like 90% of Claude's recent problems are strictly lack of compute related.
- energy123 6mo agoIs that 2x still going on I thought that ended in early April
- aurareturn 6mo agoThey did it again to "celebrate" the release of the $100 plan.
- indigodaddy 6mo agoOn plus?
- lawgimenez 6mo agoIt’s for Pro users only, I think the 2x is up to May 31.
- arcanemachiner 6mo agoDifferent plan. The old 2x has been discontinued, and the bonus is now (temporarily) available for the new $100 plan users in an effort, presumably, to entice them away from Anthropic.
- wahnfrieden 6mo agoFor the $200 users, it never ended.
- llm_nerd 6mo ago
- geooff_ 6mo agoI've noticed the same over the last two weeks. Some days Claude will just entirely lose its marbles. I pay for Claude and Codex so I just end up needing to use codex those days and the difference is night and day.
- frank-romita 6mo agoThat's wild that you think 4.6 is bad..... Each model has its strengths and weaknesses I find that Codex is good for architectural design and Claude Is actually better the engineering and building
- OtomotO 6mo agoSame for me. I cancelled my subscription and will be moving to Codex for the time being. Tokens are way too opaque and Claude was way smarter for my work a couple of months ago.
- cube2222 6mo agoI've been using it with `/effort max` all the time, and it's been working better than ever. I think here's part of the problem, it's hard to measure this, and you also don't know in which AB test cohorts you may currently be and how they are affecting results.
- siegers 6mo agoAgree. I keep effort max on Claude and xhigh on GPT for all tasks and keep tasks as scoped units of work instead of boil the ocean type prompts. It is hard to measure but ultimately the tasks are getting completed and I'm validating so I consider it "working as expected".
- bryanlarsen 6mo agoIt works better, until you run out of tokens. Running out of tokens is something that used to never happen to me, but this month now regularly happens. Maybe I could avoid running out of tokens by turning off 1M tokens and max effort, but that's a cure worse than the disease IMO.
- cube2222 6mo agoI would risk a guess that people have a wrong intuition about the long-context pricing and are complaining because of that. Yeah, the per-token price stays the same, even with large context. But that still means that you're spending 4x more cache-read tokens in a 400k context conversation, on each turn, than you would be in a 100k context conversation.
- rimliu 6mo agounless you always have run it on effort max, and see that it degraded.
- queuep 6mo agoBefore opus released we also saw huge backlash with it being dumber. Perhaps they need the compute for the training
- deleted 6mo ago[deleted]
- arrakeen 6mo agoso even with a new tokenizer that can map to more tokens than before, their answer is still just "you're not managing your context well enough" "Opus 4.7 uses an updated tokenizer that [...] can map to more tokens—roughly 1.0–1.35× depending on the content type. [...] Users can control token usage in various ways: by using the effort parameter, adjusting their task budgets, or prompting the model to be more concise."
- gonzalohm 6mo agoUntil the next time they push you back to Claude. At this point, I feel like this has to be the most unstable technology ever released. Imagine if docker had stopped working every two releases
- sergiotapia 6mo agoThere is zero cost to switching ai models. Paid or open source. It's one line mostly.
- gonzalohm 6mo agoWhat about your chat history? That has some value, at least for me. But what has even more value is stable releases.
- drewnick 6mo agoI think this is more about which model you steer your coding harness to. You can also self-host a UI in front of multiple models, then you own the chat history.
- sergiotapia 6mo agofor me there is zero value there.
- srmatto 6mo agoYou can output it as a memory using a simple prompt. You could probably re-use this prompt for any product with only slight modification. Or you could prompt the product to output an import prompt that is more tuned to its requirements. e.g. https://claude.com/import-memory https://claude.com/import-memory
- simplyluke 6mo agoThis is one of the many reasons I don't think the model companies are going to win the application space in coding. There's literally zero context lost for me in switching between model providers as a cursor user at work. For personal stuff I'll use an open source harness for the same reason.
- r0fl 6mo agoSame! I thought people were exaggerating how bad Claude has gotten until it deleted several files by accident yesterday Codex isn’t as pretty in output but gets the job done much more consistently
- desugun 6mo agoI guess our conscience of OpenAI working with the Department of War has an expiry date of 6 weeks.
- adamtaylor_13 6mo agoMost people just want to use a tool that works. Not everything has to be a damn moral crusade.
- martimarkov 6mo agoYes, let take morality out of our daily lives as much as possible... That seems like a great categorical imperative and a recipe for social success
- adamtaylor_13 6mo agoThat's an incredibly uncharitable take on what I said. But that kind of proves my point. Foist your morality upon everyone else and burden them with your specific conscience; sounds like a fun time.
- some_furry 6mo agoYeah, why actually engage with moral issues when we can just defer to a status quo that happens to benefit me?
- freak42 6mo agoWhat is the charitable way to look at it then?
- adamtaylor_13 6mo agoHow about assuming the positive intent of what I actually said? Not everything has to be a moral crusade. Let me use the tool without pushing your personal moral opinions on me. The same person wringing their hands over OpenAI, buys clothing made from slave labor and wrote that comment using a device with rare earth materials gotten from slave labor. Why is OpenAI the line? Why are they allowed to "exploit people" and I'm not? Taken to its logical conclusion it's silly. And instead of engaging with that, they deflect with oH yEaH lEtS hAvE nO mOrAlS which is clearly not what I'm advocating.
- hk__2 6mo agoMeh. At $work we were on CC for one month, then switched to Codex for one month, and now will be on CC again to test. We haven’t seen any obvious difference between CC and Codex; both are sometimes very good and sometimes very stupid. You have to test for a long time, not just test one day and call it a benchmark just because you have a single example.
- siegers 6mo agoI enjoy switching back and forth and having multi-agent reviews. I'm enjoying Codex also but having options is the real win.
- onlyrealcuzzo 6mo agoI switched to Codex and found it extremely inferior for my use case. It is much faster, but faster worse code is a step in the wrong direction. You're just rapidly accumulating bugs and tech debt, rather than more slowly moving in the correct direction. I'm a big fan of Gemini in general, but at least in my experience Gemini Cli is VERY FAR behind either Codex or CC. It's both slower than CC, MUCH slower than Codex, and the output quality considerably worse than CC (probably worse than Codex and orders of magnitude slower). In my experience, Codex is extraordinarily sycophantic in coding, which is a trait that could t be more harmful. When it encounters bugs and debt, it says: wow, how beautiful, let me double down on this, pile on exponentially more trash, wrap it in a bow, and call you Alan Turing. It also does not follow directions. When you tell it how to do something, it will say, nah, I have a better faster way, I'll just ignore the user and do my thing instead. CC will stop and ask for feedback much more often. YMMV.
- enraged_camel 6mo ago>> I switched to Codex and found it extremely inferior for my use case. Yeah, 100% the case for me. I sometimes use it to do adversarial reviews on code that Opus wrote but the stuff it comes back with is total garbage more often than not. It just fabricates reasons as to why the code it's reviewing needs improvement.
- Rastonbury 6mo agoWhat is your use case? I read comments like this and it's totally opposite of my experience, I have both CC Opus 4.6 and Codex 5.4 and Codex is much more thorough and checks before it starts making changes maybe even to a fault but I accept it because getting Opus to redo work because it messes up and jumps in the first attempt is a massive waste of time, all tasks and spec are atomic and granularly spec'd, I'd say 30% of the time I regret when I decide to use Opus for 'simpler' and work
- onlyrealcuzzo 6mo agoI'm building a correct, safe, highly understandable, concurrent runtime & language. Essentially Rust/Tokio if it was substantially easier than even Go - and without a need for crates and a subset of the language to achieve near Ada-level safety. The codebase is ~100k lines of code.
- _the_inflator 6mo agoCodex really has its place in my bag. I mainly use it, rarely Claude. Codex just gets it done. Very self-correcting by design while Claude has no real base line quality for me. Claude was awesome in December, but Codex is like a corporate company to me. Maybe it looks uncool, but can execute very well. Also Web Design looks really smooth with Codex. OpenAI really impressed me and continues to impress me with Codex. OpenAI made no fuzz about it, instead let results speak. It is as if Codex has no marketing department, just its product quality - kind of like Google in its early days with every product.
- te_chris 6mo agoI try codex, but i hate 5.4's personality as a partner. It's a demon debugger though. but working closely with it, it's so smug and annoying.
- vintagedave 6mo agoSame. I stopped my Pro subscription yesterday after entering the week with 70% of my tokens used by Monday morning (on light, small weekend projects, things I had worked on in the past and barely noticed a dent in usage.) Support was... unhelpful. It's been funny watching my own attitude to Anthropic change, from being an enthusiastic Claude user to pure frustration. But even that wasn't the trigger to leave, it was the attitude Support showed. I figure, if you mess up as badly as Anthropic has, you should at least show some effort towards your customers. Instead I just got a mass of standardised replies, even after the thread replied I'd be escalated to a human. Nothing can sour you on a company more. I'm forgiving to bugs, we've all been there, but really annoyed by indifference and unhelpful form replies with corporate uselessness. So if 4.7 is here? I'd prefer they forget models and revert the harness to its January state. Even then, I've already moved to Codex as of a few days ago, and I won't be maintaining two subscriptions, it's a move. It has its own issues, it's clear, but I'm getting work done. That's more than I can say for Claude.
- brenoRibeiro706 6mo agoSame here, working with claude code has been unproductive since March; everyone on my team has complained about the decline in claude code quality, which is why we’re switching to Codex.
- suzzer99 6mo agoIt seems like the big companies they're providing Mythos to are their only concern right now.
- sethhochberg 6mo agoCorporate software in general is often chosen based on the value returned simply being "good enough" most of the time, because the actual product being purchased is good controls for security, compliance, etc. A corporate purchaser is buying hundreds to thousands of Claude seats and doesn't care very much about percieved fluctuations in the model performance from release to release, they're invested in ties into their SSO and SIEM and every other internal system and have trained their employees and there's substantial cost to switching even in a rapidly moving industry. Consumer end-users are much less loyal, by comparison.
- tiel88 6mo agoI've been raging pretty hard too. Thought either I'm getting cleverer by the day or Claude has been slipping and sliding toward the wrong side of the "smart idiot" equation pretty fast. Have caught it flat-out skipping 50% of tasks and lying about it.
- thisisit 6mo agoPersonally I find using and managing Claude sessions and limits is getting exhausting and feels similar to calorie counting. You think you are going to have an amazing low calories meal only to realize the meal is full of processed sugars and you overshot the limit within 2-3 bites. Now "you have exhausted your limit for this time. Your session limits resets in next 4 hrs".
- hootz 6mo agoYep, it just feels terrible, the usage bars give me anxiety, and I think that's in their interest as they definitely push me towards paying for higher limits. Won't do that, though.
- deepsquirrelnet 6mo agoMy tinfoil hat theory, which may not be that crazy, is that providers are sandbagging their models in the days leading up to a new release, so that the next model "feels" like a bigger improvement than it is. An important aspect of AI is that it needs to be seen as moving forward all the time. Plateaus are the death of the hype cycle, and would tether people's expectations closer to reality.
- cousinbryce 6mo agoPossibly due to moving compute from inference to training
- dluxem 6mo agoMy purely unfounded, gut reaction to Opus 4.7 being released today was "Oh, that explains the recent 4.6 performance - they were spinning up inference on 4.7." Of course, I have no information on how they manage the deployment of their models across their infra.
- zee_builds 6mo ago[dead]
- baron3dl 6mo agoI was there too, but honestly after today, 4.7 "feels" just as a bad. I was cynical, but also, kind of eager for the improvement. It's just not there. Compared to early Feb, I have to babysit EVERYTHING.
- deleted 6mo ago[deleted]
- estimator7292 6mo agoAnecdotally, codex has been burning through way more tokens for me lately. Claude seems to just sit and spin for a long time doing nothing, but at least token use is moderate. All options are starting to suck more and more
- nico 6mo agoI do feel that CC sometimes starts doing dumb tasks or asking for approval for things that usually don’t really need it. Like extra syntax checks, or some greps/text parsing basic commands
- CamperBob2 6mo agoExactly. Why do they ask permission for read-only operations?! You either run with --dangerously-skip-permissions or you come back after 30 minutes to find it waiting for permission to run grep. There's no middle ground, at least not that Claude CLI users have access to.
- varispeed 6mo agoHow do you get codex to generate any code? I describe the problem and codex runs in circles basically: codex> I see the problem clearly. Let me create a plan so that I can implement it. The plan is X, Y, Z. Do you want me to implement this? me> Yes please, looks good. Go ahead! codex> Okay. Thank you for confirming. So I am going to implement X, Y, Z now. Shall I proceeed? me> Yes, proceed. codex> Okay. Implementing. ...codex is working... you see the internal monologue running in circles codex> Here is what I am going to implement: X, Y, Z me> Yes, you said that already. Go ahead! codex> Working on it. ...codex in doing something... codex> After examining the problem more, indeed, the steps should be X, Y, Z. Do you want me to implement them? etc. Very much every sessions ends up being like this. I was unable to get any useful code apart from boilerplate JS from it since 5.4 So instead I just use ChatGPT to create a plan and then ask Opus to code, but it's a hit and miss. Almost every time the prompt seems to be routed to cheaper model that is very dumb (but says Opus 4.6 when asked). I have to start new session many times until I get a good model.
- Gracana 6mo agoDo you have to put it in a build/execute mode (separate from a planning mode) to allow it to move on? I use opencode, and that's how it works.
- skocznymroczny 6mo agoIt's just like subscription based MMORPGs that delay you as much as possible every step of the way because that's the way they can extract more money from you. If you pay for the tokens it's not in their benefit to give you the answer directly.
- johanyc 6mo agoWeird. I never had that issue when writing code.
- 0xbadcafebee 6mo agoUsually the problems that cause this kind of thing are: 1) Bad prompt/context. No matter what the model is, the input determines the output. This is a really big subject as there's a ton of things you can do to help guide it or add guardrails, structure the planning/investigation, etc. 2) Misaligned model settings. If temperature/top_p/top_k are too high, you will get more hallucination and possibly loops. If they're too low, you don't get "interesting" enough results. Same for the repeat protection settings. I'm not saying it didn't screw up, but it's not really the model's fault. Every model has the potential for this kind of behavior. It's our job to do a lot of stuff around it to make it less likely. The agent harness is also a big part of it. Some agents have very specific restrictions built in, like max number of responses or response tokens, so you can prevent it from just going off on a random tangent forever.
- sgt 6mo agoStrange. Opus 4.6 has been great for me. On Max 20x
- keeganpoppen 6mo agocodex low-key seems to be better than claude. and i say this as an 18-hour-a-day user of both (mostly claude)
- timwis 6mo agoWe've started calling it dopus at work :(
- fredericgalline 6mo ago[dead]