5 ms·
Codex on AWS bedrock bug causing 10x charges
- deleted 1mo ago[deleted]
- TheP1000 1mo agoOur codex on AWS Bedrock read / write cache ratio was less than 5%. Cache writes are very expensive and they were never being used. This results in codex on Bedrock causing ~10x what it should due to no caching and massive writes. The workaround in issue resolved for me: web_search = "disabled"
- otterley 1mo agoIf you’ve got a workaround, I’d suggest updating the issue description to have it up top there so similarly impacted users can spot it quickly and benefit.
- yablak 1mo agoWay to bury the lede..
- chrisweekly 1mo ago"causing" -> "costing", right?
- vorticalbox 1mo agoIn this case yeah. If it’s not reading the cache then it has to compute all the context window again and not just the newest tokens.
- edoceo 1mo agoLoaded question: would an openrouter or similar solution caught this before the $BigProblem showed up?
- amluto 1mo agoWow, that whole thread is borderline incoherent, presumably generated by an AI without adequate oversight. Here are the docs: https://developers.openai.com/api/docs/guides/prompt-caching#prompt-caching-for-gpt-56-and-later-models https://developers.openai.com/api/docs/guides/prompt-caching... The thread has little explanation as to what weird thing they’re doing to Codex that is making the default work poorly, and it kind of seems like it’s getting confused about whether it wants to set the caching mode or the breakpoint or both. In any case, I find the behavior change interesting. It sounds to be like 5.5 and below may have been using a conventional attention scheme where a cached KV sequence can be easily used to restore a prefix of itself, but perhaps 5.6 is using linear attention or LSTM or another recurrent scheme where you cannot rewind the model state by just truncating it.
- xiphias2 1mo agoIt’s really cool that we have this proof that US companies are half year behind Chinese models in architecture.
- geysersam 1mo agoWhat is the proof?
- cma 1mo agoNemotron was using hybrid with recurrence via mamba layers since around April 2025.
- firecall 1mo ago[dead]
- weird-eye-issue 1mo agoThen why are they (US frontier models) still so far ahead whenever I test them against the latest Chinese models? No bias here, I'd love them to be better for my own personal gain, but I haven't seen it
- hk1337 1mo agoI wonder if it's related to Codex wearing out SSDs.
- vee-kay 1mo ago[dead]
- spacedoutman 1mo agoSomething is wrong with the codex app too, burning usage like crazy lately.
- zuzululu 1mo agoindeed it has anybody know whats going on at openai ??
- dgellow 1mo agoMaybe preparing for their IPO?
- ac29 1mo agoThere haven't been any free resets in the past week, there were 4 in the first half of the month
- deleted 1mo ago[deleted]
- CSMastermind 1mo agoYeah regardless of comments by the team to the contrary (https://x.com/thsottiaux/status/2090675027670978569 https://x.com/thsottiaux/status/2090675027670978569) I have observed this in the cdoex app. My conspiratorial mind thinks they're doing this deliberately and using the resets to mask things so people can't tell their limits are reduced. The $200 / month plan covers about 2 days of usage for me right now.
- hahuhs 1mo ago[dead]
- moralestapia 1mo agoFunny how it is always more charges but never less or no charges. "Random" accidents that always go against you, too biased to be random. But don't notice that too much, you might start to see patterns here and there that you're not allowed to, might get you banned from places, etc.
- varjag 1mo agoThis can be a reporting bias. Noone opens an issue when they were billed too low.
- andrewchambers 1mo agoI doubt anyone announces when they have under billed. OpenAI has also done many low price deals and quota resets.
- evalystai 1mo agoUsually it's user's incentive to control over-billing and company's one to make sure there's no under-billing :)
- moralestapia 1mo agoBut you have eyes, right? And a brain, and live through your own experience and can reason about it, right? When was the last time someone undercharged you or didn't charge you at all by mistake? Is this a common occurrence? What's the proportion of overcharges vs. undercharges you have observed in your life?
- deleted 1mo ago[deleted]
- catlifeonmars 1mo agoApplying Occam’s razor, which do you think is more likely: 1. OpenAI intentionally adds random overcharges. 2. OpenAI deprioritizes fixing actual bugs that cause occasional overcharges because doing so won’t affect their bottom line.
- ryanjshaw 1mo agoThe fact that companies with access to SOTA non-public AI keep having these kinds of dumb bugs in something as simple as a chat app helps validate my disbelief of people claiming AI is ready to replace all software development.
- echelon 1mo agoAI is ready to replace both jobs and companies. The companies you see struggling are ripe for disruption.
- palmotea 1mo ago> The fact that companies with access to SOTA non-public AI keep having these kinds of dumb bugs in something as simple as a chat app helps validate my disbelief of people claiming AI is ready to replace all software development. Except leadership doesn't care about dumb bugs. Never has, never will (until it's too late). It cares about velocity and cost. They'd totally replace all software development with worse AI software development in a heartbeat.
- luciana1u 1mo ago[flagged]
- bigbuppo 1mo agoSounds like a path to profitability rather than a bug.
- bflesch 1mo agoRookie mistake - it seems like they didn't follow manufacturers' guidance when installing the 10x engineers. One needs to clearly define which metric should be 10x'd before powering them up.
- prtmnth 1mo agoCodex usage feels exorbitantly high since today. They [0] are denying it, but the number of anecdotal users who decided to raise this as an issue (as a result it's trending on X) says otherwise. [0] https://x.com/thsottiaux/status/2090675027670978569 https://x.com/thsottiaux/status/2090675027670978569
- shidesheng 1mo ago[flagged]
- ike_sh 1mo agoPrompt edits leaking into the cache and affecting model responses is exactly the kind of billing-relevant behavior change that should be in release notes, not discovered by users.
- ushiro35 1mo ago[flagged]
- i2km 1mo agoBut but according to Mr Altman, we've like entered the singularity. Right? Who cares about a billing issue?
- cedws 1mo agoIt’s called the singularity because all your money vanishes into a black hole.
- cmiles8 1mo agoIt would be ironic if this bug exists because it was vibe coded.
- GuestFAUniverse 1mo agoI love the fact that devs are still complaining that invoices are able to grow from $300 to $1000+ How can anyone use a platform where this is even an issue? Just because AWS is a failure in this regard, doesn't mean there aren't alternatives with fixed prices or others with easily settable limits.
- spwa4 1mo agoYou know, installing unsloth studio lets you use codex against a local qwen 3.8 instance, which does ~10 tok/sec without GPU on a modern machine, and 100+ tok/sec on a 5090, and is incredibly good.
- lostmsu 1mo agoI got a usage reset today. Possibly related.