15 ms·
The big change here is: > Standard pricing now applies across the full 1M window for both models, with no long-context premium. Media limits expand to 600 imag
by dimitri-vs 7mo ago
The big change here is:
> Standard pricing now applies across the full 1M window for both models, with no long-context premium. Media limits expand to 600 images or PDF pages.
For Claude Code users this is huge - assuming coherence remains strong past 200k tok.
- MikeNotThePope 7mo agoIs it ever useful to have a context window that full? I try to keep usage under 40%, or about 80k tokens, to avoid what Dex Horthy calls the dumb zone in his research-plan-implement approach. Works well for me so far. No vibes allowed: https://youtu.be/rmvDxxNubIg?is=adMmmKdVxraYO2yQ https://youtu.be/rmvDxxNubIg?is=adMmmKdVxraYO2yQ
- ogig 7mo agoWhen running long autonomous tasks it is quite frequent to fill the context, even several times. You are out of the loop so it just happens if Claude goes a bit in circles, or it needs to iterate over CI reds, or the task was too complex. I'm hoping a long context > small context + 2 compacts.
- boredtofears 7mo agoAll of those things are smells imo, you should be very weary of any code output from a task that causes that much thrashing to occur. In most cases it’s better to rewind or reset and adapt your prompt to avoid the looping (which usually means a more narrowly defined scope)
- grafmax 7mo agoA person has a supervision budget. They can supervise one agent in a hands-on way or many mostly-hands-off agents. Even though theres some thrashing assistants still get farther as a team than a single micromanaged agent. At least that’s my experience.
- not_kurt_godel 7mo agoJust curious, what kind of work are you doing where agentic workflows are consistently able to make notable progress semi-autonomously in parallel? Hearing people are doing this, supposedly productively/successfully, kind of blows my mind given my near-daily in-depth LLM usage on complex codebases spanning the full stack from backend to frontend. It's rare for me to have a conversation where the LLM (usually Opus 4.6 these days) lasts 30 minutes without losing the plot. And when it does last that long, I usually become the bottleneck in terms of having to think about design/product/engineering decisions; having more agents wouldn't be helpful even if they all functioned perfectly.
- avereveard 7mo agoI've passed that bottleneck with a review task that produces engineering recommendations along six axis (encapsulation, decoupling, simplification, dedoupling, security, reduce documentation drift) and a ideation tasks that gives per component a new feature idea, an idea to improve an existing feature, an idea to expand a feature to be more useful. These two generate constant bulk work that I move into new chat where it's grouped by changeset and sent to sub agent for protecting the context window. What I'm doing mostly these days is maintaining a goal.md (project direction) and spec.md (coding and process standards, global across projects). And new macro tasks development, I've one under work that is meant to automatically build png mockup and self review.
- not_kurt_godel 7mo agoWhat are you using to orchestrate/apply changes? Claude CLI?
- avereveard 7mo agoI prefer in IDE tools because I can review changes and pull in context faster. At home I use roo code, at work kiro. Tbh as long as it has task delegation I'm happy with it.
- 7mo ago
- chrisweekly 7mo agoweary (tired) -> wary (cautious)
- saaaaaam 7mo agoWary, not weary. Wary: cautious. Weary: tired.
- dentalnanobot 7mo agoThis is really common, I think because there’s also “leery” - cautious, distrustful, suspicious.
- SequoiaHope 7mo agoYep I have an autonomous task where it has been running for 8 hours now and counting. It compacts context all the time. I’m pretty skeptical of the quality in long sessions like this so I have to run a follow on session to critically examine everything that was done. Long context will be great for this.
- lukan 7mo agoAre those long unsupervised sessions useful? In the sense, do they produce useful code or do you throw most of it away?
- brookst 7mo agoI get very useful code from long sessions. It’s all about having a framework of clear documentation, a clear multi-step plan including validation against docs and critical code reviews, acceptance criteria, and closed-loop debugging (it can launch/restsart the app, control it, and monitor logs) I am heavily involved in developing those, and then routinely let opus run overnight and have either flawless or nearly flawless product in the morning.
- MikeNotThePope 7mo agoI haven't figured out how to make use of tasks running that long yet, or maybe I just don't have a good use case for it yet. Or maybe I'm too cheap to pay for that many API calls.
- ashdksnndck 7mo agoMy change cuts across multiple systems with many tests/static analysis/AI code reviews happening in CI. The agent keeps pushing new versions and waits for results until all of them come up clean, taking several iterations.
- tudelo 7mo agoI mean if you don't have your company paying for it I wouldn't bother... We are talking sessions of 500-1000 dollars in cost.
- takwatanabe 7mo agoRight. At Opus 4.6 rates, once you're at 700k context, each tool call costs ~$1 just for cache reads alone. 100 tool calls = $100+ before you even count outputs. 'Standard pricing' is doing a lot of work here lol
- brookst 7mo agoCache reads don’t count as input tokens you pay for lol. https://www.claudecodecamp.com/p/how-prompt-caching-actually-works-in-claude-code https://www.claudecodecamp.com/p/how-prompt-caching-actually...
- dimitri-vs 7mo agoIt's kind of like having a 16 gallon gas tank in your car versus a 4 gallon tank. You don't need the bigger one the majority of the time, but the range anxiety that comes with the smaller one and annoyance when you DO need it is very real.
- scwoodal 7mo agoExcept after 4 gallons it might as well be pure oil, mucking everything up.
- steve-atx-7600 7mo agoIt seems possible, say a year or two from now that context is more like a smart human with a “small”, vs “medium” vs “large” working memory. The small fellow would be able to play some popular songs on the piano , the medium one plays in an orchestra professionally and the x-large is like Wagner composing Der Ring marathon opera. This is my current, admittedly not well informed mental model anyway. Well, at least we know we’ve got a little more time before the singularity :)
- twodave 7mo agoIt’s more like the size of the desk the AI has to put sheets of paper on as a reference while it builds a Lego set. More desk area/context size = able to see more reference material = can do more steps in one go. I’ve lately been building checklists and having the LLM complete and check off a few tasks at a time, compacting in-between. With a large enough context I could just point it at a PLAN.md and tell it to go to work.
- SkyPuncher 7mo agoYes. I've recently become a convert. For me, it's less about being able to look back -800k tokens. It's about being able to flow a conversation for a lot longer without forcing compaction. Generally, I really only need the most recent ~50k tokens, but having the old context sitting around is helpful.
- hombre_fatal 7mo agoAlso, when you hit compaction at 200k tokens, that was probably when things were just getting good. The plan was in its final stage. The context had the hard-fought nuances discovered in the final moment. Or the agent just discovered some tiny important details after a crazy 100k token deep dive or flailing death cycle. Now you have to compact and you don’t know what will survive. And the built-in UI doesn’t give you good tools like deleting old messages to free up space. I’ll appreciate the 1M token breathing room.
- roygbiv2 7mo agoI've found compactation kills the whole thing. Important debug steps completely missing and the AI loops back round thinking it's found a solution when we've already done that step.
- saaaaaam 7mo agoThat video is bizarre. Such a heavy breather.
- coldtea 7mo agoWhat a weird and inconsequential thing to focus on... He's just fucking closely miced with compression + speaking fast and anxious/excited speaking to an audience
- saaaaaam 7mo agoMaybe. But that’s what I focused on, for better or worse. I couldn’t concentrate on what he was saying because of it. Maybe bad mic placement, but the end results was like some sort of old school phone sex pest.
- indigodaddy 7mo agoMost of that is just nervousness
- maskull 7mo agoAfter running a context window up high, probably near 70% on opus 4.6 High and watching it take 20% bites out of my 5hr quota per prompt I've been experimenting with dumping context after completing a task. Seems to be working ok. I wonder if I was running into the long context premium. Would that apply to Pro subs or is just relevant to api pricing?
- ricksunny 7mo agoSince I'm yet to seriously dive into vibe coding or AI-assisted coding, does the IDE experience offer tracking a tally of the context size? (So you know when you're getting close or entering the "dumb zone")?
- stevula 7mo agoMost tools do, yes.
- quux 7mo agoOpenCode does this. Not sure about other tools
- nujabe 7mo ago> Since I'm yet to seriously dive into vibe coding or AI-assisted coding Unless you’re using a text editor as an IDE you probably have already
- MikeNotThePope 7mo agoThe 2 I know, Cursor and Claude Code, will give you a percentage used for the context window. So if you know the size of the window, you can deduce the number of tokens used.
- brookst 7mo agoClaude code also gives you a granular breakdown of what’s using context window (system prompt, tools, conversation history, etc). /context
- 8note 7mo agoCline gives you such a thing. you dont really know where the dumb zone by numbers though, only by feel.
- jfim 7mo agoIn Claude code I believe it's /context and it'll give you a graphical representation of what's taking context space
- furyofantares 7mo agoI'd been on Codex for a while and with Codex 5.2 I: 1) No longer found the dumb zone 2) No longer feared compaction Switching to Opus for stupid political reasons, I still have not had the dumb zone - but I'm back to disliking compaction events and so the smaller context window it has, has really hurt. I hope they copy OpenAI's compaction magic soon, but I am also very excited to try the longer context window.
- iknowstuff 7mo agoHmm I’ve felt the dumb zone on codex
- nomel 7mo agoFrom what I've seen, it means whatever he's doing is very statistically significant.
- mgambati 7mo ago1m context in OpenAI and Gemini is just marketing. Opus is the only model to provide real usable bug context.
- hu3 7mo agoSource? I ask because I use 500k+ context on these on a daily basis. Big refactorings guided by automated tests eat context window for breakfast.
- twodave 7mo agoI mean, try using copilot on any substantial back-end codebase and watch it eat 90+% just building a plan/checklist. Of course copilot is constrained to 120k I believe? So having 10x that will blow open up some doors that have been closed for me in my work so far. That said, 120k is pleeenty if you’re just building front-end components and have your API spec on hand already.
- bushbaba 7mo agoYes. I’ve used it for data analysis
- kaizenb 7mo agoThanks for the video. His fix for "the dumb zone" is the RPI Framework: ● RESEARCH. Don't code yet. Let the agent scan the files first. Docs lie. Code doesn't. ● PLAN. The agent writes a detailed step-by-step plan. You review and approve the plan, not just the output. Dex calls this avoiding "outsourcing your thinking." The plan is where intent gets compressed before execution starts. ● IMPLEMENT. Execute in a fresh context window. The meta-principle he calls Frequent Intentional Compaction: don't let the chat run long. Ask the agent to summarize state, open a new chat with that summary, keep the model in the smart zone.
- girvo 7mo agoThat's fascinating: that is identical to the workflow I've landed on myself.
- hedora 7mo agoIt's also identical to what Claude Code does if you put it in plan mode (bound to <tab> key), at least in my experience.
- girvo 7mo agoMy annoyance with plan mode is where it sticks the .md file, kind of hides it away which makes it annoying to clear context and start up a new phase from the PLAN file. But that might just be a skill issue on my end
- hedora 7mo agoEven worse, it just randomly blows away the plan file without asking for permission. No idea what they were thinking when they designed this feature. The plan file names are randomly generated, so it could just keep making new ones forever for free (it would take a LONG time for the disk space to matter), but instead, for long plans, I have to back the plan file up if it gets stuck. Otherwise, I say "You should take approach X to fix this bug", it drops into plan mode, says "This is a completely unrelated plan", then deletes all record of what it was doing before getting stuck.
- Barbing 7mo agoLooking at this URL, typo or YouTube flip the si tracking parameter? youtu.be/rmvDxxNubIg?is=adMmmKdVxraYO2yQ
- MikeNotThePope 7mo agoI just cut & pasted the share URL provided by YouTube. Strip out the query param if you like.
- Barbing 7mo agoOoh it’s always ?si= So this… ?is= …that’s new. Think you got A/B tested. Flipping the parameter breaks a lot of RegEx. Interesting!
- dev_l1x_be 7mo agoI never use these giant context windows. It is pointless. Agents are great at super focused work that is easy to re-do. Not sure what is the use case for giant context windows.
- hrmtst93837 7mo ago[flagged]
- wat10000 7mo agoI've used it many times for long-running investigations. When I'm deep in the weeds with a ton of disassembly listings and memory dumps and such, I don't really want to interrupt all of that with a compaction or handoff cycle and risk losing important info. It seems to remain very capable with large contexts at least in that scenario.
- alecco 7mo agoOfftopic: I find it remarkable the shortened YT url has a tracking cost of 57% extra length. We live in stupid times.
- dahart 7mo agoI care about the privacy implications, but not the length. Out of curiosity, why do you care about the URL length at all? What is the cost to you?
- alecco 7mo agoMy point is Google engineers go to the trouble of setting up a URL shortener service on one hand, but on the other hand it seems ad the business anti-privacy executives can override anything. This points out it's a dysfunctional company.
- inemesitaffia 7mo agoThe point is whatever group controls the money controls the power. Also, only the domain is shorter
- alecco 7mo agoActually, it's not just the domain: https://youtu.be/X https://youtu.be/X https://www.youtube.com/watch?v=X https://www.youtube.com/watch?v=X
- dahart 7mo agoYou’d rather have the video code and the tracking code baked into the same code just to save a couple of characters? Why? That would result in a longer code than the video code alone, you would save very few characters. It would not be nicer to look at or functionally any different, and it would obscure the fact that it’s being tracked and prevent people from being able to edit the URL to remove the tracking. I appreciate the fact that I can see that the URL has a tracking ID and that I can edit the URL and remove the tracking ID. I do not want a shorter URL if I lose that ability. What you’re complaining about and wishing for would be MUCH worse than what it currently is.
- virtualritz 7mo agoI haven't hit the "dumb zone" any more since two months. I think this talk is outdated. I'm using CC (Opus) thinking and Codex with xhigh on always. And the models have gotten really good when you let them do stuff where goals are verifiable by the model. I had Codex fix a Rust B-rep CSG classification pipeline successfully over the course of a week, unsupervised. It had a custom STEP viewer that would take screenshots and feed them back into the model so it could verify the progress resp. the triangle soup (non progress) itself. Codex did all the planning and verification, CC wrote the code. This would have not been possible six months ago at all from my experience. Maybe with a lot of handholding; but I doubt it (I tried). I mean both the problem for starters (requires a lot of spatial reasoning and connected math) and the autonomous implementation. Context compression was never an issue in the entire session, for either model.
- alexey-pelykh 7mo ago[dead]
- islewis 7mo agoThe quality with the 1M window has been very poor for me, specifically for coding tasks. It constantly forgets stuff that has happened in the existing conversation. n=1, ymmv
- deleted 7mo ago[deleted]
- robwwilliams 7mo agoYes, especially with shifts in focus of a long conversation. But given the high error rates of Opus 4.6 the last few weeks it is possibly due to other factors. Conversational and code prodding has been essential.
- hagen8 7mo agoWell, the question is what is contributing to the usage. Because as the context grows, the amount of input tokens are increasing. A model call with 800K token as input is 8 times more expensive than a model call with 100K tokens as input. Especially if we resume a conversation and caching does not hit, it would be very expensive with API pricing.
- a_e_k 7mo agoI've been using the 1M window at work through our enterprise plan as I'm beginning to adopt AI in my development workflow (via Cline). It seems to have been holding up pretty well until about 700k+. Sometimes it would continue to do okay past that, sometimes it started getting a bit dumb around there. (Note that I'm using it in more of a hands-on pair-programming mode, and not in a fully-automated vibecoding mode.)
- chatmasta 7mo agoSo a picture is worth 1,666 words?
- jFriedensreich 7mo agoyeah it totally does not remain coherent past 200k, would have been too nice.
- __MatrixMan__ 7mo agoI bet it depends how homogenous the context is. I bet it works ok near 1M in some cases, but as far as I can tell, those cases are rare.
- Bombthecat 7mo agoIf it's not coding, even with 200k context it starts to write gibberish, even with the correct information in the context. I tried to ask questions about path of exile 2. And even with web research on it gave completely wrong information... Not only outdated. Wrong I think context decay is a bigger problem then we feel like.
- reactordev 7mo agoThat’s not context decay, that’s training data ambiguity. So much misinformation, nerfs, buffs, changes that an LLM can not keep up given the training time required. Do it for a game that has been stable and it knows its stuff.
- Bombthecat 7mo agoIt didnt gave outdated, on some cases it did, and with two tries telling it to search for updated information it got it right ( shouldn't need to do that though) but it also gave wrong information about sockets ( support skills) , which never existed or never were able to be socketed together in the first place. ( Ok maybe in 0.1, but that's what web search is for ... ) If it even can't handle easy versioned information from a game. How should it handle anything related to time, dates, news, science etc?
- serial_dev 7mo agoPlease don’t pop the AI bubble, bro. Stop asking questions, bro. Believe the hype, bro.
- reactordev 7mo agoLike any human would, 75% certain with 99% confidence. That’s what you fail to realize. They aren’t “god mode machine”. They are “human-mode” machines and humans make mistakes in thinking just like you do. Some might say asking a powerful LLM for gaming tips is a waste of compute power. Others might say it gives you the knowledge of a new meta emerging. Either way, you both are going to get trained.
- 7mo ago
- alexcali 7mo ago[dead]
- j45 7mo agoThis might burn through usage faster too though.