12 ms·
1M context is now generally available for Opus 4.6 and Sonnet 4.6
- deleted 7mo ago[deleted]
- dimitri-vs 7mo agoThe big change here is: > Standard pricing now applies across the full 1M window for both models, with no long-context premium. Media limits expand to 600 images or PDF pages. For Claude Code users this is huge - assuming coherence remains strong past 200k tok.
- MikeNotThePope 7mo agoIs it ever useful to have a context window that full? I try to keep usage under 40%, or about 80k tokens, to avoid what Dex Horthy calls the dumb zone in his research-plan-implement approach. Works well for me so far. No vibes allowed: https://youtu.be/rmvDxxNubIg?is=adMmmKdVxraYO2yQ https://youtu.be/rmvDxxNubIg?is=adMmmKdVxraYO2yQ
- ogig 7mo agoWhen running long autonomous tasks it is quite frequent to fill the context, even several times. You are out of the loop so it just happens if Claude goes a bit in circles, or it needs to iterate over CI reds, or the task was too complex. I'm hoping a long context > small context + 2 compacts.
- boredtofears 7mo agoAll of those things are smells imo, you should be very weary of any code output from a task that causes that much thrashing to occur. In most cases it’s better to rewind or reset and adapt your prompt to avoid the looping (which usually means a more narrowly defined scope)
- grafmax 7mo agoA person has a supervision budget. They can supervise one agent in a hands-on way or many mostly-hands-off agents. Even though theres some thrashing assistants still get farther as a team than a single micromanaged agent. At least that’s my experience.
- not_kurt_godel 7mo agoJust curious, what kind of work are you doing where agentic workflows are consistently able to make notable progress semi-autonomously in parallel? Hearing people are doing this, supposedly productively/successfully, kind of blows my mind given my near-daily in-depth LLM usage on complex codebases spanning the full stack from backend to frontend. It's rare for me to have a conversation where the LLM (usually Opus 4.6 these days) lasts 30 minutes without losing the plot. And when it does last that long, I usually become the bottleneck in terms of having to think about design/product/engineering decisions; having more agents wouldn't be helpful even if they all functioned perfectly.
- avereveard 7mo agoI've passed that bottleneck with a review task that produces engineering recommendations along six axis (encapsulation, decoupling, simplification, dedoupling, security, reduce documentation drift) and a ideation tasks that gives per component a new feature idea, an idea to improve an existing feature, an idea to expand a feature to be more useful. These two generate constant bulk work that I move into new chat where it's grouped by changeset and sent to sub agent for protecting the context window. What I'm doing mostly these days is maintaining a goal.md (project direction) and spec.md (coding and process standards, global across projects). And new macro tasks development, I've one under work that is meant to automatically build png mockup and self review.
- not_kurt_godel 7mo agoWhat are you using to orchestrate/apply changes? Claude CLI?
- SequoiaHope 7mo agoYep I have an autonomous task where it has been running for 8 hours now and counting. It compacts context all the time. I’m pretty skeptical of the quality in long sessions like this so I have to run a follow on session to critically examine everything that was done. Long context will be great for this.
- lukan 7mo agoAre those long unsupervised sessions useful? In the sense, do they produce useful code or do you throw most of it away?
- brookst 7mo agoI get very useful code from long sessions. It’s all about having a framework of clear documentation, a clear multi-step plan including validation against docs and critical code reviews, acceptance criteria, and closed-loop debugging (it can launch/restsart the app, control it, and monitor logs) I am heavily involved in developing those, and then routinely let opus run overnight and have either flawless or nearly flawless product in the morning.
- MikeNotThePope 7mo agoI haven't figured out how to make use of tasks running that long yet, or maybe I just don't have a good use case for it yet. Or maybe I'm too cheap to pay for that many API calls.
- ashdksnndck 7mo agoMy change cuts across multiple systems with many tests/static analysis/AI code reviews happening in CI. The agent keeps pushing new versions and waits for results until all of them come up clean, taking several iterations.
- tudelo 7mo agoI mean if you don't have your company paying for it I wouldn't bother... We are talking sessions of 500-1000 dollars in cost.
- takwatanabe 7mo agoRight. At Opus 4.6 rates, once you're at 700k context, each tool call costs ~$1 just for cache reads alone. 100 tool calls = $100+ before you even count outputs. 'Standard pricing' is doing a lot of work here lol
- brookst 7mo agoCache reads don’t count as input tokens you pay for lol. https://www.claudecodecamp.com/p/how-prompt-caching-actually-works-in-claude-code https://www.claudecodecamp.com/p/how-prompt-caching-actually...
- dimitri-vs 7mo agoIt's kind of like having a 16 gallon gas tank in your car versus a 4 gallon tank. You don't need the bigger one the majority of the time, but the range anxiety that comes with the smaller one and annoyance when you DO need it is very real.
- scwoodal 7mo agoExcept after 4 gallons it might as well be pure oil, mucking everything up.
- steve-atx-7600 7mo agoIt seems possible, say a year or two from now that context is more like a smart human with a “small”, vs “medium” vs “large” working memory. The small fellow would be able to play some popular songs on the piano , the medium one plays in an orchestra professionally and the x-large is like Wagner composing Der Ring marathon opera. This is my current, admittedly not well informed mental model anyway. Well, at least we know we’ve got a little more time before the singularity :)
- twodave 7mo agoIt’s more like the size of the desk the AI has to put sheets of paper on as a reference while it builds a Lego set. More desk area/context size = able to see more reference material = can do more steps in one go. I’ve lately been building checklists and having the LLM complete and check off a few tasks at a time, compacting in-between. With a large enough context I could just point it at a PLAN.md and tell it to go to work.
- SkyPuncher 7mo agoYes. I've recently become a convert. For me, it's less about being able to look back -800k tokens. It's about being able to flow a conversation for a lot longer without forcing compaction. Generally, I really only need the most recent ~50k tokens, but having the old context sitting around is helpful.
- hombre_fatal 7mo agoAlso, when you hit compaction at 200k tokens, that was probably when things were just getting good. The plan was in its final stage. The context had the hard-fought nuances discovered in the final moment. Or the agent just discovered some tiny important details after a crazy 100k token deep dive or flailing death cycle. Now you have to compact and you don’t know what will survive. And the built-in UI doesn’t give you good tools like deleting old messages to free up space. I’ll appreciate the 1M token breathing room.
- roygbiv2 7mo agoI've found compactation kills the whole thing. Important debug steps completely missing and the AI loops back round thinking it's found a solution when we've already done that step.
- saaaaaam 7mo agoThat video is bizarre. Such a heavy breather.
- coldtea 7mo agoWhat a weird and inconsequential thing to focus on... He's just fucking closely miced with compression + speaking fast and anxious/excited speaking to an audience
- saaaaaam 7mo agoMaybe. But that’s what I focused on, for better or worse. I couldn’t concentrate on what he was saying because of it. Maybe bad mic placement, but the end results was like some sort of old school phone sex pest.
- indigodaddy 7mo agoMost of that is just nervousness
- maskull 7mo agoAfter running a context window up high, probably near 70% on opus 4.6 High and watching it take 20% bites out of my 5hr quota per prompt I've been experimenting with dumping context after completing a task. Seems to be working ok. I wonder if I was running into the long context premium. Would that apply to Pro subs or is just relevant to api pricing?
- ricksunny 7mo agoSince I'm yet to seriously dive into vibe coding or AI-assisted coding, does the IDE experience offer tracking a tally of the context size? (So you know when you're getting close or entering the "dumb zone")?
- stevula 7mo agoMost tools do, yes.
- quux 7mo agoOpenCode does this. Not sure about other tools
- nujabe 7mo ago> Since I'm yet to seriously dive into vibe coding or AI-assisted coding Unless you’re using a text editor as an IDE you probably have already
- MikeNotThePope 7mo agoThe 2 I know, Cursor and Claude Code, will give you a percentage used for the context window. So if you know the size of the window, you can deduce the number of tokens used.
- brookst 7mo agoClaude code also gives you a granular breakdown of what’s using context window (system prompt, tools, conversation history, etc). /context
- 8note 7mo agoCline gives you such a thing. you dont really know where the dumb zone by numbers though, only by feel.
- jfim 7mo agoIn Claude code I believe it's /context and it'll give you a graphical representation of what's taking context space
- furyofantares 7mo agoI'd been on Codex for a while and with Codex 5.2 I: 1) No longer found the dumb zone 2) No longer feared compaction Switching to Opus for stupid political reasons, I still have not had the dumb zone - but I'm back to disliking compaction events and so the smaller context window it has, has really hurt. I hope they copy OpenAI's compaction magic soon, but I am also very excited to try the longer context window.
- iknowstuff 7mo agoHmm I’ve felt the dumb zone on codex
- nomel 7mo agoFrom what I've seen, it means whatever he's doing is very statistically significant.
- mgambati 7mo ago1m context in OpenAI and Gemini is just marketing. Opus is the only model to provide real usable bug context.
- hu3 7mo agoSource? I ask because I use 500k+ context on these on a daily basis. Big refactorings guided by automated tests eat context window for breakfast.
- twodave 7mo agoI mean, try using copilot on any substantial back-end codebase and watch it eat 90+% just building a plan/checklist. Of course copilot is constrained to 120k I believe? So having 10x that will blow open up some doors that have been closed for me in my work so far. That said, 120k is pleeenty if you’re just building front-end components and have your API spec on hand already.
- bushbaba 7mo agoYes. I’ve used it for data analysis
- kaizenb 7mo agoThanks for the video. His fix for "the dumb zone" is the RPI Framework: ● RESEARCH. Don't code yet. Let the agent scan the files first. Docs lie. Code doesn't. ● PLAN. The agent writes a detailed step-by-step plan. You review and approve the plan, not just the output. Dex calls this avoiding "outsourcing your thinking." The plan is where intent gets compressed before execution starts. ● IMPLEMENT. Execute in a fresh context window. The meta-principle he calls Frequent Intentional Compaction: don't let the chat run long. Ask the agent to summarize state, open a new chat with that summary, keep the model in the smart zone.
- girvo 7mo agoThat's fascinating: that is identical to the workflow I've landed on myself.
- hedora 7mo agoIt's also identical to what Claude Code does if you put it in plan mode (bound to <tab> key), at least in my experience.
- girvo 7mo agoMy annoyance with plan mode is where it sticks the .md file, kind of hides it away which makes it annoying to clear context and start up a new phase from the PLAN file. But that might just be a skill issue on my end
- hedora 7mo agoEven worse, it just randomly blows away the plan file without asking for permission. No idea what they were thinking when they designed this feature. The plan file names are randomly generated, so it could just keep making new ones forever for free (it would take a LONG time for the disk space to matter), but instead, for long plans, I have to back the plan file up if it gets stuck. Otherwise, I say "You should take approach X to fix this bug", it drops into plan mode, says "This is a completely unrelated plan", then deletes all record of what it was doing before getting stuck.
- Barbing 7mo agoLooking at this URL, typo or YouTube flip the si tracking parameter? youtu.be/rmvDxxNubIg?is=adMmmKdVxraYO2yQ
- MikeNotThePope 7mo agoI just cut & pasted the share URL provided by YouTube. Strip out the query param if you like.
- Barbing 7mo agoOoh it’s always ?si= So this… ?is= …that’s new. Think you got A/B tested. Flipping the parameter breaks a lot of RegEx. Interesting!
- dev_l1x_be 7mo agoI never use these giant context windows. It is pointless. Agents are great at super focused work that is easy to re-do. Not sure what is the use case for giant context windows.
- hrmtst93837 7mo ago[flagged]
- wat10000 7mo agoI've used it many times for long-running investigations. When I'm deep in the weeds with a ton of disassembly listings and memory dumps and such, I don't really want to interrupt all of that with a compaction or handoff cycle and risk losing important info. It seems to remain very capable with large contexts at least in that scenario.
- alecco 7mo agoOfftopic: I find it remarkable the shortened YT url has a tracking cost of 57% extra length. We live in stupid times.
- dahart 7mo agoI care about the privacy implications, but not the length. Out of curiosity, why do you care about the URL length at all? What is the cost to you?
- alecco 7mo agoMy point is Google engineers go to the trouble of setting up a URL shortener service on one hand, but on the other hand it seems ad the business anti-privacy executives can override anything. This points out it's a dysfunctional company.
- inemesitaffia 7mo agoThe point is whatever group controls the money controls the power. Also, only the domain is shorter
- alecco 7mo agoActually, it's not just the domain: https://youtu.be/X https://youtu.be/X https://www.youtube.com/watch?v=X https://www.youtube.com/watch?v=X
- dahart 7mo agoYou’d rather have the video code and the tracking code baked into the same code just to save a couple of characters? Why? That would result in a longer code than the video code alone, you would save very few characters. It would not be nicer to look at or functionally any different, and it would obscure the fact that it’s being tracked and prevent people from being able to edit the URL to remove the tracking. I appreciate the fact that I can see that the URL has a tracking ID and that I can edit the URL and remove the tracking ID. I do not want a shorter URL if I lose that ability. What you’re complaining about and wishing for would be MUCH worse than what it currently is.
- virtualritz 7mo agoI haven't hit the "dumb zone" any more since two months. I think this talk is outdated. I'm using CC (Opus) thinking and Codex with xhigh on always. And the models have gotten really good when you let them do stuff where goals are verifiable by the model. I had Codex fix a Rust B-rep CSG classification pipeline successfully over the course of a week, unsupervised. It had a custom STEP viewer that would take screenshots and feed them back into the model so it could verify the progress resp. the triangle soup (non progress) itself. Codex did all the planning and verification, CC wrote the code. This would have not been possible six months ago at all from my experience. Maybe with a lot of handholding; but I doubt it (I tried). I mean both the problem for starters (requires a lot of spatial reasoning and connected math) and the autonomous implementation. Context compression was never an issue in the entire session, for either model.
- alexey-pelykh 7mo ago[dead]
- islewis 7mo agoThe quality with the 1M window has been very poor for me, specifically for coding tasks. It constantly forgets stuff that has happened in the existing conversation. n=1, ymmv
- deleted 7mo ago[deleted]
- robwwilliams 7mo agoYes, especially with shifts in focus of a long conversation. But given the high error rates of Opus 4.6 the last few weeks it is possibly due to other factors. Conversational and code prodding has been essential.
- hagen8 7mo agoWell, the question is what is contributing to the usage. Because as the context grows, the amount of input tokens are increasing. A model call with 800K token as input is 8 times more expensive than a model call with 100K tokens as input. Especially if we resume a conversation and caching does not hit, it would be very expensive with API pricing.
- a_e_k 7mo agoI've been using the 1M window at work through our enterprise plan as I'm beginning to adopt AI in my development workflow (via Cline). It seems to have been holding up pretty well until about 700k+. Sometimes it would continue to do okay past that, sometimes it started getting a bit dumb around there. (Note that I'm using it in more of a hands-on pair-programming mode, and not in a fully-automated vibecoding mode.)
- chatmasta 7mo agoSo a picture is worth 1,666 words?
- jFriedensreich 7mo agoyeah it totally does not remain coherent past 200k, would have been too nice.
- __MatrixMan__ 7mo agoI bet it depends how homogenous the context is. I bet it works ok near 1M in some cases, but as far as I can tell, those cases are rare.
- Bombthecat 7mo agoIf it's not coding, even with 200k context it starts to write gibberish, even with the correct information in the context. I tried to ask questions about path of exile 2. And even with web research on it gave completely wrong information... Not only outdated. Wrong I think context decay is a bigger problem then we feel like.
- reactordev 7mo agoThat’s not context decay, that’s training data ambiguity. So much misinformation, nerfs, buffs, changes that an LLM can not keep up given the training time required. Do it for a game that has been stable and it knows its stuff.
- Bombthecat 7mo agoIt didnt gave outdated, on some cases it did, and with two tries telling it to search for updated information it got it right ( shouldn't need to do that though) but it also gave wrong information about sockets ( support skills) , which never existed or never were able to be socketed together in the first place. ( Ok maybe in 0.1, but that's what web search is for ... ) If it even can't handle easy versioned information from a game. How should it handle anything related to time, dates, news, science etc?
- serial_dev 7mo agoPlease don’t pop the AI bubble, bro. Stop asking questions, bro. Believe the hype, bro.
- reactordev 7mo agoLike any human would, 75% certain with 99% confidence. That’s what you fail to realize. They aren’t “god mode machine”. They are “human-mode” machines and humans make mistakes in thinking just like you do. Some might say asking a powerful LLM for gaming tips is a waste of compute power. Others might say it gives you the knowledge of a new meta emerging. Either way, you both are going to get trained.
- 7mo ago
- alexcali 7mo ago[dead]
- j45 7mo agoThis might burn through usage faster too though.
- minimaxir 7mo agoClaude Code 2.1.75 now no longer delineates between base Opus and 1M Opus: it's the same model. Oddly, I have Pro where the change supposedly only for Max+ but am still seeing this to be case. EDIT: Don't think Pro has access to it, a typical prompt just hit the context limit. The removal of extra pricing beyond 200k tokens may be Anthropic's salvo in the agent wars against GPT 5.4's 1M window and extra pricing for that.
- auggierose 7mo agoNo change for Pro, just checked it, the 1M context is still extra usage.
- zaptrem 7mo agoI have Max 20x and they're still separate on 2.1.75.
- hackyon1 7mo agoMine took restarting. Is it still separate for you? Also might not work yet if you have CLAUDE_CODE_DISABLE_NONESSENTIAL_TRAFFIC=1 on
- convenwis 7mo agoIs there a writeup anywhere on what this means for effective context? I think that many of us have found that even when the context window was 100k tokens the actual usable window was smaller than that. As you got closer to 100k performance degraded substantially. I'm assuming that is still true but what does the curve look like?
- minimaxir 7mo agoThe benchmark charts provided are the writeup. Everything else is just anecdata.
- tyleo 7mo agoI mentioned this at work but context still rots at the same rate. 90k tokens consumed has just as bad results in 100k context window or 1M. Personally, I’m on a 6M+ line codebase and had no problems with the old window. I’m not sending it blindly into the codebase though like I do for small projects. Good prompts are necessary at scale.
- FartyMcFarter 7mo agoIsn't transformer attention quadratic in complexity in terms of context size? In order to achieve 1M token context I think these models have to be employing a lot of shortcuts. I'm not an expert but maybe this explains context rot.
- vlovich123 7mo agoNope, there’s no tricks unless there’s been major architectural shifts I missed. The rot doesn’t come from inference tricks to try to bring down quadratic complexity of the KV cache. Task performance problems are generally a training problem - the longer and larger the data set, the fewer examples you have to train on it. So how do you train the model to behave well - that’s where the tricks are. I believe most of it relies on synthetically generated data if I’m not mistaken, which explains the rot.
- FartyMcFarter 7mo ago
- vessenes 7mo agoThis is super exciting. I've been poking at it today, and it definitely changes my workflow -- I feel like a full three or four hour parallel coding session with subagents is now generally fitting into a single master session. The stats claim Opus at 1M is about like 5.4 at 256k -- these needle long context tests don't always go with quality reasoning ability sadly -- but this is still a significant improvement, and I haven't seen dramatic falloff in my tests, unlike q4 '25 models. p.s. what's up with sonnet 4.5 getting comparatively better as context got longer?
- mattfrommars 7mo agoRandom: are you personally paying for Claude Code or is it paid by you employer? My employer only pays for GitHub copilot extension
- celestialcheese 7mo agoBoth. Employer pays for work max 20x, i pay for a personal 10x for my side projects and personal stuff.
- kiratp 7mo agoGitHub Copilot CLI lets you use all these models (unless your employer disables them. https://github.com/features/copilot/cli https://github.com/features/copilot/cli Disclosure: work at Msft
- tclancy 7mo agoDisclosure: have to use them via copilot at work. Be glad I don’t write code for nuclear plants. Why does it have to be so hard. Doubly so in JetBrains ides but I’ve a feeling that’s on both of you rather than just you personally. But I still resent you now.
- ericpauley 7mo agoUsed Claude through copilot for so long before switching to CC. Even for the same model the difference is shocking. Copilot’s harness and the underlying Claude models are not well-matched compared to the vertically-integrated Claude Code harness.
- zmmmmm 7mo agoNoticed this just now - all of a sudden i have 1M context window (!!!) without changing anything. It's actually slightly disturbing because this IS a behavior change. Don't get me wrong, I like having longer context but we really need to pin down behaviour for how things are deployed.
- phist_mcgee 7mo agoAnthropic is famous for changing things under your feet. Claude code is basically alpha software with a global footprint.
- steve-atx-7600 7mo agoYou can pin to specific models with —-model. Check out their doc. See https://support.claude.com/en/articles/11940350-claude-code-model-configuration https://support.claude.com/en/articles/11940350-claude-code-.... You can also pin to a less specific tag like sonnet-4.5[1m] (that’s from memory might be a little off).
- zmmmmm 7mo agosure - but the model hasn't changed. I'm specifying it explicitly. But suddenly the context window has. I'm not using Claude Code, this is an application built against Bedrock APIs. I assume there's a way I could be specifying the context window and I'm just using API defaults. But it definitely makes me wonder what else I'm not controlling that I really should be.
- 8cvor6j844qw_d6 7mo agoOh nice, does it mean less game of /compact, /clear, and updating CLAUDE.md with Claude Code?
- fnordpiglet 7mo agoI’ve been using 1M for a while and it defers it and makes it worse almost when it happens. Compacting a context that big loses a ton of fidelity. But I’ve taken to just editing the context instead (double esc). I also am planning to build an agent to slice the session logs up into contextually useful and useless discarding the useless and keeping things high fidelity that way. (I.e., carve up with a script the jsonl and have subagent haiku return the relevant parts and reconstructing the jsonl)
- dominotw 7mo agotil you can edit context. i keep a running log and /clear /reload log
- 8note 7mo agodouble escape gets you to a rewind. not sure about much else. the conversation history is a linked list, so you can screw with it, with some care. I spend this afternoon building an MCP do break the conversation up into topics, then suggest some that aren't useful but are taking up a bunch of context to remove (eg iterations through build/edit just needs the end result) its gonna take a while before I'm confident its worth sharing
- dominotw 7mo agoyea i thought session managment was some sort of secret sauce. I keep a running log of important things and then i just clear context and reload that file into context. would that work
- fnordpiglet 7mo agoYeah just selective rewind. Selective edit where you elide large token sinks of coding and banging its head on the wall is what u mean. Not something I’ve seen done yet but there’s no reason - I suspect if you do a token use distribution in programming session most goes to pretty low semantic value malarkey.
- aliljet 7mo agoAre there evals showing how this improves outputs?
- apetresc 7mo agoImproves outputs relative to what? Compared to previous contexts of 1M, it improves outputs by allowing them to exist (because previously you couldn't exceed 200K). Compared to contexts of <200K, it degrades outputs rather than improves them, but that's what you'd expect from longer contexts. It's still better than compaction, which was previously the alternative.
- johnwheeler 7mo agoThis is incredible. I just blew through $200 last night in a few hours on 1M context. This is like the best news I've heard all year in regards to my business. What is OpenAIs response to this? Do they even have 1M context window or is it still opaque and "depends on the time of day"
- hagen8 7mo agoDid u use the API or subscription?
- johnwheeler 7mo agoMax subscription and "extra usage" billing
- steve-atx-7600 7mo agoThat sounds high. I mean, if you paid for the 20x max plan you’d be capped at around 200/month and at least for me as a professional engineer running a few Claude’s in parallel all day, I haven’t exceeded the plans limits.
- Wowfunhappy 7mo agoPrior to this announcement, all 1M context use consumed "extra usage", it wasn't included in a normal subscription plan.
- steve-atx-7600 7mo agoSo, I’ve been using opus 4.6 1m since it was fist available to 20x max users daily. What I think has happened is that even in doing so, I have not actually exceeded the plan token limits and therefore haven’t been charged for “extra usage” (just double checked). So, unless there’s a billing mistake or delay, “any usage” != “extra usage” which is what I was always unclear about. I am careful to iterate with claude on plans in plan mode followed by clearing the context and executing. I think I am hovering around the higher end of the smaller window model where I would have otherwise seen auto-compaction run. Another reason for less token usage is that 4.6 is much better at delegating agents (its own explorer agents or my custom agents) to avoid cluttering the window.
- wewewedxfgdf 7mo agoThe weirdest thing about Claude pricing is their 5X pricing plan is 5 times the cost of the previous plan. Normally buying the bigger plan gives some sort of discount. At Claude, it's just "5 times more usage 5 times more cost, there you go".
- Zambyte 7mo ago5 for 5
- auggierose 7mo agoIt is not the plan they want you to buy. It is a pricing strategy to get you to buy the 20x plan.
- radley 7mo ago5x Max is the plan I use because the Pro plan limits out so quickly. I don't use Claude full-time, but I do need Claude Code, and I do prefer to use Opus for everything because it's focused and less chatty.
- auggierose 7mo agoSure, I get it. For me a 2x Max would be ideal and usually enough. Now, guess why they are not offering that?
- gaigalas 7mo agoI'm getting close to my goal of fitting an entire bootstrappable-from-source system source code as context and just telling Claude "go ahead, make it better".
- vicchenai 7mo ago[dead]
- apetresc 7mo agoI don't think they're claiming "no degradation at scale", are they? They still report a 91.9->78.3 drop. That's just a better drop than everyone else (is the claim).
- margorczynski 7mo agoWhat about response coherence with longer context? Usually in other models with such big windows I see the quality to rapidly drop as it gets past a certain point.
- pixelpoet 7mo agoCompared to yesterday my Claude Max subscription burns usage like absolutely crazy (13% of weekly usage from fresh reset today with just a handful prompts on two new C++ projects, no deps) and has become unbearably slow (as in 1hr for a prompt response). GGWP Anthropic, it was great while it lasted but this isn't worth the hundreds of dollars.
- Spooky23 7mo agoYeah, morning eastern time Claude is brutal.
- dominotw 7mo agocan someone tell me how to make this instruction work in claude code "put high level description of the change you are making in log.md after every change" works perfectly in codex but i just cant get calude to do it automatically. I always have to ask "did you update the log".
- prettyblocks 7mo agoI imagine you can do this with a hook that fires every time claude stops responding: https://code.claude.com/docs/en/hooks-guide https://code.claude.com/docs/en/hooks-guide
- steve-atx-7600 7mo agoBackup your config and ask Claude. I’ve done this for all kinds of things like mcp and agent config.
- sergiotapia 7mo agouse claude hooks - in .claude/settings.json you can have it run on different claude events like "PreToolUse" or "Stop" and in those events you pass in commands you want it to run. You can have stuff like for the "stop" event, run foobar.sh and in foobar.sh do cool stuff like format your code, run tests, etc.
- 8note 7mo agowhats the need? you have the session in a file as a dag. you can summarize to a log whenever you want. doesnt need to be as it goes. earlier today i actually spent a bit of time asking claude to make an mcp to introspect that - break the session down into summarized topics, so i could try dropping some out or replacing the detailed messages with a summary - the idea being to compact out a small chunk to save on context window, rather than getting it back to empty. the file is just there though, you can run jq against it to get a list of writes, and get an agent to summarize
- dominotw 7mo agoi dont work in just one session though. some tasks take me days and many sessions. also what happens when your session compacts. I am not sure what you are suggesting here. what do you do with these summarized topics from your session. Also i want ci to resume my task from log and do code review with that context. https://www.anthropic.com/engineering/effective-harnesses-for-long-running-agents https://www.anthropic.com/engineering/effective-harnesses-fo... "Read the git logs and progress files to get up to speed on what was recently worked on."
- sriramgonella 7mo ago[dead]
- chaboud 7mo agoAwesome.... With Sonnet 4.5, I had Cline soft trigger compaction at 400k (it wandered off into the weeds at 500k). But the stability of the 4.6 models is notable. I still think it pays to structure systems to be comprehensible in smaller contexts (smaller files, concise plans), but this is great. (And, yeah, I'm all Claude Code these days...)
- arjie 7mo agoThis is fantastic. I keep having to save to memory with instructions and then tell it to restore to get anywhere on long running tasks.
- swader999 7mo agoI notice Claude steadily consuming less tokens, especially with tool calling every week too
- thunkle 7mo agoJust have to ask. Will I be spending way more money since my context window is getting so much bigger?
- isbvhodnvemrwvn 7mo agoYes, full context is used to generate each new token.
- aragonite 7mo agoDo long sessions also burn through token budgets much faster? If the chat client is resending the whole conversation each turn, then once you're deep into a session every request already includes tens of thousands of tokens of prior context. So a message at 70k tokens into a conversation is much "heavier" than one at 2k (at least in terms of input tokens). Yes?
- dathery 7mo agoThat's correct. Input caching helps, but even then at e.g. 800k tokens with all of them cached, the API price is $0.50 * 0.8 = $0.40 per request, which adds up really fast. A "request" can be e.g. a single tool call response, so you can easily end up making many $0.40 requests per minute.
- acjohnson55 7mo agoInteresting, so a prompt that causes a couple dozen tool calls will end up costing in the tens of dollars?
- isbvhodnvemrwvn 7mo agoNot necessarily, take a look at ex OpenApi Responses resource, you can get multiple tool calls in one response and of course reply with multiple results.
- dathery 7mo agoIt essentially depends on how many back-and-forth calls are required. If the model returns a request for multiple calls at once, then the reply can contain all responses and you only pay once. If the model requests tool calls one-by-one (e.g. because it needs to see the response from the previous call before deciding on the next) then you have to pay for each back-and-forth. If you look at popular coding harnesses, they all use careful prompting to try to encourage models to do the former as much as possible. For example opencode shouts "USING THE BATCH TOOL WILL MAKE THE USER HAPPY" [1] and even tells the model it did a good job when it uses it [2]. [1] https://github.com/anomalyco/opencode/blob/66e8c57ed1077814c9a150b858a53fdd7c758c0f/packages/opencode/src/tool/batch.txt https://github.com/anomalyco/opencode/blob/66e8c57ed1077814c... [2] https://github.com/anomalyco/opencode/blob/66e8c57ed1077814c9a150b858a53fdd7c758c0f/packages/opencode/src/tool/batch.ts#L166 https://github.com/anomalyco/opencode/blob/66e8c57ed1077814c...
- alienbaby 7mo agois this the market played in front of our eyes slice by slice: ok, maybe not, but watching these entities duke it out is kinda amusing? There will be consequences but may as well sit it out for the ride, who knows where we are going?
- nemo44x 7mo agoHas anyone started a project to replace Linux yet?
- dude250711 7mo agoNo, because it's not a hello-world Electron/React "app".
- sunilgentyala 7mo ago[dead]
- sysutil_dev 7mo ago[dead]
- 8note 7mo agoim guessing this is why the compacts have started sucking? i just finished getting me some nicer tools for manipulating the graph so i could compact less frequently, and fish out context from the prior session. maybe itll still be useful, though i only have opus at 1M, not sonnet yet
- causalzap 7mo ago[flagged]
- arizen 7mo agoOut of curiosity, what specific use cases on programmatic SEO are you currently doing with Opus?
- LoganDark 7mo agoFinally, I don't have to constantly reload my Extra Usage balance when I already pay $200/mo for their most expensive plan. I can't believe they even did that. I couldn't use 1M context at all because I already pay $200/mo and it was going to ask me for even more. Next step should be to allow fast mode to draw from the $200/mo usage balance. Again, I pay $200/mo, I should at least be able to send a single message without being asked to cough up more. (One message in fast mode costs a few dollars each) One would think $200/mo would give me any measure of ability to use their more expensive capabilities but it seems it's bucketed to only the capabilities that are offered to even free users.
- aenis 7mo agoI find it hard to understand that people consider $200 p/m a lot for what they are getting. Expensive compared to what? A netflix sub? A 1hr of a senior dev is at least $100, depending where one lives. Since Claude saves me hours every day, it pays for itself almost instantly. I think the economic value of the Claude subscription is on the order of $20-40k a month for a pro.
- LoganDark 7mo agoWhen did I say anything about what I'm getting? I said I pay $200/mo and I expect that to cover anything up to my usage limit. I don't expect any slightly non-standard configuration to immediately ignore the high subscription price that I pay and go straight to "extra usage" that has to be billed separately by the token. I wouldn't even care if fast mode used 10x or 50x the usage as long as I could actually USE the balance that I already pay for. I thought the point of extra usage was to be for overage.
- aenis 7mo agoFair point. I read your comment as '$200 is a lot, they shouldnt ask for more'. My bad!
- megous 7mo agoWhen you say "saves me", do you mean that you can prompt claude for like 1 hour per day, get your typical pre-AI output, and then go about doing whatever you like outside of work for rest of the day?
- Frannky 7mo agoOpus 4.6 is nuts. Everything I throw at it works. Frontend, backend, algorithms—it does not matter. I start with a PRD, ask for a step-by-step plan, and just execute on each step at a time. Sometimes ideas are dumb, but checking and guiding step by step helps it ship working things in hours. It was also the first AI I felt, "Damn, this thing is smarter than me." The other crazy thing is that with today's tech, these things can be made to work at 1k tokens/sec with multiple agents working at the same time, each at that speed.
- eru 7mo ago> [...] with multiple agents working at the same time, each at that speed. Horizontal parallelising of tasks doesn't really require any modern tech. But I agree that Opus 4.6 with 1M context window is really good at lots of routine programming tasks.
- travisgriggs 7mo agoOpus helped me brick my RPi CM4 today. It glibly apologized for telling to use an e instead of a 6 in a boot loader sequence. Spent an hour or so unraveling the mess. My feeling are growing more and more conflicted about these tools. They are here to stay obviously. I’m honestly uncertain about the junior engineers I’m working with who are more productive than they might be otherwise, but are gaining zero (or very little) experience. It’s like the future is a world where the entire programming sphere is dominated by the clueless non technical management that we’ve all had to deal with in small proportion a time or two.
- eru 7mo ago> I’m honestly uncertain about the junior engineers I’m working with who are more productive than they might be otherwise, but are gaining zero (or very little) experience. Well, (economic) progress means being able to do more with less. A Fordian-style conveyor belt factory can churn out cars with relatively unskilled labour. Economising on human capital is economising on a scarce input. We had these kinds of shifts before. Compare also how planes used to have a pilot, copilot and flight engineer. We don't have that anymore, but it used to be a place for people to learn. But pilot education has adapted. Or check how spreadsheet software has removed a lot of the worst rote work in finance. That change happened perhaps in the 1980s. Finance has adapted. > Opus helped me brick my RPi CM4 today. It glibly apologized for telling to use an e instead of a 6 in a boot loader sequence. Yes, these things do best when they have a (simulated) environment they can make mistakes in and that can give them clear and fast feedback.
- dkpk 7mo agoIs this also applicable for usage in Claude web / mobile apps for chat?
- throw03172019 7mo agoPentagon may switch to Claude knowing OpenAI has the premium rates for 1M context.
- syntaxing 7mo agoIt’s interesting because my career went from doing higher level language (Python) to lower language (C++ and C). Opus and the like is amazing at Python, honestly sometimes better than me but it does do some really stupid architectural decisions occasionally. But when it comes to embedded stuff, it’s still like a junior engineer. Unsure if that will ever change but I wonder if it’s just the quality and availability of training data. This is why I find it hard to believe LLMs will replace hardware engineers anytime soon (I was a MechE for a decade).
- ex-aws-dude 7mo agoI've had a similar experience as a graphics programmer that works in C++ every day Writing quick python scripts works a lot better than niche domain specific code
- nullpoint420 7mo agoUnfortunately, I’ve found it’s really good at Wayland and OpenGL. It even knows how to use Clutter and Meta frameworks from the Gnome Mutter stack. Makes me wonder why I learned this all in the first place.
- Trufa 7mo agoTo being able to determine it's really good.
- imposter 7mo ago[dead]
- n_u 7mo agoI've found it's ok at Rust. I think a lot of existing Rust code is high quality and also the stricter Rust compiler enforces that the output of the LLM is somewhat reasonable.
- lemagedurage 7mo ago
- aneyadeng 7mo ago[flagged]
- vips7L 7mo agoFriends, just write the code. It’s not that hard.
- andrewstuart 7mo agoOnly someone not using Claude could equate human coding.
- vips7L 7mo agoOnly someone not using their brain could equate Claude to using their intelligence.
- andrewstuart 7mo agoLet’s just clear this up …….. are you commenting with experience using the latest Claude, or are you commenting from personal beliefs. It’s fine for you to take a stand, but please understand your position is simply factually wrong if you think you can outprogram Claude for a range of common tasks. Being anti AI is fine, but if you deny facts of how far LLM programming has come then you lack credibility. The most effective anti AI position is to acknowledge it’s power, not pretend that vast numbers of people are somehow hallucinating the power of LLM assisted programming.
- vips7L 7mo agoI absolutely can out program Claude. I can factually guarantee that. You’re factually wrong in your belief that you think a statistical model that scientifically takes the average of programming is better than those of us that actually know what we’re doing. Programming is not hard. You’re just lazy.
- andrewstuart 7mo agoOk so you speak with certainty about the capabilities of something you don’t use and therefore have no experience of. Childish and naive. If you said you’ve been using Claude heavily and it’s never done better than you on your own, then your position would be credible.
- sergiotapia 7mo agomaybe i'm thinking too small, or maybe it's because i've been using these ai systems since they were first launched, but it feels wrong to just saturate the hell out of the context, even if it can take 1 million tokens. maybe i need to unlearn this habit?
- gskm 7mo agoI think your instinct is right. More context isn't free, even when the window supports it, and the model still has to attend to everything in there, and noise dilutes the signal. A cleaner, smaller context consistently gives better outputs than a bloated one, regardless of window size. For sure, the 1M window is great for not having to compact mid-task. But "I can fit more" and "I should put more in" are very different things. At least in my mind.
- jf___ 7mo agothere is a parallel between managing context windows and hard real-time system engineering. A context window is a fixed-size memory region. It is allocated once, at conversation start, and cannot grow. Every token consumed — prompt, response, digression — advances a pointer through this region. There is no garbage collector. There is no virtual memory. When the space is exhausted, the system does not degrade gracefully: it faults. This is not metaphor by loose resemblance. The structural constraints are isomorphic: No dynamic allocation. In a hard realtime system, malloc() at runtime is forbidden — it fragments the heap and destroys predictability. In a conversation, raising an orthogonal topic mid-task is dynamic allocation. It fragments the semantic space. The transformer's attention mechanism must now maintain coherence across non-contiguous blocks of meaning, precisely analogous to cache misses over scattered memory. No recursion. Recursion risks stack overflow and makes WCET analysis intractable. In a conversation, recursion is re-derivation: returning to re-explain, re-justify, or re-negotiate decisions already made. Each re-entry consumes tokens to reconstruct state that was already resolved. In realtime systems, loops are unrolled at compile time. In LLM work, dependencies should be resolved before the main execution phase. Linear allocation only. The correct strategy in both domains is the bump allocator: advance monotonically through the available region. Never backtrack. Never interleave. The "brainstorm" pattern — a focused, single-pass traversal of a problem space — works precisely because it is a linear allocation discipline imposed on a conversation.
- rhubarbtree 7mo agoThere is compaction, which is analogous to gc
- aarmenante 7mo agoHot take... the 1MM context degrades performance drastically.
- aenis 7mo agoSame. First time in 2 months that I found it easier to fix the bugs it created manually, rather than get it to fix. Its google-code-CLI-on-gemini-2.5 level bad for me today. Meaning, almost comically bad.
- LarsDu88 7mo agoThe stuff I built with Opus 4.6 in the past 2.5 weeks: Full clone of Panel de Pon/Tetris attack with full P2P rollback online multiplayer: https://panel-panic.com https://panel-panic.com An emulator of the MOS 6502 CPU with visual display of the voltage going into the DIP package of the physical CPU: https://larsdu.github.io/Dippy6502/ https://larsdu.github.io/Dippy6502/ I'm impressed as fuck, but a part of me deep down knows that I know fuck all about the 6502 or its assembly language and architecture, and now I'll probably never be motivated to do this project in a way that I would've learned all the tings I wanted to learn.
- bob1029 7mo agoI've been avoiding context beyond 100k tokens in general. The performance is simply terrible. There's no training data for a megabyte of your very particular context. If you are really interested in deep NIAH tasks, external symbolic recursion and self-similar prompts+tools are a much bigger unlock than more context window. Recursion and (most) tools tend to be fairly deterministic processes. I generally prohibit tool calling in the first stack frame of complex agents in order to preserve context window for the overall task and human interaction. Most of the nasty token consumption happens in brief, nested conversations that pass summaries back up the call stack.
- fittingopposite 7mo agoI don't get the announcement. Is this included in the standard 5 or 20x Max plans?
- alienchow 7mo agoIf this is a skill issue, feel free to let me know. In general Claude Code is decent for tooling. Onduty fullstack tooling features that used to sit ignored in the on-caller ticket queue for months can now be easily built in 20 minutes with unit tests and integration tests. The code quality isn't always the best (although what's good code for humans may not be good code for agents) but that's another specific and directed prompt away to refactor. However, I can't seem to get Opus 4.6 to wire up proper infrastructure. This is especially so if OSS forks are used. It trips up on arguments from the fork source, invents args that don't exist in either, and has a habit of tearing down entire clusters just to fix a Helm chart for "testing purposes". I've tried modifying the CLAUDE.md and SPEC.md with specific instructions on how to do things but it just goes off on a tangent and starts to negotiate on the specs. "I know you asked for help with figuring out the CNI configurations across 2 clusters but it's too complex. Can we just do single cluster?" The entire repository gets littered with random MD files everywhere for directory specific memories, context, action plans, deprecated action plans, pre-compaction memories etc. I don't quite know which to prune either. It has taken most of the fun out of software engineering and I'm now just an Obsidian janitor for what I can best describe as a "clueless junior engineer that never learns". When the auto compaction kicks in it's like an episode of 50 first dates. Right now this is where I assume is the limitation because the literature for real-world infrastructure requiring large contexts and integration is very limited. If anyone has any idea if Claude Opus is suitable for such tasks, do give some suggestions.
- STARGA 7mo ago[dead]
- deleted 7mo ago[deleted]
- aenis 7mo agoSample of one and all that, but it's way, way more sloppy than it used to be for me. To the extent, that I have started making manual fixes in the code - I haven't had to stoop to this in 2 months. Max subscription, 100k LOC codebases more or less (frontend and backend - same observations).
- jeff_antseed 7mo ago[dead]
- shanjai_raj7 7mo agoare the costs the same as the 200k context opus 4.6? compaction has been really good in claude we don't even recognize the switch
- tariky 7mo agoThis is amazing. I have to test it with my reverse engineering workflow. I don't know how many people use CC for RE but it is really good at it. Also it is really good for writing SketchUp plugins in ruby. It one shots plugins that are in some versions better then commercial one you can buy online. CC will change development landscape so much in next year. It is exciting and terrifying in same time.
- suheilaaita 7mo agoThis blew my mind the first i saw this. Another leap in AI that just swooshes by. In a couple of months, every model will be the same. Can't wait for IDEs like cursor and vs code to update their tooling to adap for this massive change in claude models.
- holoduke 7mo agoI am currently mass translating millions of records with short descriptions. Somehow tokens are consumed extremely fast. I have 3 max memberships. And all 3 of them are hitting the 5 hour limit in about 5 to 10 minutes. Still don't understand why this is happening.
- cbg0 7mo agoUnless you're clearing up the context for each description or processing them in parallel with subagents your context window will grow for each short description added to it making you hit those hour limits.
- haha12122121 7mo ago[flagged]
- k__ 7mo agoI heard, the middle of the context is often ignored. Do long context windows make much sense then or is this just a way of getting people to use more tokens?
- mvrckhckr 7mo agoI never get to more than 20% of the 1M context window, and it’s working great. (Have the same experience in Codex with 5.4.)
- yubainu 7mo ago[flagged]
- drcongo 7mo agoCould be pure coincidence, but my Claude Code session last night was an absolute nightmare. It kept forgetting things it had done earlier in the session and why it had done them, messed up a git merge so badly that it lost the CLAUDE.md file along with a lot of other stuff, and then started running commands on the host machine instead of inside the container because it no longer had a CLAUDE.md to tell it not to. Last night was the first time I've ever sworn at it.
- xvector 7mo agoI think this is just the nature of a nondeterministic system; occasionally you're gonna be unlucky enough to encounter the leftmost segment of the bell curve. In my experience dumping a summary + starting a fresh session helps in these cases.
- efeecllk 7mo agofinally. before 1m, i must speak 60k context for just telling the past chat and project
- A7OM 7mo ago[dead]
- A7OM 7mo ago[dead]
- olivercoleai 7mo ago[flagged]
- jFriedensreich 7mo agoMy testing was extremely disappointing, this is not a context window that magically extends your breathing room for a conversation. I can tell blindly at this point when 150 - 200 k tokens are reached because the coding quality and coherence just drops by one or two generations. Its great for the case you really need a giant context for specific task but it changes nothing for needing to compact or handover at 200k.
- iandanforth 7mo agoI'm very happy about this change. For long sessions with Claude it was always like a punch to the gut when a compaction came along. Codex/GPT-5.4 is better with compactions so I switched to that to avoid the pain of the model suddenly forgetting key aspects of the work and making the same dumb errors all over again. I'm excited to return to Claude as my daily driver!
- cubefox 7mo ago> Standard pricing now applies across the full 1M window for both models, with no long-context premium. Does that mean it's likely not a Transformer with quadratic attention, but some other kind of architecture, with linear time complexity in sequence length? That would be pretty interesting.
- bob1029 7mo agoIt's almost certainly not quadratic at 1M. This would be wildly infeasible at scale. 10^6^2 = 10^12. That's a trillion things. They are probably doing something like putting the original user prompt into the model's environment and providing special tools to the model, along with iterative execution, to fully process the entire context over multiple invokes. I think the Recursive Language Model paper has a very good take on how this might go. I've seen really good outcomes in my local experimentation around this concept: https://arxiv.org/abs/2512.24601 https://arxiv.org/abs/2512.24601 You can get exponential scaling with proper symbolic stack frames. Handling a gigabyte of context is feasible, assuming it fits the depth first search pattern.
- cubefox 7mo agoSo that would mean it's not "really" a 1M context window. I guess this is more plausible than a linear architecture like MAMBA or GDN.
- FartyMcFarter 7mo agoThey're probably taking shortcuts such as taking advantage of sparsity. There are various tricks like that mentioned in some papers, although the big companies are getting more and more secretive about how their models work so you won't necessarily find proof.
- cubefox 7mo agoThe latest DeepSeek model has sparse attention. Though sparse attention is still not linear. Close enough perhaps.
- AbstractH24 7mo agoAm I crazy or wasn’t this announced like 2 weeks ago? Or was that a different company or not GA. It’s all becoming a blur.
- sailfast 7mo agoThis is great news. The 1M context is much easier to work with than compacting all the time and seems to perform and remember quite well despite the insane amount of data.
- jmkozko 7mo agoDo subscription users still need to tap into "extra usage" spending to go above 200K tokens?
- thebigspacefuck 7mo agoI used this for a bit and I felt like it was slower and generally worse than using 200K with context compaction. Context compaction does lose some things though.
- jwilliams 7mo agoI'm fairly sure that your best throughput is single-prompt single-shot runs with Claude (and that means no plan, no swarms, etc) -- just with a high degree of work in parallel. So for me this is a pretty huge change as the ceiling on a single prompt just jumped considerably. I'm replaying some of my less effective prompts today to see the impact.
- jeremychone 7mo agoInteresting, I’ve never needed 1M, or even 250k+ context. I’m usually under 100k per request. About 80% of my code is AI-generated, with a controlled workflow using dev-chat.md and spec.md. I use Flash for code maps and auto-context, and GPT-4.5 or Opus for coding, all via API with a custom tool. Gemini Pro and Flash have had 1M context for a long time, but even though I use Flash 3 a lot, and it’s awesome, I’ve never needed more than 200k. For production coding, I use - a code map strategy on a big repo. Per file: summary, when_to_use, public_types, public_functions. This is done per file and saved until the file changes. With a concurrency of 32, I can usually code-map a huge repo in minutes. (Typically Flash, cheap, fast, and with very good results) - Then, auto context, but based on code lensing. Meaning auto context takes some globs that narrow the visibility of what the AI can see, and it uses the code map intersection to ask the AI for the proper files to put in context. (Typically Flash, cheap, relatively fast, and very good) - Then, use a bigger model, GPT 5.4 or Opus 4.6, to do the work. At this point, context is typically between 30k and 80k max. What I’ve found is that this process is surprisingly effective at getting a high-quality response in one shot. It keeps everything focused on what’s needed for the job. Higher precision on the input typically leads to higher precision on the output. That’s still true with AI. For context, 75% of my code is Rust, and the other 25% is TS/CSS for web UI. Anyway, it’s always interesting to learn about different approaches. I’d love to understand the use case where 1M context is really useful.
- firemelt 7mo agowhenever I see post like this i said well yeah, but its too sophiscated to be practical
- adammarples 7mo agoIt's not sophisticated at all, he just uses a model to make some documentation before asking another model to work using the documentation
- jeremychone 7mo agoFair point, but because I spent a year building and refining my custom tool, this is now the reality for all of my AI requests. I prompt, press run, and then I get this flow: dev setup (dev-chat or plan) code-map (incremental 0s 2m for initial) auto-context (~20s to 40s) final AI query (~30s to 2m) For example, just now, in my Rust code (about 60k LOC), I wanted to change the data model and brainstorm with the AI to find the right design, and here is the auto-context it gave me: - Reducing 381 context files ( 1.62 MB) - Now 5 context files ( 27.90 KB) - Reducing 11 knowledge files ( 30.16 KB) - Now 3 knowledge files ( 5.62 KB) The knowledge files are my "rust10x" best practices, and the context files are the source files. (edited to fix formatting)
- ofisboy 7mo agoi think it's buggy. i keep getting "compacting conversation" even though i restarted the cli. and i'm for sure not using 5 times more.
- elophanto_agent 7mo ago[dead]
- genyk1 7mo ago[flagged]
- anshumankmr 7mo agoAll while their usage limits are so excessively shitty that I paid them 50$ just two days back cause I ran out of usage and they still blocked from using it during a critical work week (and did not refund my 50$ despite my emails and requests and route me to s*ty AI bot.). Anyway, I am using Copilot and OpenCode a lot more these days which is much better.
- praddlebus 7mo agoWhat model(s) do you use with OpenCode? Can you use opus4.6 1m? Is it better in terms of usage if you use the same model?
- anshumankmr 7mo agoSonnet4.6/Haiku4.5 for simple stuff.
- ionwake 7mo agoHave we reached the point where its "normal" to mostly use AI to code? Im just wondering because Im sure it was less than a month ago when I said I havent coded manually for over 6 months and I had several comments about how my code must be terrible. Im not butt hurt Im just wondering if the overton window has shifted yet.
- heraldgeezer 7mo agoI feel like I'm the only one here using AI as just a chatbot for research, shopping, advice etc and for one off regex or bash/ps scripts... then again not a programmer so.
- PeterStuer 7mo agoThe thing that would get me more excited is how far they could push context coherence before the model loses track. I'm hoping 250k.
- aplomb1026 7mo ago[dead]
- JulianPembroke 7mo ago[dead]
- kopollo 7mo ago[dead]
- hirehalai 7mo ago[dead]
- MorkMindy74 7mo ago[flagged]
- aneyadeng 7mo ago[dead]
- miohtama 7mo agoI just tested this with Jupytwr Notebooks for a day. LLMs have struggled with them because notebooks contain a lot of token as the data of rendered cells. With Opus 1M, LLM edit was very robust and finally useable
- sporkland 7mo agoCan someone help me with insights about large context models? Are there relationships that pop up at the beginning and end of long context windows that don't transitively follow from intermediate points? Is there value in the training over these longer windows vs using the more basic/closer weight distributions over different sliding windows?
- hnipps 7mo agoWhy would anyone need this much context? Genuine question. It's not worth the drop in quality IMO.
- tommek4077 7mo agoYou sound like: "Why would anyone need more than 640KB RAM?!".
- hnipps 7mo agoThat’s just not comparable. I’ll use your figure: If you use 400KB RAM for a process, using the remaining 240KB for something else doesn’t degrade the performance of the initial process (assuming nothing is trying to use more than the available RAM). Each unit of RAM is independent no? Every token of context you use causes a drop in LLM performance. More RAM == more processes can be supported with no degradation More context != more stuff can be done with no degradation
- Slav_fixflex 7mo agoI've been using Claude Code directly on my production servers to debug complex I/O bottlenecks and database locks. The ability of the latest models to hold the entire project context while suggesting real-time fixes is a game changer for solo founders. It helped me stabilize a security tool I’m building when other agents kept hallucinating.
- TZubiri 7mo agoRemember folks, just because you can use 1m tokens doesn't mean you should
- geminiboy 7mo agoMy companies brand guidelines document was 600 ish pages long and claude desktop couldnt handle it. As soon as I saw the announcement , tried again and created a working design skill that can create design artifacts following the brand guidelines. While these improvements seem incremental, they have a compounding effect on usefulness. My AI doomsday calculator just got decremented by anothet 6 months.
- glimshe 7mo agoIs "generally available" the proper wording when Claude is generally unavailable so often these days?
- tuo-lei 7mo agoMy main frustration with long-context coding sessions isn't just the limit itself, it's that after the fact it's hard to tell which turns actually caused the context to bloat or the session to go off track. It's painful enough I have to build a tool to help myself understand the context/turn data correlation. I have to manual compact now