13 ms·
The Pro plan quota seems to be getting worse. I can get maybe 20-30 minutes work done before I hit my 4 hour quota. I found myself using it more just for the pl
by laserDinosaur 8mo ago
The Pro plan quota seems to be getting worse. I can get maybe 20-30 minutes work done before I hit my 4 hour quota. I found myself using it more just for the planning phase to get a little bit more time out of it, but yesterday I managed to ask it ONE question in plan mode (from a fresh quota window), and while it was thinking it ran out of quota. I'm assuming it probably pulled in a ton of references from my project automatically and blew out the token count. I find I get good answers from it when it does work, but it's getting very annoying to use.
(on the flip side, Codex seems like it's being SO efficient with the tokens it can be hard to understand its answers sometimes, it rarely includes files without you doing it manually, and often takes quite a few attempts to get the right answer because it's so strict what it's doing each iteration. But I never run out of quota!)
- stareatgoats 8mo agoClaude Code allegedly auto-includes the currently active file and often all visible tabs and sometimes neighboring files it thinks are 'related' - on every prompt. The advice I got when scouring the internets was primarily to close everything except the file you’re editing and maybe one reference file (before asking Claude anything). For added effect add something like 'Only use the currently open file. Do not read or reference any other files' to the prompt. I don't have any hard facts to back this up, but I'm sure going to try it myself tomorrow (when my weekly cap is lifted ...).
- sigseg1v 8mo agoWhat does "all visible tabs" mean in the context of Claude Code in a terminal window? Are you saying it's reading other terminals open on the system? Also how do you determine "currently active file"? It just greps files as needed.
- adobrawy 8mo agoYou can install VSCode extension and use "/ide" to connect them.
- withinboredom 8mo agoDo people actually use this mode? Having to approve diffs in the ide is too annoying.
- solumunus 8mo agoDepends on my task. If it’s complex and my expectation is for Claude to get things wrong the diff preview is helpful.
- vidarh 8mo agoEven then, I'd wait until it's had a chance to iterate and correct itself in a loop before I'd even consider looking at the output, or I end up babysitting it to prevent it from making mistakes it'd often recognise and fix itself if given the chance.
- solumunus 8mo agoTrue. I’ve been strictly in the terminal for weeks and I have a stop hook which commits each iteration after successful rust compilation and frontend typechecks, then I have a small command line tool to quickly review last commit. It’s a pretty good flow!
- HumanOstrich 8mo agoYou can tell it not to do that and it will show inline diffs.
- DANmode 8mo ago13 days ago on HN: https://news.ycombinator.com/item?id=46566292 https://news.ycombinator.com/item?id=46566292
- idonotknowwhy 8mo agoYes, it does exactly that. It also sends other prompts like generating 3 options to choose from, prefilling a reply like 'compile the code', etc. (I can confirm this because I connect CC to llama.cpp and use it with GLM-4.7. I see all these requests/prompts in the llama-server verbose log.) You can stop most of this with export DISABLE_NON_ESSENTIAL_MODEL_CALLS=1 And might as well disable telemetry, etc: export CLAUDE_CODE_DISABLE_NONESSENTIAL_TRAFFIC=1 I also noticed every time you start CC, it sends off > 10k tokens preparing the different agents. So try not to close / re-open it too often. source: https://code.claude.com/docs/en/settings https://code.claude.com/docs/en/settings
- JLCarveth 8mo agoI would always close claude to start a new chat... Guess I should stop doing that. Thanks for bringing my attention to those two env vars.
- aanet 8mo ago^ THIS I've run out of quota on my Pro plan so many times in the past 2-3 weeks. This seems to be a recent occurrence. And I'm not even that active. Just one project, execute in Plan > Develop > Test mode, just one terminal. That's it. I keep getting a quota reset every few hours. What's happening @Anthropic ?? Anybody here who can answer??
- genewitch 8mo agosounds like the "thinking tokens" are a mechanism to extract more money from users?
- mystraline 8mo agoIts the clanker version of the "Check Wallet Light" (check engine light).
- vunderba 8mo agoAnecdotally but it definitely feels like in the last couple weeks CC tends to be more aggressive at pulling in significantly larger chunks of an existing code base - even for some simple queries I'll see it easily ramp up to 50-60k token usage.
- genewitch 8mo agoI'm curious if anyone has logged the number of thinking tokens over time. My implication was the "thinking/reasoning" modes are a way for LLM providers to put their thumb on the scale for how much the service costs. they get to see (if not opted-out) your context, idea, source code, etc. and in return you give them $220 and they give you back "out of tokens"
- throwup238 8mo ago> My implication was the "thinking/reasoning" modes are a way for LLM providers to put their thumb on the scale for how much the service costs. It's also a way to improve performance on the things their customers care about. I'm not paying Anthropic more than I do for car insurance every month because I want to pinch ~~pennies~~ tokens, I do it because I can finally offload a ton of tedious work on Opus 4.5 without hand holding it and reviewing every line. The subscription is already such a great value over paying by the token, they've got plenty of space to find the right balance.
- ChicagoDave 8mo agoI never run out of this mysterious quota thing. I close Claude Code at 10% context and restart. I work for hours and it never says anything. No clue why you’re hitting this. $230 pro max.
- croes 8mo agoPro is 20x less than Max
- yjtpesesu2 8mo agoAny clue why you might be a favored/favoured high value user?
- 0xack 8mo agoThe entire conversation is fed in as context effectively compounding your token usage over the course of a session. Sessions are most efficient when used for one task only.
- ChicagoDave 8mo agoI get a decent amount of work in before restarts.
- fluidcruft 8mo agoDoes closing claude code do something that running /clear does not?
- idonotknowwhy 8mo agoYeah, it re-sends all the agent system prompts.
- nwatson 8mo agoSelf-hosted might be the way to go soon. I'm getting 2x Olares One boxes, each with an RTX 5090 GPU (NVIDIA 24GB VRAM), and a built-in ecosystem of AI apps, many of which should be useful, and Kubernetes + Docker will let me deploy whatever else I want. Presumably I will manage to host a good coding model and use Claude Code as the framework (or some other). There will be many good options out there soon.
- NitpickLawyer 8mo agoI've been using local LLMs since before chatgpt launched (gpt-j, gpt-neox for those that remember), and have tried all the promising models as they launch. While things are improving faster than I thought ~3 years ago, we're still not there in terms of 1-1 comparison with the SotA models. For "consumer" local at least. The best you can get today with consumer hardware is something like devstral2-small(24B) or qwen-coder30b(underwhelming) or glm-4.7-flash (promising but buggy atm). And you'll still need beefy workstations ~5-10k. If you want open-SotA you have to get hardware worth 80-100k to run the big boys (dsv3.2, glm4.7, minimax2.1, devstral2-123b, etc). It's ok for small office setups, but out of range for most local deployments (esp considering that the workstations need lots of power if you go 8x GPUs, even with something like 8x 6000pro @ 300w).
- asno3030 8mo agoIt's amazing how inefficient and cost prohibitive this tech is currently both for consumers, and large host. Yes I like my crack at 20, bucks a month, but the tech is going to need to improve quick to keep that up.
- behnamoh 8mo ago> Self-hosted might be the way to go soon. As someone with 2x RTX Pro 6000 and a 512GB M3 Ultra, I have yet to find these machines usable for "agentic" tasks. Sure, they can be great chat bots, but agentic work involves huge context sent to the system. That already rules out the Mac Studio because it lacks tensor cores and it's painfully slow to process even relatively large CLAUDE.md files, let alone a big project. The RTX setup is much faster but can only support models ≤192GB, which severely limits its capabilities as you're limited to low Q GLM 4.7, GLM 4.7 Flash/Air/ GPT OSS 120b, etc.
- thunfischtoast 8mo agoI've used the Anthropic models mostly through Openrouter using aider. With so much buzz around Claude Code I wantes to try it out and thought that a subscription might be more cost efficient for me. I was kinda disappointed by how quickly I hit the quota limit. Claude Code gives me a lot more freedom than what aider can do, on the other side I have the feeling that pure coding tasks work better through aider or Roo Code. The API version is also much much faster that the subscription one.
- aja12 8mo agoBeing in the same boat as you I switched to OpenCode with z.ai GLM 4.7 Pro plan and it's quite ok. Not as smart as Opus but smart enough for my needs, and the pricing is unbeatable
- davidwritesbugs 8mo agoDitto. It is very very slow but I never hit quota limits but people on Discord are complaining like mad it is slow even on the Pro plans. I tend to use glm-*air a lot for planning before using 4.7
- thunfischtoast 8mo agoI've also see OpenCode around, but have yet to try it. I wonder how it compares to Roo Code
- rasmus1610 8mo agoVery happy to see that I am not the only one. My pro subscription lasts maybe 30 minutes for the 5 hour limit. It is completely unusable and that's why I actually switched to OpenCode + GLM 4.7 for my personal projects and. It's not as clever as Opus 4.5 but it often gets the job done anyway