4 ms·
I don’t get how Claude is considered providing “unlimited” quotas. I use up my 5h on Max $100 and Team Premium in 2-3h of relaxed use of Opus 5 high+ on fresh s
by jmaker 1mo ago
I don’t get how Claude is considered providing “unlimited” quotas. I use up my 5h on Max $100 and Team Premium in 2-3h of relaxed use of Opus 5 high+ on fresh sessions with just a couple skills/plugins. And my weekly quotas are gone in 3 days of such relaxed use. With Codex my $100 weekly quota is used up within 3 days with Sol high+ too.
Im not convinced to pay $200 for Claude’s models.
With Claude, I have to intervene every 15-20 minutes, it’s non-autonomous and it’s incredibly unreliable at self-correction. GPT is strong at self-correction but it tends to drift away from the plan to self-correct in a loop very often - a lot of tokens and time burnt on aimless churn. Opus tends to push its uninformed opinions and fake retrieval, drifting every turn increasingly farther from the intended and approved design. Opus skims over specs and makes too many mistakes.
As for closed frontier models, I prefer the GPT models over Claude’s.
I’ve started relying more on Grok, GLM, Kimi and DeepSeek models for subagents - I’ve ended up with a factory and am seeking to reduce my reliance on the closed frontier models - they’re just not SoTA on their own for development anymore.
- greenavocado 1mo agoI'd start by disabling all plugins and MCPs. You'd be surprised how quickly those can annihilate usage and cache coherency
- zarzavat 1mo agoYou may be causing a lot of cache misses. You have to use the caching efficiently otherwise you can burn up any plan in any amount of time.
- oblio 1mo agoHow do you use the cache efficiently?
- eulo__ 1mo agoI’m currently running /compact “explain the core next step” whenever over 20% context. Also doing /clear with a md file handover if I think the next input is diverse enough from the previous work. You get a lot more out of it. Not sure if this is best practice though.
- zarzavat 1mo agoYou have to keep your session warm in cache. Keep the AI talking/thinking. If you have not touched a session for few minutes then /clear and start a new session. Providers will generally keep your session in cache for at least 5 minutes, possibly hours. The exact cache policy depends on the provider. If your session expires from cache then the next time you send a message you will have to pay for all the tokens you had used in context up until that point again. e.g. if you have 200k tokens in context then if your session goes cold and you send a message after expiry you will have to pay for those 200k tokens again. With 1M contexts especially you have to be extremely careful that you don't end up resubmitting requests for hundreds of thousands of tokens again and again. Try to get yourself and the model to use disk for medium-term context rather than model context, that way it's much easier to /clear and restart if you need to go to the bathroom or something.