7 ms·
> That forces people to look beyond the walled garden. Every time I get "you used your quota, come back in 3 hours, or 2 days" -> that is experimentation time
by visarga 2mo ago
> That forces people to look beyond the walled garden.
Every time I get "you used your quota, come back in 3 hours, or 2 days" -> that is experimentation time with their competition, leading to changed service plans. When they said "claude -p" will be billed at API pricing even for plan users I moved my harness off claude. After I integrated codex, then it was never going to be a full claude project again.
What business encourages users to try their competition and adapt their usage to the competing products?
- LaurensBER 2mo ago>™Every time I get "you used your quota, come back in 3 hours, or 2 days" -> that is experimentation time with their competition, leading to changed service plans. I guess this is why they're pushing Claude code hard (not supporting agents.md, not allowing third party harnesses, etc) but when switching to another provider is as easy as opening a new terminal and typing omp/pi/codex your moat is effectively zero. They can compete on price, quality or value but anything else is just madness. Currently they (arguably) own quality but this won't last.
- zymhan 2mo agoIt drove me to setup Qwen 3.8 this weekend. I couldn't see the value in just giving them money for a higher tier plan instead. I've never run a local LLM model before. Certainly won't take as long to iterate on this.
- dexterlagan 2mo agoQwen 3.8 is excellent. With the right harness, it does about 95% of what Opus can do, in my case automation software development. Since 3.8 came out, I have significantly revised my expectations for a local model. Give it another year or two, and we'll be running fast and free local models for nearly everything that matters, and these costly subscriptions will be a thing of the past. I've always believed that AI should be free for everybody, like TV and radio. We're almost there.
- mrtsepelev 2mo agoWhat harness would you recommend? I’ve tried Pi but the model struggled to stay on track after the compaction. I have only 48gb of ram, so can fit only 80k context max, so good compaction is must.
- Scaled 2mo agoNot op, but check out open code; you can turn on K/V quantization to help with increasing context if you have not already. I think K needs to stay at least 8 but I hear V can go down to 4?
- npodbielski 1mo agoI am running it on 32GB and I did not saw model loosing it context even after 4-5 compactions in pi. I am running sessions for few days sometimes. I think it looped once, but loop police extension stopped it. The only problem I have know is how pi compaction works, which is forcing full prefill which takes time and it is erroring a lot. I wrote my own compaction that should remove full prefil but it does not work. But this is the only problem with this setup and it is more problem with pi then the model. I much more prefer it to use Qwen then paid models: Claude forces me to do reauth every other day and codex models either are too costly or not capable enough.
- wccrawford 2mo agoI'd also love to hear your setup? How much VRAM/RAM, I assume Qwen 3.8 27b, what harness, are you using any particular skill set?
- c16 2mo agoI've a 32gb and 64gb (work) MBP. 32 works - just and sits at around 28/29gb of 32. 64 works great, so the 48gb laptop with MLX + MTP should be fine. I'm using Ollama. I initially used the Claude Code harness on 3.6 A3B, but found that tooling would break as Claude released new versions and things would go weird. I've since written my own harness which has basic operations: read, find, bash (which can write files, python etc...) & web_fetch, all within a mac container. Works amazing. You don't need anything complicated to go very far. Low hanging fruit would be Pi or OpenCode. If you really want a much better understanding of what your hardware is capable of then give writing your own a go. Additional tip: Low Power mode reduces some token speed, but stops the laptop over heating and the fans going crazy.
- walthamstow 2mo agoThe claude -p thing was doubly stupid because they quietly allowed it again a couple of weeks later, so they pissed off developers for nothing
- airspresso 2mo agoWait, it's allowed again? Completely missed that. Been avoiding to use it and trying to find workarounds, not great.
- walthamstow 2mo agoI can't find the page now but yes they quietly "paused" the June 15th rollout of API pricing for -p headless. Presumably to come back again one day.
- smoe 2mo agoI think they didn't even bother making a separate post about their backpedaling, they just slapped some disclaimers onto the existing page: https://support.claude.com/en/articles/15036540-use-the-claude-agent-sdk-with-your-claude-plan https://support.claude.com/en/articles/15036540-use-the-clau... Update June 15: We're pausing the changes to Claude Agent SDK usage described below. For now, nothing has changed: Claude Agent SDK, claude -p, and third-party app usage still draw from your subscription's usage limits. The previously announced monthly credit, which would have been available to eligible claimants in connection with these changes, isn't available. We’re working to update the plan to better support how users build with Claude subscriptions. When we have an update, we'll share it before anything takes effect.
- PufPufPuf 2mo agoYes, they sent an email about it. You can also use the Claude SDK with an OAuth token ("claude setup-token" output) and it counts against the regular limit. Maybe they were afraid of losing users dependent on ACP (Zed editor and other compatible tools), since Claude Code does not have native ACP support and integrates only through the SDK?
- dan_ggggg 2mo ago> What business encourages users to try their competition and adapt their usage to the competing products? They are high on their own supply. The people running these companies are delusional imbeciles who have been placed in charge of billions of dollars.
- lelanthran 2mo ago> What business encourages users to try their competition and adapt their usage to the competing products? If you are selling something, and losing $10 on each sale, you also would want to limit how much you sell. I mean, sure, you are losing money on each sale so you can landgrab, but you still have to balance the land-grabbing with how much money you can actually lose.