3 ms·
For the past month, I've been claiming that $20/mo codex is the best deal in AI. Now I'm going to have to find the new best deal.
by __mharrison__ 6mo ago
For the past month, I've been claiming that $20/mo codex is the best deal in AI.
Now I'm going to have to find the new best deal.
- verdverm 6mo agoWe are exiting a hype cycle, well into the adoption curve. Subscriptions were never going to last. My next step is going to be evaluating open and local models to see if they are sufficiently close to par with frontier models. My hope is that the end of seat based pricing comes with this tech cycle. I was looking for document signing provider that doesn't charge a monthly, I only need a few docs a year.
- __mharrison__ 6mo agoI recently experimented creating a Python library from scratch with Codex. After I was done, I took the PRD and Task list that was generated and fed them to opencode with Qwen 3.5 running locally. Opencode was able to create the library as well. It just took about 2x longer.
- selectodude 6mo agoWhich version of Qwen 3.5 did you use?
- verdverm 6mo agowhich quant as well
- __mharrison__ 6mo agoNot at my computer now, either 27 or 35b not quantized. Next week I will be trying qwopus 27b.
- alifeinbinary 6mo agoI'm developing software in this area right now, so I try a lot of the new models. They're not even close for coding tasks. It basically comes down to 26b parameters vs 1T parameters / quantisation / smaller context sizs, there's no comparison. However, for agentic work, tool calling, text summarisation, local LLMs can be quite capable. Workloads that run as background tasks where you're not concerned about TTFB, cold starts, tok/s etc., this is where local AI is useful. If you have an M processor then I would recommend that you ditch Ollama because it performs slowly. We get double or triple tok/s using omlx or vmlx, respectively, but vmlx doesn't have extensive support for some models like gpt-oss.
- AstroBen 6mo agoKimi K2.5 (as an example) is an open model with 1T params. I don't see a reason it has to be local for most use cases- the fact that it's open is what's important.
- Art9681 6mo agoThat is just idealism. Being "open" doesnt get you any advantage in the real world. You're not going to meaningfully compete in the new economy using "lesser" models. The economy does not care about principles or ethics. No one is going to build a long term business that provides actual value on open models. They can try. They can hype. And they can swindle and grift and scalp some profit before they become irrelevant. But it will not last. Why? Because what was built with an open model can be sneezed into existence by a frontier model ran via first party API with the best practice configurations the providers publish in usage guides that no one seems to know exist. The difference between the best frontier model (gpt-5.4-xhigh or opus 4.6) and the best open model is vast. But that is only obvious when your use case is actually pushing the frontier. If you're building a crud app, or the modern equivalent of a TODO app, even a lemon can produce that nowadays so you will assume open has caught up to closed because your use case never required frontier intelligence.
- adrian_b 6mo agoA model with open weights gives you a huge advantage in the real world. You can run it on your own hardware, with perfectly predictable costs and predictable quality, without having to worry about how many tokens you use, or whether your subscription limits will be reached in the most inconvenient moment, forcing you to wait until they will be reset, or whether the token price will be increased, or your subscription limits will be decreased, or whether your AI provider will switch the model with a worse one, and so on. Moreover, no matter how good a "frontier model" may be, it can still produce worse results than a worse model when the programmer who manages it does not also have "frontier intelligence". When liberated of the constraints of a paid API, you may be able to use an AI coding assistant in much more efficient ways, exactly like when the time-sharing access to powerful mainframes has been replaced with the unconstrained use of personal computers. When I was very young I have passed through the transition from using remotely a mainframe to using my own computer. I certainly do not want to return to that straitjacket style of work.
- piyh 6mo agoAlready paying for Google photo storage, AI pro for an extra $7 is a steal with anti-gravity.
- matt_heimer 6mo agoThat's only good for the web based UI. If you want Gemini API access which is what this article is about then you must go the AIStudio route and pricing is API usage based. It does have a free usage tier and new signups can get $300 in free credits for the paid tier so it's I think it's still a good deal, just not as good as using the subscriptions would be.
- spijdar 6mo agoNo? Isn't the article about Codex, which is roughly equivalent to "Gemini CLI" and Google's Antigravity? Google's subscriptions include quotas for both of those, albeit the $20 monthly "Pro" plan has had its "Pro" model quota slashed in the last few weeks. You still get a large number of "Gemini 3 Flash" queries, which has been good enough for the projects I've toyed with in Antigravity.
- matt_heimer 6mo agoI guess that's true but I find Google's models better than their public tooling. The Pro subscription includes "Gemini Code Assist and Gemini CLI" but the Gemini Code Assist plugin for IntelliJ which is my daily driver is broken most of the time to the degree that it's completely unusable. Sometimes you can't even type in the input box. The only way I can do serious development with Gemini models is with other tooling (Cline, etc) that requires API based access which isn't available as part of the subscription.
- bethekind 6mo agoI agree. Gemini models are held back by their segmentation of usage between multiple products, combined with their awful harnesses and tooling. Gemini cli, antigravity, Gemini code assist, Jules.... The list goes on. Each of these products has only a small limit and they must share usage. It gets worse than that though. Most harnesses that are made to handle codex and Claude cannot handle Gemini 3.1 correctly. Google has trained Gemini 3.1 to return different json keys than most harnesses expect resulting in awful results and failure. (Based on me perusing multiple harness GitHub issues after Gemini 3.1 came out)
- scosman 6mo agoCheck out z.ai coder plan. The $27/mo plan is roughly the same usage as the 20x $200 Claude plan. I have both and Claude is a little better, but GLM 5.1 is much better value.
- rustyhancock 6mo agoAgreed, I use Z.ai and the usage is fantastic the only temper that recommendation that it's often unreliable. Perhaps a few times per week it's unresponsive. Maybe more often it seems to become flakey. It's very variable though recently I'm noticing it's more reliable but there was a patch where it was nearly unusable some days. I guess I won't complain for the price and YMMV.
- scosman 6mo agoAgreed. They had a rough patch around the 4.7 to 5 upgrade. New architecture required hardware migration. The 5 to 5.1 upgrade was much smoother (same architecture new weights). As you say, little rough around edges, but still great value. Trick I learned is that it's max 2 parallel requests per user. You can put a billion tokens a month through it, but need to manage your parallelism.
- mickeyp 6mo agoIf you're ok with a model provider that goes down all the time and has such a poor inference engine setup that once you get past 50k tokens you're going to get stuck in endless reasoning loops.
- aulin 6mo agoGH Copilot is still the best deal, while it lasts
- __mharrison__ 6mo agoYeah, it's really good. Probably going to be the next best deal until they cut back. I need to try the command line version.
- aulin 6mo ago> I need to try the command line version. Is there any other?
- hokkos 6mo agoI feel they will go token base at some point, currently if you only use it with precise prompts and not random suggestions, switch between models 5.4 and 5.4 mini depending on the work, it is the best deal.
- muyuu 6mo agoWhat has actually changed? It's unclear how much can you do right now, unless they've already switched you to the new plan and you're speaking from experience.