6 ms·
I'm afraid the music may be slowly fading at this party, and the lights will soon be turned on. We may very well look back on the last couple years as the golde
by geeky4qwerty 6mo ago
I'm afraid the music may be slowly fading at this party, and the lights will soon be turned on. We may very well look back on the last couple years as the golden era of subsidized GenAI compute.
For those not in the Google Gemini/Antigravity sphere, over the last month or so that community has been experiencing nothing short of contempt from Google when attempting to address an apparent bait and switch on quota expectations for their pro and ultra customers (myself included). [1]
While I continue to pay for my Google Pro subscription, probably out of some Stockholm Syndrome, beaten wife level loyalty and false hope that it is just a bug and not Google being Google and self-immolating a good product,
I have since moved to Kiro for my IDE and Codex for my CLI and am as happy as clam with this new setup.
[1] https://github.com/google-gemini/gemini-cli/issues/24937 https://github.com/google-gemini/gemini-cli/issues/24937
- dgellow 6mo agoFor what it’s worth, that was pretty obvious from the get go it wasn’t a realistic long term deal. I’ve been building all the libraries I hoped existed over the past 1-2y to have something neat to work with whenever the free compute era ends. I feel that’s the approach that makes sense. Take the free tokens, build everything you would want to exist if you don’t have access to the service anymore. If it goes away you’re back to enjoying writing code by hand but with all the building blocks you dreamt of. If it never goes away, nothing wasted, you still have cool libs
- apgwoz 6mo agoYes! I’ve been trying (and failing!) to get people to understand this. Build the high leverage tools while the tokens are cheap. Unfortunately, I haven’t figured out the right set of high leverage tools. :)
- 1970-01-01 6mo agoLights on = Ads in your output. EOY latest; they can't keep kicking the massive costs down the road.
- fooster 6mo agoWhere is your evidence of this "massive cost"? Inference is massively profitable for both anthropic and openai. Training is not.
- wesammikhail 6mo agosource?
- KaoruAoiShiho 6mo agoAfter googling https://www.reddit.com/r/singularity/comments/1psesym/openais_compute_margin_said_to_jump_to_70/ https://www.reddit.com/r/singularity/comments/1psesym/openai...
- caminante 6mo ago>OpenAI's compute margin, referring to the share of revenue excluding the costs of running its AI models for paying users Huh? The reddit summary comment makes no sense. How are they getting revenues without ads or paying customers? "After" makes more sense. FTA: >The company has yet to show a profit and is searching for ways to make money to cover its high computing costs and infrastructure plans.
- wesammikhail 6mo agoI've seen sources like this before. It's all hearsay and promo. I was asking for any publicly available verifiable information regarding the cost of inference at scale. I haven't seen any such info personally which is why I asked. I'm dying to see S-1 filing for Anthropic or OpenAI. I don't actually think inference is as cheap as people say if you consider the total cost (hardware, energy, capex, etc)
- KaoruAoiShiho 6mo agoWell they're not public yet so you'll have to put up with rumors. But the numbers are available for companies like DeepSeek say they have an 80% profit margin, so it stands to reason OAI etc would do similar numbers considering they charge much more.
- faangguyindia 6mo agoUltimately we'll find more efficient techniques and hardware and AI companies will end up owning Nuclear Power Stations and continue providing models capable of 10x of what they are now. Valuation have already reached point where these companies can run their nuclear power station, fund developement of new hardware and techniques and boost capabilities of their models by 10x
- croes 6mo agoToo bad the models collapse because the lack of nee good training data. How many companies will generate profit in the end, what will happen with all those power stations and data centers ?
- Root_Denied 6mo agoThere's not enough nuclear to go around, and the approval/permitting process for new nuclear power plants is nothing to sneeze at, both in terms of time and cost. That's also ignoring that nuclear power plants also consume quite a bit of water, which may be a more difficult bottleneck in and of itself even without trying to add nuclear into the mix.
- asdfasgasdgasdg 6mo agoSo, antigravity will definitely quickly eat up your pro quota. You can run out of it in an hour (at least on the $20/mo plan) and then you'll be waiting five days for it to refresh. However, I've found that the flash quota is much more generous. I have been building a trio drive FOC system for the STM32G474 and basically prompting my way through the process. I have yet to be able to run completely out of flash quota in a given five hour time window. It is definitely completing the work a lot faster than I could do myself -- mainly due to its patience with trying different things to get to the bottom of problems. It's not perfect but it's pretty good. You do often have to pop back in and clean up debris left from debugging or attempts that went nowhere, or prompt the AI to do so, but that's a lot easier than figuring things out in the first place as long as you keep up with it. I say this as someone who was really skeptical of AI coding until fairly recently. A friend gave me a tutorial last weekend, basically pointing out that you need to instruct the AI to test everything. Getting hardware-in-loop unit tests up and running was a big turning point for productivity on this project. I also self-wired a bunch of the peripherals on my dev board so that the unit tests could pretend to be connected to real external devices. I think it helps a lot that I've been programming for the last twenty years, so I can sometimes jump in when it looks like the AI is spinning its wheels. But anyway, that's my experience. I'm just using flash and plan mode for everything and not running out of the $20/mo quota, probably getting things done 3x as fast as I could if I were writing everything myself.
- rzkyif 6mo agoFellow annoyed Google AI Pro subscriber here! Can confirm, I initially enjoyed the 5-hour limits on Gemini CLI and Antigravity so much that I paid for a full year, thinking it was a great decision In the following months, they significantly cut the 5-hour limits (not sure if it even exists anymore), introduced the unrealistically bad weekly limit that I can fully consume in 1-2 hour, introduced the monthly AI credits system, and added ads to upgrade to Ultra everywhere At the very least the Gemini mobile app / web app is still kinda useful for project planning and day-to-day use I guess. They also bumped the storage from 2TB to 5TB, but I don't even use that
- stavros 6mo agoIt should be illegal to change the terms of the subscription mid-period. If you paid for the full year, you should get that plan for the whole year. I don't understand how it's ok for corporations to just change the terms mid-way, and we just have to accept it.
- bobmcnamara 6mo agoT&C?
- stavros 6mo agoI'm sure the T&C say something like "you're going to pay us money, and we reserve the right to give you something for it, or maybe nothing, and you should thank us for the privilege".
- bachmeier 6mo ago> It should be illegal to change the terms of the subscription mid-period Unfortunately, at least for those of us in the US, there isn't legally much that can be done. It's simply not possible to make a contract that would obligate a company to fulfill its promises on this type of sale.
- nprateem 6mo agoDon't bother upgrading to ultra. It's also now easy to burn all your credits where in Jan it was almost impossible
- palata 6mo ago> We may very well look back on the last couple years as the golden era of subsidized GenAI compute. Looks like enshittification on steroids, honestly.
- omosubi 6mo agoGetting $5000 worth of product essentially free and then being told to pay is not enshittification.
- zzzoom 6mo agoIt's predatory pricing.
- Chaosvex 6mo agoAnother take: perhaps they shouldn't have been pricing it at that point if they weren't capable of actually delivering.
- quikoa 6mo agoThe cost for AI companies might be $5000 but the "essentially free" could be close to the limit of what people are willing to spend. If that's the case then enshittification will continue and/or many AI companies will never be profitable.
- tvbusy 6mo agoWe have seen this before. Companies using VC money to take over the market and then increase prices. In the end, we're worse off without these scumbags but some will still sing that we got free service do it's bot enshitification.
- knollimar 6mo agoIt absolutely is. Loss leading is their fault and anticompetitive.
- byzantinegene 6mo agoit's not worth $5000 if people are not willing to pay that amount for it
- rr808 6mo agoI still remember those $3 uber rides.
- elephanlemon 6mo agoIMO we are currently in the ENIAC era of LLMs. Perhaps there will be a brief moment where things get worse, but long term the cost of these things will go way down.
- pier25 6mo agoCost will probably go down but nobody knows when or how. It might take 10 years for all we know as training costs have only been rising. A huge difference is early computers were not subsidized. It took decades until most people could afford to own a computer at home.
- croes 6mo agoOr we are in the early Netflix era where profit wasn’t as important as customer growth.
- pbmonster 6mo agoI assume the "briefly gets worse" is when a buch of hyperscalers do a complete write-off of their entire AI investments, bankrupting several of them (which, in turn, bankrupts several large banks and most current venture capital firms)? Cumulative AI capex will hit $2T this year. Cumulative opex is on the same order. Unless the models get real good (as in: can fully replace many engineers) right quick, nobody is even going to see interest getting paid on those investments. The only alternative is model access costing 5 figures per (replaced) seat. But yes, once GPU racks can be had at auction for pennies on the dollar, inference of open source models might be an... OK low margin commodity business.
- alecco 6mo ago> I'm afraid the music may be slowly fading at this party, and the lights will soon be turned on. We may very well look back on the last couple years as the golden era of subsidized GenAI compute. Indeed. Anthropic is just leading the pack switching to juicy corporate users who are happy to pay thousands per month per dev and leave the fans behind. And now OpenAI is following suit. They lowered significantly the limits for the Plus $20 plan and answered concerns with vague confusing tweets about promotions. All this is pushed by the fastest rising demand (Codex growing +50% monthly) while having a serious bottleneck building data centers and getting parts (permits, energy, memory, flash, etc). Users on reddit and Discord are trying to switch to open models or Chinese alternatives. But there's no real replacement.
- muyuu 6mo agoI don't know about users on reddit and discord, but the open models are essentially at SotA with a 3-4 months delay. That puts a hard backstop at what OpenAI and Anthropic can do before I personally can cut them off entirely without losing too much. Granted the experience can be worse, esp. if you're using it very hands-off and not like a junior assistant who's extremely fast but doesn't know what he's doing at the architecture and strategy level. But even for that I'm relatively confident the Chinese will be competitive pretty soon, and they won't be too expensive. And we know this because we can see their current models and we know what it takes to run them. Currently my Strix Halo computer that costed me under £3k can do a lot of LLM stuff that is perfectly useful. In some ways, it's better than "cloud" models, I have models that essentially don't say "no" and I have relatively predictable setups. If you want to get fancy, you can right now rent compute to run models that are extremely capable like the latest ones from Kimi, GLM, Qwen, Minimax at full size from providers that are not operating at a loss and it won't be too expensive. You can pool resources to do the same locally. You can do stuff that cloud providers are unlikely to market, like distillation and abliteration to serve your specific needs. I'm very optimistic about open weights models just the way they are right now. But I agree with you that OpenAI will likely play similar games to Anthropic and it could be soon.
- rachel_rig 6mo ago[dead]
- hacker_homie 6mo agoMaybe I missed the party, but it feels like it's just starting. I have only been running local models and we are finally at the point with gemma4 and Qwen3.5 where they can start doing coding work. And the quota can't change.
- ainiriand 6mo agoThe only viable future-proof solution to this hellscape is what you mention, local models and/or corporate models for work.
- Gareth321 6mo agoI am surprisingly optimistic about local LLMs. Their progress (especially with regards to distillation) over the last year has been remarkable. Qwen 3.5 is amazing for what it is. It think it's production capable - for many use cases, but not all. It does require more careful alignment of instructions, and offers a smaller context (even with very large unified memory). But with some care, one can code all day, every day, without limits. The Mac Mini 64GB is probably sufficient for Qwen 3.5 35B. Go larger for larger contexts. Of course it's not as easy as pointing February Opus 4.6 at a folder and giving it one-sentence instructions.