13 ms·
The era of subsidised inference is truly ending. The new model multipliers (https://docs.github.com/en/copilot/reference/copilot-billing/models-and-pricing#mode
by my002 5mo ago
The era of subsidised inference is truly ending. The new model multipliers (https://docs.github.com/en/copilot/reference/copilot-billing/models-and-pricing#model-multipliers-for-annual-copilot-pro-and-copilot-pro-subscribers https://docs.github.com/en/copilot/reference/copilot-billing...) seem like a huge leap, though. From 1x to 6x for new-ish GPT and Sonnet models. 27x for Opus...
Seems like folks would be better off with OpenRouter instead.
- ItsClo688 5mo ago27x for Opus is genuinely shocking. at that point you're not paying for convenience anymore, you're just paying a GitHub tax. OpenRouter or direct API makes way more sense unless you're really glued to the IDE integration.
- thrdbndndn 5mo agoI keep seeing people mention OpenRouter. Does it effectively bypass regional restrictions for you, so you can use something like the Claude API from unsupported regions such as Hong Kong, or does it still enforce the official providers' geo-restrictions?
- rvnx 5mo agoOpenRouter is great for budget control, but as they are indirect APIs, your experience with cached tokens may vary, eventually costing much more than in direct depending on the providers. You can pay with crypto though, which seems to be convenient for people under sanctions or with limited access, or if you are in low-tax jurisdiction (e.g. HK)
- jauntywundrkind 5mo agoCaching is advertised per model+provider. That said I think few people using openrouter are actually being selective about providers. It took half a day to get my opencode setup, was not friendly. A lot of manually cross referencing model and providers. I was actually mainly optimizing for relatively fast providers. It all is super fragile and I'm sure half out of date; I have no idea if these picks are still fast, no promises they are still the same price (pretty terrifying honestly). I'm mostly on coding plans so it doesn't super affect me. But man is it a bother to maintain.
- schneehertz 5mo agoEven when using OpenRouter in Hong Kong, it is still not possible to connect to region-restricted models like Gemini
- minimaxir 5mo agoWhat's annoying is that it's obvious. In the case of GPT 5.5, if Copilot is going to charge 7.5x what GPT 5.4 costs while OpenAI themselves via the API/Codex only charges 2x of what GPT 5.4 costs, that will immediately raise an eyebrow.
- boothby 5mo agoTo anybody who's been watching the tech sector with a critical eye for pretty much any period from the late 90s and onward, this is just the enshittification process. For most of OpenAI's existence it's been obvious, to me, that investors were burning insane levels of capital to build the market, and now that folks are locked in, you're seeing higher fees, ads, etc. Yet again, the user is the product; the investors want to siphon your data, attention and once you're hooked, money. And for companies like Microsoft and Apple, those hooks can dig deep.
- Incipient 5mo agoI'd call it a straight up "bait and switch".
- BearOso 5mo agoIf you paid attention to the power requirements and amount of hardware being put into data centers, you should have realized that it cost them an order of magnitude more than you were being charged. To rework your analogy: they hooked you, now they're gonna see if they can reel you in.
- AntiUSAbah 5mo agoThey can only reel you in if its worth it. I still can code. And while i do not spend 200$ privat, in my startup we discussed this and our current mental model is, that instead of hiring someone new, we prefer to have more money for tokens. This is easier for us and has a bigger benefit. The cost of a new / first employee is very high, a 200$ subscription is not. Upgrading that to lets say 400 or 800$ is still alot easier and if i can run multiply and better agents with that money, lets goooo.
- specproc 5mo agoYeah, totally. The recent pricing changes have just made my Copilot subscription go from great deal to awful value over night. I've been wanting to get off MS more generally and this is good motivation. Will be playing round with OR this week.
- cedws 5mo agoJust be aware OpenRouter charges a 5.5% fee, I didn’t know until recently. I like the product, and I think the fee is fair, but if you want the absolute best pricing then go direct.
- ffsm8 5mo agoBut with open router you can always just use the latest model. If you're committed to eg Claude opus then you're better off going directly to anthropic for sure, but if not, varying other models may be fine too, depending on use case and be massively cheaper. Eg new deep seek model with same mio context window or Kimi k2.6 with 270k context window for subagents which implement
- gruez 5mo ago>but if not, varying other models may be fine too, depending on use case and be massively cheaper Do inference providers have standardized endpoints, or at least endpoints compatible with claude code? Otherwise to pay 5.5% on all your tokens just so it's slightly easier to swap providers (ie. changing a few urls?)
- swiftcoder 5mo ago> Do inference providers have standardized endpoints, or at least endpoints compatible with claude code? Yep, you can plug deepseek/kimi/minimax into claude code just fine. Or run everything through another harness like opencode instead.
- AntiUSAbah 5mo agoWow thats a lot for routing traffic.
- nacs 5mo agoEven Sonnet 4.6 is 9x multiplier (previously 1x)! The only model I even used on Copilot was Sonnet and now its got a ridiculous multiplier. At this point they might as well just charge per Million tokens like every other provider instead of having a subscription.
- altmanaltman 5mo ago> At this point they might as well just charge per Million tokens like every other provider instead of having a subscription. Pretty sure that's what they will eventually do
- tjoff 5mo ago... that is exactly what they will do. Just click the link in this thread, or read the headline.
- hrpnk 5mo agoWhy the multipliers then at all?
- lexone 5mo agoThe multipliers are there only for current annual plan customers. After 2026 its all tokens.
- MattBDev 5mo agoI thought I was smart for buying the annual plan after I graduated and lost my student plan and then GitHub taking away my Copilot Pro I got for free for being a author of a popular OSS project. Turns out I'm being punished for making that year commitment to them. I like to think I'm only a moderate user of GHCP so this is just terrible for me. I'm honestly thinking about cancelling and switching to alternatives while also looking at investing in a local LLM setup.
- 5mo ago
- rvnx 5mo agoOne theory of the play of SpaceX might do if everyone migrates to query-based billing: Provide cheap and unlimited access to Grok for programmers (hence the Cursor partnership/purchase for distribution). -> This would drag massive revenue right before the IPO announcement, like if the company is super growing -> At a loss, but don't worry, we need these funds to build the biggest datacenter of the universe. This announcement would create enough momentum to increase valuation, and because of the merge of his companies, would save his X/Twitter investors from a tragedy. -> Would also be a great service to Cursor investors and so, who are stuck with their VSCode fork
- minimaxir 5mo agoIt takes longer to build a datacenter with that much capacity than it does for the market to respond.
- 2ndorderthought 5mo agoBuying real estate in imaginary places is lucrative at first
- gigiogigione 5mo agoI don’t get the SpaceX reference. I thought they made rockets?
- deleted 5mo ago[deleted]
- vizzier 5mo agoThey now also own xAI
- 0xffff2 5mo agoWhich in turn owns Twitter. SpaceX is now a social media company in addition to a rocket company. One theory I think Matt Levine posited, is that SpaceX will go public with dual-class stock that gives Elon control even with a minority ownership stake, and will subsequently buy Tesla, which doesn't have dual class stock, making SpaceX the singular "Elon Musk company", with him having operational control despite being public.
- giwook 5mo agoLots of us have noticed that usage limits for Claude have been nerfed in recent weeks/months. If anything, these new multipliers are more transparent than anything OpenAI or Anthropic have communicated regarding actual costs and give us a more realistic understanding of what it's costing these providers. The fact that we were able to get such a substantial amount of usage for $20/$100/$200 a month was never meant to last and to think otherwise was perhaps a bit naive. This feels like a strategy from the ZIRP era of tech growth where companies burned investor capital and gave away their products and services for free (or subsidized them heavily) in order to prioritize user acquisition initially. Then once they'd gained enough traction and stickiness they'd then implement a monetization strategy to capitalize on said user base.
- dualvariable 5mo agoHowever, inference costs for entirely good enough models are likely to keep declining in the future. We're probably hitting diminishing returns on model size and training. The new generations aren't quantum leaps anymore, and newer generations of open source models like DeepSeek are likely to start getting good enough. There's going to be a limit to how much they can raise prices, because someone can always build out a datacenter and fill it up with open source DeepSeek inference and undercut your prices by 10x while still making a very good ROI--and that's a business model right there. Right now I'm sure there's a lot of people who will protest that they couldn't do their jobs with lesser models, but as time goes on that will get less and less. Already right now the consumers who are using AI for writing presentations, cooking recipe generation and ELI5 answers for common things, aren't going to be missing much from a lesser model. That'll actually only start to get cheaper over time. Also for business needs, as AI inference costs escalate there comes a point where businesses rediscover human intelligence again, and start hiring/training people to do more work to use lesser models--if that is more productive in the end than shelling out large amounts of cash for inference on the latest models. [Although given how much companies waste on AWS, there's a lot of tolerance for overspending in corporations...]
- Fire-Dragon-DoL 5mo agoI hope it's true, but right now hardware prices are insane
- whateveracct 5mo ago"eras" tend to not be so short lol
- Mattwmaster58 5mo agoFYI, these are the multipliers for annual plan. I would hazard a guess most people are not on an annual plan
- mitjam 5mo agoI am and I see it as stopping the music at a party when you want everyone to go home without telling them to go home. There is also the offer to quit with prorated refund for the remaining time. I think I am going to take it.
- skeeter2020 5mo ago"This change aligns Copilot pricing with actual usage and is an important step toward a sustainable, reliable Copilot business and experience for all users." I see statements like this as strong indicators that the sales people are wrapping up their work and the accountants are taking over. The land rush is switching to an operational efficiency play.
- fsniper 5mo agoAnd enshitification starts.
- torben-friis 5mo agoThe sooner the better. Let's take a look at the long term, enshittified, viable product before we get too dependent on the trial version.
- siva7 5mo agoThat's so unfair to us hard working developers. A month ago i could buy for .4$ a turn with Sonnet. Now i have to pay at least .9$ for this turn. Weeks ago i could buy for .12$ an Opus turn after they already raised prices and now they want .27$ from me for the same product! They are stealing from us!
- croes 5mo agoThe already stole when they trained their models on the data. Now they just increase the price to buy it back
- asdfasgasdgasdg 5mo agoThey aren't stealing from us, for several reasons. First of all, it's a voluntary transaction. If you don't like the prices, use something else. Or don't use AI at all. Second, you have no idea what their costs are. It is most likely that they are simply passing on their costs to you. If that was not the setup, users would just go to another service provider who was providing tokens at a cheaper rate. It's not like there is a dearth of competitors in this business.
- recitedropper 5mo agoEveryone seems to believe OpenRouter isn't subsidizing but, until they publish audited financials, I personally doubt it.
- deaux 5mo agoOpenRouter doesn't even have hardware. What are they possibly subsidizing? The platform costs? OpenRouter is guaranteed to be about the highest margin operator in the business right now. Everyone wishes they'd be them, skimming 5% off as the middleman without any OpEx.
- recitedropper 5mo agoStreaming, caching, and tool calling can get pretty expensive with scale, even when you don't touch inference. Maybe they're doing something clever and are quite profitable.. or maybe they've already taken $40mm from VCs and are currently trying to raise $120mm at a 1.3B evaluation. They also show headline prices for the cheapest provider of whatever model, but then need to hit different backends some of which may be more expensive. For now they absorb those costs, but the VCs always come knocking. Just my opinion though. Totally agreed that they have one of the best positions amongst all AI providers from a financial standpoint.
- vanviegen 5mo ago> They also show headline prices for the cheapest provider of whatever model, but then need to hit different backends some of which may be more expensive. For now they absorb those costs, [..] They do?? I was under the impression I was just playing the price for whatever provider they deemed 'best' for each completion.
- recitedropper 5mo agoThat is what I had heard. Checking now: The way they describe it in their FAQ is that if the price changes, then they will bill you the new price. But I read that as regarding if the primary model provider changes their headline token cost; not in the case of pricing differences for models that have many different backends that host them. Regardless, I would be more concerned about the streaming costs if the service continues to blow up and they scale aggressively through VC investments. If their 5.5% skim accounted for what they needed, you'd think they could effectively grow organically..
- johndough 5mo agoIt's interesting that the cost multiplier for Claude Sonnet 4/4.5/4.6 varies so much (1/6/9), while the API cost is exactly the same for all three models. Also, the multiplier of 27 for Claude Opus 4.6/4. is way higher than the increase in API price would suggest. I wonder why that is.
- vanviegen 5mo agoOn GitHub copilot you pay per prompt. More powerful models can do a lot more work (consuming a lot more tokens) per prompt. Also, they tend to use more thinking tokens.
- johndough 5mo ago> More powerful models can do a lot more work (consuming a lot more tokens) per prompt. That is not my experience. Each model since at least GPT-4 can fill up an entire context window. In fact, more powerful models can solve tasks faster, so their ratio of multiplier to API price should decrease, not increase. For example, Claude Sonnet 4.6 has a multiplier of 9 and an API price of $15, which is 0.6 multiplier per dollar. Claude Opus 4.7 has an API price of $25, so it should have a multiplier of 25 * 0.6 = 15 when extrapolating from Sonnet, but the multiplier is 27. > Also, they tend to use more thinking tokens. That might be it. Is there any data on this somewhere?
- port11 5mo ago> Is there any data on this somewhere? Anecdata: for me, this is exactly the case with Opus. It _really_ thinks, looks into more sources, more of the codebase, etc. Sonnet is 80% thorough, but Opus can go the extra mile and burns a ton of tokens doing so.
- mkhalil 5mo agoWhy would folks be better paying 5.5% fee to OpenRouter ("Open") if most people just use one or two providers? Just use the provider's API.
- djeastm 5mo agoThe routing automatically routes you to other inference providers (for the same model) if/when the original provider goes down. It's a convenience cost, for sure, but it's not valueless in a fast-moving world. Certainly if you're comfortable with one provider and it's cheaper, do that.
- sally_glance 5mo agoFor me the largest value-add is the unified API. Being able to instantly start trialling a new model with zero code changes is well worth 5%. The other part is not having to deal with billing for multiple platforms.
- themafia 5mo agoWe can't even get slop delivery worked out. So use SlopAggregator instead.
- 2ndorderthought 5mo agoI don't know if it's just me but copilot kind of sucks. I've been running local models with like 9b parameters and they are about as good if not better. Obviously there's no integrations or whatever and I get most people are probably paying for that than anything else but eh. Big no thanks from me.
- mullingitover 5mo agoThe point of this loss leading is to properly hoover up the money in the pockets of enterprise customers, get them locked into the idea that they need the latest and greatest cloud-based model, while simultaneously starving everyone of the memory they'd need in order to run competent models locally. In not-too-distant future we're going to be running better models on our phones than we can buy access to today in the cloud. Skate where the puck is going: soak the customers until that day comes.
- port11 5mo agoI think your first paragraph is spot on, while the second is fairly incorrect. Hardware isn’t getting cheaper at a reasonable pace, and datacenters will keep depleting the market. State-of-the-art models are very, very far from being run on your own hardware.
- mullingitover 5mo ago> State-of-the-art models are very, very far from being run on your own hardware. Still, the models will only get smarter and more efficient as the hardware gets cheaper. The timeframe may be debatable but the outcome really isn't.
- port11 5mo agoI could accept that the end result might be what you propose, but training models is getting tricky, running them more so, now that hardware is becoming pricier. The future might simply be a few feudal lords permitting you to run the best models on their equipment, and a few good open models that most will struggle to run.
- joelthelion 5mo agoCan't wait for people to migrate to open tools (opencode/openrouter). This will unlock a lot of innovation. (I know openrouter is not open, but it allows competition and should be easily replaceable if needed)
- krzyk 5mo agoThose multiplier are only for grandfathered Pro an Pro+ plans that had annual billing, basically a way to scare people of out of those plans. Ant new ones (and bussiness+enterprise plans) will be on token based billing since June 1.
- youwangd 5mo agoShow HN timing matters more than people think. Monday-Thursday, 9-11am Pacific, is when the front page has the most engaged readers. Weekend posts get less competition but also less engagement.
- sandos 5mo agoWow, having a corp. account I do wonder WHEN we are getting some kind of resctriction of usage, or require us to justify our usage. That GPT4-mini change is going to be brutal! Its much better than 5-mini, which was itself much better than earlier free models.
- darqis 5mo agoIT'S NOT SUBSIDIZED