13 ms·
Higher usage limits for Claude and a compute deal with SpaceX
- htrp 5mo ago>Higher usage limits >The following three changes—all effective today—are aimed at improving the experience of using Claude for our most dedicated customers. >First, we’re doubling Claude Code’s five-hour rate limits for Pro, Max, Team, and seat-based Enterprise plans. >Second, we’re removing the peak hours limit reduction on Claude Code for Pro and Max accounts. >Third, we’re raising our API rate limits considerably for Claude Opus models, Looks like Elon's finally giving up on XAI and just selling the compute
- JustSkyfall 5mo agoProbably a good idea in all honesty. xAI is a deeply unserious lab
- cyanydeez 5mo agoThere's only so much determinism you can create when you try not to filter (read CENSOR) your LLM.
- throwa356262 5mo agoFrom a technical standpoint xAI is basically Gemini team B who were give A+ salaries to join the company. But even then, I suspect their hands were tied in some areas because Elon had some expectations from his AI.
- fancyfredbot 5mo agoDid Google outbid Elon for team A? Or A team just don't like Elon?
- throwa356262 5mo agoIt's an internal jokes since very few high profile Deepmind engineers accepted his offer despite some serious cash being thrown at them. Meta engineers on the other hand, couldn't wait to jump ship. But that only reinforces the B team theory.
- lostmsu 5mo agoLLaMA was pretty good at the time
- petercooper 5mo agoI don't know if it relates to the same data centers, but this also comes hours after several still recent Grok models were deprecated at short notice. Grok 4.1 Fast is the cheapest way to do research on X (cheaper than the X API!) and it's gone on May 15: https://docs.x.ai/developers/models https://docs.x.ai/developers/models - freeing up compute to sell?
- peder 5mo ago> Looks like Elon's finally giving up on XAI and just selling the compute I don't think that's certain yet, but I do think that the open-source models like Gemma and Qwen are getting so good so fast that even Anthropic has real risk around the long-term value of their models and tooling. Basically, if I'm Anthropic or xAI, I try to get revenue whenever and wherever possible and see what sticks. There's no value in playing for monopolistic control when everything is so volatile.
- swalsh 5mo agoThere's always money in the giggawatt datacenter
- kingstnap 5mo agoThe details are secret. It very well could be wasted GPU time but Anthropic could have made a killer offering as well. I'm just speculating, but a particularly killer offering Elon wouldnt be able to refuse would be if Anthropic agreed to give them some training data / technology.
- swalsh 5mo agoBillions in revenue just before your IPO isn't a bad deal either.
- fancyfredbot 5mo agoThe icing on the cake for Elon is that it strengthens the competition to OpenAI. Or is that actually his main motivation. Hard to know. Either way it's a win win win for him.
- throwa356262 5mo agoThat's certainly one way one could spin this. I guess loosing a ton of money then trying to get some if it back makes you a genius...
- scottyah 5mo agoYeah real geniuses go down with the ship and never change what they set out to do
- fancyfredbot 5mo agoElon has many many faults but "loosing" money doesn't appear to be one of them. He's literally the richest person alive!
- Rover222 5mo ago[dead]
- croes 5mo agoOr he just got leverage on a competitor
- spikels 5mo agoNo I don't ever give up. I would have to be dead or completely incapacitated. -Elon https://x.com/XFreeze/status/2012390928221094335 https://x.com/XFreeze/status/2012390928221094335
- AlexCoventry 5mo agoI don't think this is giving up. He's getting inside information on how Claude works, and a huge stream of Claude usage data. This will all inform future grok development, IMO.
- vagab0nd 5mo agoGiving Musk the benefit of the doubt, here's a thought experiment: It doesn't seem like any of the big labs in the US can keep a lead for more than 3 months. The Chinese models are closing in. Even if xAI comes up with the best model, so what? On the other hand, power and compute are limited. Ridiculous as orbital compute sounds, land/power on earth is not easily scalable. There are too many limiting factors, chief among which in the US is regulation. But in space, if you make one satellite work, you just get more resources and launch more. This also leads naturally to Tesla's plan for a chip fab. So if you squint, Musk might not be that crazy.
- hn1986 5mo agoquestion is, will they buy cursor?
- Philpax 5mo agoWell, this sucks :/ https://en.wikipedia.org/wiki/Colossus_(supercomputer)#Environmental_impact https://en.wikipedia.org/wiki/Colossus_(supercomputer)#Envir...
- thrownthatway 5mo agoWhy?
- chainwax 5mo agoI think he's referring to the fact that Colossus is powered by fossil fuels.
- xienze 5mo ago[flagged]
- thrownthatway 5mo ago[flagged]
- morgoths_bane 5mo agoHonestly if he was a Nazi that would be less worse than whatever the fuck he is.
- thrownthatway 5mo agoDid you also dance in the street when Charlie got shot in the neck?
- bigyabai 5mo agoSounds like projection to me, that's a non-sequitur.
- cbg0 5mo agoThey're doubling the five hour limits, but no mention about the weekly limit. So overall it's the same maximum usage, right?
- joncik91 5mo agoSome get the reset, some don't it seems :(
- adriand 5mo agoI think so, but that's also really great because I frequently run into the five hour caps, but very rarely use my entire weekly allotment. There are lots of situations where I do things like write the plan for all the work that has to get done, and then set a reminder to execute the plan after I get home, when I'm done making dinner (because e.g. my five hour cap ends at 6pm). Higher caps for the five hour period is a lot more convenient.
- novaleaf 5mo agoI (and many others) are the opposite. I run out of quota is 4-5 days. Generally no issues with the 5hr cap. ($200 sub)
- solenoid0937 5mo agoLike 90% of people I know never hit their weekly but they hit their hourly. I'd bet your case is way rarer.
- farfatched 5mo agoIf this logic applied, then there would be no purpose in them having the 5 hourly limit.
- cbg0 5mo agoThe purpose is to control the total amount of requests they need to handle in a given timeframe. If everyone could use up their whole weekly limit in 5 hours, many would do so, thus pushing the GPU/TPU clusters to or above their capacity limits.
- minimaxir 5mo ago> First, we’re doubling Claude Code’s five-hour rate limits for Pro, Max, Team, and seat-based Enterprise plans. The fine-print-omission appears to be that weekly limits are not doubled. The progressive 5-hour rate limit shrinking was indeed an efficiency blocker that finally convinced me to cancel, but being only able to get 4 full sessions a week as opposed to 8 doesn't compell me to resubscribe.
- dw_arthur 5mo agoFor my hobbyist purposes Deepseek v4 Flash has replaced Claude Code because I was also sick of hitting 5 hour limits with Claude. Right now, the only thing I miss from Claude is multi-modal image support. I can work around no image support since I can use v4 Flash all day and spend around $1. I am aware Deepseek is currently discounting their API at 75% off so I may try out another provider once the discount is gone at the end of the month. At this point if feels like if you properly scope your work open weight LLMs are adequate.
- farfatched 5mo agoOuch, I wasn't aware they were discounting so much. There goes my subscription escape plan.
- dw_arthur 5mo agoGood news, I was wrong. The discount only applies to Deepseek v4 PRO. Deepseek v4 Flash is not currently on discount which means the dirt cheap price will stay the same.
- mostafas 5mo agoThe 75% discount only applies to Deepseek v4 Pro. Flash will stay the same price after the discount ends. It's remarkably cheap for what it delivers. https://api-docs.deepseek.com/quick_start/pricing https://api-docs.deepseek.com/quick_start/pricing
- scottyah 5mo agoThe datacenter isn't operational yet, they don't magically get more processing instantly after signing a deal.
- SilverElfin 5mo ago[flagged]
- bigyabai 5mo ago"HN pretends that companies have morals: Part 48,037,986"
- slopinthebag 5mo agoCounterpoint: Valve Which is kind of like the exception that proves the rule hahaha
- bigyabai 5mo agoValve isn't moral, they're just privately owned. CS cases have given them enough fuck-you money to rehabilitate their image in any way they see fit.
- minimaxir 5mo agoYou haven't been following the discourse around a) how Steam handles GenAI disclosures and b) how Steam handles forum/review moderation. People haven't been saying "GabeN can do no wrong" for awhile.
- slopinthebag 5mo agoI was motivated to post this because I was just reading a thread where many users were praising Valve and GabeN for how their company is run, but I'm curious to read more about A & B.
- richwater 5mo ago[flagged]
- josefresco 5mo agoIt is very much a valid argument. SpaceX has been working on this issue for years. https://theconversation.com/a-million-new-spacex-satellites-will-destroy-the-night-sky-for-everyone-on-earth-277938 https://theconversation.com/a-million-new-spacex-satellites-... FTA: "SpaceX has done a lot of engineering work to make its Starlink satellites fainter. They are still too bright for research astronomy, but thanks to new coatings, their brightness has not increased dramatically even as SpaceX has launched larger and larger satellites."
- athrow 5mo agoAnthropic taketh and Anthropic giveth.
- HarHarVeryFunny 5mo agoExactly. Today they say this, then tomorrow they'll silently reduce limits and argue with anyone who calls them on it.
- solenoid0937 5mo agoThis is obviously because of the new compute deal. I don't see them going back unless prosumer demand outgrows compute again.
- stavros 5mo ago> First, we’re doubling Claude Code’s five-hour rate limits for Pro, Max, Team, and seat-based Enterprise plans. Ok I guess, this was a bit of a hassle, but you're not increasing my weekly allowance, you're just not annoying me as often. > Second, we’re removing the peak hours limit reduction on Claude Code for Pro and Max accounts. It wasn't a limit reduction (as in, I didn't have a lower 5-hour limit), it was "tokens are more expensive" and it ate my weekly limits faster. This should never have been instituted to begin with. > Third, we’re raising our API rate limits considerably for Claude Opus models, as shown in the table below: Meh. This is why I don't care for all the "it's a subscription, you're free to not use it!" arguments here. It's not an all-you-can-eat subscription with some generous fair use limits, it's a "X tokens per month for $Y", and they keep lowering the X unilaterally and in secret.
- solenoid0937 5mo agoPeople are so cynical on HN. Just move to API billing if not getting enough subsidized compute is that big a deal for you?
- stavros 5mo agoIs that what you do when you prepay for a year to get a discount and the supplier just says "oh I'll just give you half of what you paid for"? You "just move to pay again for the rest"?
- solenoid0937 5mo agoSorry, were you told you'd be given a specific number of tokens when you subscribed?
- stavros 5mo agoYes? Weren't you? Did you think you were buying a token lottery, where you'd have a billion tokens one day and zero the next?
- iamleppert 5mo agoHopefully they will work on response time. I've been noticing it taking 5+ minutes for each turn, for not complicated requests. Seems to vary based on time of day too.
- amacbride 5mo agoAs a bonus, it looks like they reset limits a few minutes ago -- I went from 53% of my weekly allotment to 0%.
- readitalready 5mo agoNot for me. 2 Claude Max 20x accounts here both at high usage on weekly allotments.
- arian_ 5mo agoAnthropic renting out the data center Elon built for Grok is the kind of plot twist you can't make up.
- brokencode 5mo agoPretty smart for SpaceX though. They’re turning an asset they made for a money-pit (Grok) into probably a major source of revenue ahead of their IPO.
- 23rf 5mo agoIts not even that. Its better to be involved in the game with a leader/help out a competitor who is competing against someone you don't like and don't want them to win, than to sit it out.
- giwook 5mo agoThe enemy of my enemy is my friend.
- bombcar 5mo agoMaxim 29: The enemy of my enemy is my enemy's enemy. No more. No less.
- thrownthatway 5mo agoFerengi Rule of Acquisition #76: Every once in a while, declare peace. It confuses the hell out of your enemies.
- a4isms 5mo agoThis is something you say aloud, while muttering "useful idiot" under your breath.
- deleted 5mo ago[deleted]
- hparadiz 5mo agoGive them whatever they need. Time to go to the moon.
- deleted 5mo ago[deleted]
- tanh 5mo agoWouldn't trust them not to take a copy and use it to distill. Wonder what security there is
- y42 5mo agoI want to believe. A couple of weeks ago I fell into this "trap", they offered a similar thing. I subscribed to the Pro Plan. Had fun for a couple of weeks and then I entered frustration phase. I love the product, but I hate those up and downs. My rant made it to HN front page - which I am not happy of. I want the stuff I build to be seen on the front page.
- gpugreg 5mo ago> As part of this agreement, we have also expressed interest in partnering with SpaceX to develop multiple gigawatts of orbital AI compute capacity. Anthropic is either taking this space business more serious than the general public, or posting this sentence was part of the deal to get the compute.
- JMKH42 5mo agoI don't think space compute is going to work out, but I would certainly say "yes happy to buy space compute from you in the future if you offer it at a good price" If it happens it happens, if not, it doesn't.
- CamperBob2 5mo agoIt makes no sense. We're being presented with a forced choice -- put them in space, or put them in the middle of downtown Seattle. This is stupid. I don't understand what's happening... specifically, what mental virus is spreading that lowers everybody's IQ by 10-20 points, evidently including my own. Put the data centers in the ocean, powered by solar and networked with Starlink or LEO. Put them in the desert. Put them 20 miles south of Nowhere, Idaho. But space?!
- Karrot_Kream 5mo agoBecause the US has levied high tariffs on solar cells, can't build their own solar cells economically enough, and has such a torrid permitting system that it can't build transmission lines. Natural gas is the only form of generation that's easy to permit outside cities (due to pipeline agreements and this admin fast-tracking natural gas generation approval) but few cities will allow one. DCs need to be built within low latency interconnect of urban areas or else they become uncompetitive. Elon claims (which I take with a huge grain of salt because he's made endless broken promises in investor calls and interviews) that he disagrees with the administration's stance on solar and would use it to power his DCs if he could, but contends that permitting is a huge problem. The US needs to figure out how to build again. > This is stupid. I don't understand what's happening... specifically, what mental virus "Be kind. Don't be snarky. Converse curiously; don't cross-examine. Edit out swipes"
- boramalper 5mo agoI wonder if it's just Elon realising that xAI can't beat OpenAI and thus deciding to give all his compute capacity to Anthropic instead. Certainly an interesting day for xAI.
- bpodgursky 5mo agoBuilding datacenters plays to his strengths. It's a good partnership if he can stomach it.
- consumer451 5mo agoI would think that stomaching Musk would be the hard part. Just goes to show how compute-constrained Anthropic is at this time. He literally did a Nazi salute on stage, twice! Check the video, and tell me what you see. edit: https://giphy.com/gifs/elon-musk-nazi-salute-8W0ItVv7T1kRdwbTTn https://giphy.com/gifs/elon-musk-nazi-salute-8W0ItVv7T1kRdwb...
- 1234letshaveatw 5mo agomeh, he's no Graham Platner
- consumer451 5mo agoSee, if you have philosophical/morale standards and don't just subscribe to tribalism, then you could say: f both those morons. Could you join me in that statement?
- w4yai 5mo agoIf it is a Nazi salute, it's a real bad one !
- consumer451 5mo ago"The Party told you to reject the evidence of your eyes and ears. It was their final, most essential command." https://old.reddit.com/r/gifs/comments/1i7w4nz/comparison_of_elon_musks_nazi_salute_with_real/ https://old.reddit.com/r/gifs/comments/1i7w4nz/comparison_of...
- nethunters 5mo agoHopefully this filters through to Copilot's recent rate-limits
- skeledrew 5mo agoOh. Just as I'm in the process of migrating to Pi+Qwen (local). This was probably going to be my last month on the Pro sub as I'm seriously fed up with the limits and degradation that started weeks after I signed up. Let's see how this shakes out.
- flumpcakes 5mo agoHow does Pi+Qwen (local) compare to Anthropic's offerings? Surely you're not getting the same breadth and quality of output using Qwen? How is the performance?
- skeledrew 5mo agoSo far I've only really set things up and done some benchmarking (a set of capability prompts created and evaluated by Claude, HumanEval and MBPP; haven't completed the latter 2) on several local models (Qwen 1.7b, 4b, 9b & 35b a3b; 1.7b got 6/8 correct at ~14.7 tok/s on the capability set, to 35b for 8/8 at ~4.5 tok/s; can share full results if interested), and setup llama-swap so I can dynamically select them. I'll need to decide which of my projects I'll be really testing them on, with the awareness that I'll have to be even more involved.
- dymk 5mo agoIt’s a toy compared to Opus or Sonnet. Obviously the 5 trillion parameter models running on $$$$ hardware is going to outperform a local model.
- z3ratul163071 5mo agowe all are. thank god for alibaba, seems all that crap we bought from aliexpress served some indirect purpose.
- alxsuv 5mo ago[flagged]
- int32_64 5mo agoWhat's the current status of the 'biggest computer wins' vs. specialized proprietary research/data in the AI arms race? People had such high hopes for xAI because of the monster machine Elon built. Or has xAI just turned over too much staff too quickly?
- antipaul 5mo ago"All of [SpaceX]'s compute capacity at Colossus 1" SpaceX/xAI also has Colossus 2, with double or more the GPUs Seems xAI will still be around
- deleted 5mo ago[deleted]
- empath75 5mo agoThat is one take, but here is how I interpret it. They spent a lot of money training a model which isn't doing enough inference to justify continuing to use those GPUs, and now they are buying even more GPUs to build an even bigger model that also won't be very popular.
- deleted 5mo ago[deleted]
- nikolay 5mo ago[flagged]
- mirzap 5mo agoDoubling the five-hour rate limits is merely a marketing stunt if the weekly rates are not also doubled. It simply means that you can reach the weekly limits in three days instead of five.
- swalsh 5mo agoI have never come close to my weekly limit, but have hit my hourly limit frequently.
- mirzap 5mo agoFor me it's the opposite. I almost never hit hourly limit, but I hit weekly limit in about 5 days.
- extr 5mo agoWhat does your usage look like day to day? Are you using a low level amount all day long? I'm with the others here, I've never hit the weekly limit ever, only the hourly, and I consider myself a heavy user.
- mirzap 5mo agoI dedicate a significant amount of time to defining the precise actions that agents should perform (PRD/ADR). I break down the feature sets into Milestones and slices (tasks). These tasks are small, well-defined, and scoped. I have a prompt template that the “architect” agent prepares whenever I want to initiate a new feature. This ensures that the prompt structure remains consistent and standardized over time. The generated prompt is then pasted to the “orchestrator,” which performs context discovery (using Repoprompt) and finalizes the plan then proceeds to launch subagents to do the work. Based on the size and complexity of the task, as well as any inter-task dependencies, the orchestrator deploys one or more subagents (sometimes 5 or 6 subagents) to work on these mini tasks. Once all tasks are completed, the orchestrator initiates verification and launches a review workflow. This workflow uses the original prompt, acceptance criteria, repository internal guidelines, and relevant skills to conduct a thorough review of the agents’ work. Typically, there are one or two review iterations, during which the review agent identifies any issues. Sometimes, I may also notice issues and have to "steer" the orchestrator. The time required for a slice to complete ranges from 30 minutes to 4 or 5 hours, depending on its size, complexity, and the number of subtasks it contains. Only if I run about 3 such orchestration in parallel I can reach hourly limit.
- Marciplan 5mo agoIf Anthropic and SpaceX and OpenAI are all going public this year then this is a clever move to stick it to OpenAI. However, I'm kinda sus of my Claude subscription now
- lairv 5mo agoFor a space that supposedly had "no moat", the number of players still competing for frontier models seems to be shrinking pretty fast
- swader999 5mo agoWhat's going to be the hit on our atmosphere when the data centers re enter? I guess it won't matter as the AI will replace the humans by then for the GDP and tax base.
- swalsh 5mo agoModels are a commodity, let's say Elon actually figures out building datacenters in space, or maybe he continues to be the leader of building earth based datacenters. Probably better business to not have yourself as your only customer. Dogfood, and open it to all.
- mplewis 5mo agoThe first is impossible and the second isn't happening and won't happen.
- croes 5mo agoI wouldn’t say impossible but not effective
- nextstep 5mo agothe leader of building earth-based datacenters lol what are we even talking about
- driverdan 5mo ago> Elon actually figures out Elon doesn't figure out anything. He pays people to do it and then tries to take the credit.
- 0xbadcafebee 5mo agoColossus 1 datacenter is the one using illegal power, is poisoning the air for poor communities near Memphis, and is potentially poisoning the water. It's likely the additional demand on the grid will cause massive blackouts during extreme weather events, putting residents at further risk. https://en.wikipedia.org/wiki/Colossus_(supercomputer)#Environmental_impact https://en.wikipedia.org/wiki/Colossus_(supercomputer)#Envir... So you can put Anthropic on your list of companies that like to talk big about safety, but when the rubber hits the road, profits matter more than safety.
- boldlybold 5mo agoIllegal is a strong term here. While the wiki link you included indicates there might be some permitting nuances, I've seen nothing claiming the power is "illegal."
- fancyfredbot 5mo agoThe ethics are questionable, legal or not. Anthropic are tarnishing their image again here. Not sure how much it hurts then compared to blocking openclaw though.
- port11 5mo agoI find the ethics of power generation, resource use, and pollution in a world struggling with climate change to be more of a challenge than whether a few people can run some software. And that’s coming from a Claude user that’s getting tired of their shenanigans.
- DonsDiscountGas 5mo agoI don't quite understand the business logic behind "blocking" openclaw (you can still use it at API rates) but I never saw how this was unethical. Anthropic has no ethical obligation to support other people's software
- butlike 5mo ago
- Rover222 5mo agoReading the comments here again surprises me how in an anti-Elon bubble most folks are. They are renting out spare Colossus 1 capacity. Colossus 2 is still coming online. Orbital data centers are really the plan in the next few years. XAi is still behind, but not a disaster considering how late they entered (and Elon’s unfortunate fixation on anime characters). SpaceX is extremely uniquely positioned to crush the rest of the world combined in order to orbital data centers.
- HarHarVeryFunny 5mo ago> SpaceX is extremely uniquely positioned to crush the rest of the world combined in order to orbital data centers Sure, as long as your data center is 3x4m - size of a Starlink satellite (think Spinal Tap Stone Henge) . Anything bigger than that (i.e. actual data center sized) is going to require some assembly. I've heard TeslaBot is good at folding shirts, and serving drinks (at least while teleoperated) - perhaps it can help?
- everfrustrated 5mo agoOrbital DCs is blocked on starship because they _arent_ doing starlink sized satellites.
- HarHarVeryFunny 5mo agoYou're not going to fit a data center in a Starship either, unless you are talking a Tiny Corp Exabox "data center in a shipping container" sized one. Even something that small (1MW) would still need 4x the solar capacity of the ISS, and therefore likely some assembly required. Then you've got latency from satellite to satellite ... In any case, it appears that Musk can't even generate enough AI demand to utilize his own ground based data center. Maybe he can add "data centers in space" to part of his Mars colonization plan. Maybe have Tesla Bots driving around in Cybertrucks too ?
- Rover222 5mo agoOkay well you’re clearly not understanding the basic concept of the orbital architecture.
- 4b11b4 5mo agoI mean... seems like a no-brainer
- 2001zhaozhao 5mo agoI mean, as someone who has the Max 20x plan and uses it only outside work (so I could not hit anywhere close to the weekly limit at all), I'll gladly take the 5-hour limit doubling. My first impression to this post is "what the hell are they thinking?", but actually it seems like a decent move by them. They basically made it so that normal users can better utilize their plan while not benefitting the backgroundagentmaxxers and stealth openclaw abusers in the ranks of their subscription audience. Making their plan more attractive to the people they actually want to sell to. Hopefully this leads to a loosening of harness restrictions later.
- everfrustrated 5mo agoFor those who haven't been following the build out. xAI has added about 500MW of nvidia gpu capacity in ~April and will add another 500MW before the end of the year totaling about 2GW.
- aurareturn 5mo agoSource on this?
- everfrustrated 5mo agoFollowing various X posts, asking Grok to check my memory. https://grokipedia.com/page/Colossus_supercomputer https://grokipedia.com/page/Colossus_supercomputer backs this up.
- solarkraft 5mo agoElon Musk is not a reliable source.
- hirvi74 5mo agoI would have preferred an increase in the weekly limits instead of the 5-hour limits.
- lossolo 5mo agoxAI (part of SpaceX) using just 11 percent of GPUs[1] 1. https://wccftech.com/xai-using-just-11-percent-gpus-while-meta-google-squeeze-out-much-more/ https://wccftech.com/xai-using-just-11-percent-gpus-while-me...
- stuaxo 5mo agoOh is this the polluting gas powered data centre Elon made, that's making local residents unhealthy ? This might be a good time to drop Claude.
- Geee 5mo agoFor context, xAI GPU utilization is at 11% and they're also expanding.[0] Renting one datacenter to Anthropic doesn't mean that they would be shutting xAI / Grok down. [0] https://wccftech.com/xai-using-just-11-percent-gpus-while-meta-google-squeeze-out-much-more/ https://wccftech.com/xai-using-just-11-percent-gpus-while-me...
- redox99 5mo ago11% is the MFU. Not how many GPUs are running. Big misunderstanding.
- mark_l_watson 5mo agoThe politics and economics of Musk throwing some support towards Anthropic is interesting (samma is probably pissed). But, if you will pardon a little rant: I hate the idea of subscription inference plans and also 'dumping' by subsidizing non-profitable products. Inferencing should be pay as you go and dumping illegal.
- AlexCoventry 5mo agoSo now Elon Musk gets to read all of our Claude conversations?? :-(
- mandeepj 5mo ago"use all of the compute capacity at their Colossus 1 data center" So, they handed out all of their data center to Anthropic; Grok wasn't using it much?
- jizzywizzy 5mo agoMoved to Colossus 2. Though I guess you could still frame it as 'don't they need Collosus 1 AND 2' if you want...
- gck1 5mo agoLimits were the last straw that made me cancel my subscription and make my workflow completely model agnostic with pi. While this is good news, I'm not coming back. Anthropic just lost me with too many wrongs in too short of a time period. Opus has been replaced with GPT 5.5, DeepSeek, Kimi, Qwen and they all allow me to use my own, single harness and switch models easily if any of them start treating me the same.
- farfatched 5mo agoSame, though I'm reconsidering, in light of the recent bugs (which can happen to any provider) and the increased limits. I guess that's at least 3x more Opus for my usecase.
- sergiotapia 5mo agoI wouldn't make any grand stand declarations like this honestly. The models themselves are all hot swappable with minimum effort. The AI labs american or chinese don't really have a moat. Today anthropic is bad and openai is good. Last month it was the other way around. Next month it may be google. The only certainty is that you can swap models quickly and painlessly.
- maelito 5mo agoThese energy figures are gigantic. It's getting absurd.
- losvedir 5mo ago> 300 megawatts of new capacity (over 220,000 NVIDIA GPUs) The scale is just mindboggling here. Are there any blog posts or anything discussing what kind of infrastructure is used for even just the inference side (nevermind the training) for SotA models like Opus? I would have thought it might be secret, but given that you can actually run the models yourself on AWS Bedrock doesn't that give an indication?
- giwook 5mo agoHow many instances of Doom can it run though?
- epistasis 5mo agoI know you're probably talking about the compute infrastructure, but I think the electricity infrastructure side is interesting too, data centers are doing things in dumb ways because the need for operational expansion speed is greater than the need dollars: > It’s regulation with the utilities. There are ramp rates, there are all of these things that you’re supposed to do to not screw up the grid. Data centers have been in gross violation of that. When you think about what’s wrong with data centers, they have load volatility, which we just talked about, then they decide to power it with behind-the-meter natural gas generators. These natural gas generators, their shaft is supposed to last for seven years. It’s lasting 10 months because of all the cycling. https://www.volts.wtf/p/doing-data-centers-the-not-dumb-way https://www.volts.wtf/p/doing-data-centers-the-not-dumb-way On the compute infrastructure, there are standard NVIDIA reference designs like this: https://www.nvidia.com/en-us/technologies/enterprise-reference-architecture/ https://www.nvidia.com/en-us/technologies/enterprise-referen... I haven't bothered to look but I'd guess Mellanox GPU-to-GPU networks, and massive custom code for splitting tensors across GPUs, and for shuttling activations across GPU nodes.
- airspresso 5mo ago> but given that you can actually run the models yourself on AWS Bedrock That's not exactly how it works. Anthropic are hosting their models in AWS Bedrock as a managed service. Customers call those LLMs just like calling any other API. There's no visibility into what kind of AWS infrastructure is serving that API request.
- kristianp 5mo agoWhat GPUs does Colossus run? Old H100s?
- breakingcups 5mo agoI could have used this news 2 days ago. I've been trying out Claude Code for a few days and kept running into the limit, so I wanted to upgrade to Max. In the upgrade-flow they hit me with an identity verification through Persona. No problem, I thought, I'll just cancel the upgrade. Nope, all access to Claude Code on the old plan was now also blocked and can't be unblocked without completing Identity Verification, which I'll never do. What a bad experience. On the plus-side, it told me how much cheaper Deepseek is and that it's on parity for reverse engineering work.
- exabrial 5mo ago> Within the month To me this is the mind-bending piece. It's not a like a datacenter has a plug-and-play with well written spec and an international standard interface.
- exabrial 5mo agoThis is where I see the economy of AI going: * Inference becomes cheap - speciality accelerators hit the market and race to the bottom begins * Training remains expensive - This works out for Anthropic/OpenAI, they go into the business of training * Models become rental units or purchasable assets, you run on inference hardware - Rent or own inference hardware * Or you pay someone to do all of the above for you, at a premium
- kcb 5mo agoThere's no magic bullet for inference on cheap accelerators. Any accelerator will still require large amounts of high bandwidth memory.
- exabrial 5mo agoThe way to do it _today_ requires enormous amounts of HBM! However, we've never designed inference accelerators, which is actually a quite "trivial" problem, but we've just never had a need. Groq (acqui-hired by NVidia) came up with a different processor architecture: metric shit-tons of SRAM attached to a modest single core deterministic processor. No HBM needed on this card, and 32x faster inference than today's best GPUs at inference! These LPUs are pretty useless for training though, which is useful for companies training models! Training is expensive, inference is cheap (someday, not now). There's also a Canadian company that _literally burned the model as a silicon mask_ on a chip. It's unbelievably (1000x) fast, but not flexible of course: https://chatjimmy.ai https://chatjimmy.ai
- kcb 5mo agoThe point is metric shit-tons of SRAM is still large amounts of expensive memory.
- exabrial 5mo agoSRAM and HBM are two completely different things though... SRAM is what your L1,L2,L3 caches are made of (most of the time, asterisks exist). This is something we've been doing for years and is a proven technology thats unbelievably cheap. It's all part of the processor. HBM are their own chips and dies.
- logicalappeals 5mo agoAnthropic looking to garner some good will after the recent issues. I’ll gladly take the higher rate limits
- LNSY 5mo agoToo little, too late. I'd rather have a consistent, dumber model than sometimes excellent but often miserable Claude. Staying with Claude is like going back to the restaurant where you got food poisoning: you kinda get what you deserve next time you get sick.
- gigatexal 5mo agoI would gladly take a worse experience than to have my favored LLM vendor partner with an Elon company.
- sourcegrift 5mo agoAnthropic aligning with the guy who got Trump elected means they are dead to me. I'm posting immediately after cancelling my claude subscriptions.
- tintor 5mo agoSo, when you use Claude Code from now on, you will be poisoning people in the rural parts of Memphis with xAIs unpermitted gas generators.
- dagi3d 5mo agodoes that mean this data center was way overprovisionedo or that grok is barely used and they could potentially kill it and just use claude?
- ilia-a 5mo agoInteresting that the 5h limits are raised, but if I understand announcement correctly, the weekly limit is not. So all this means is that you can burn through your weekly limit faster and be locked out entirely, or having to buy tokens
- nl 5mo agoSay what you like about Sam Altman, but given how Anthropic is scrambling to sign capacity deals for compute we can sure say he was right about the capcity build out needed.
- stingraycharles 5mo agoThat’s correct, but from what I understand his move was also strategic: to choke the market. Having said that, Anthropic’s position is fully understandable, as Sam took a very large risk here, and OpenAI’s future is all but certain.
- lmm 5mo agoOr the bubble he was pumping hasn't popped yet. We won't be able to say how much of this capacity was actually "needed" until 10 years in the future, if ever.
- nl 5mo agoThe point is that Anthropic is already a decent way into eating through all that capacity, and it's based on real revenue.
- lmm 5mo agoSome entities are paying for it, sure. I'm still not convinced that's because it's "needed".
- nl 5mo agoTrue. No one "needs" the internet or computers, right?
- lmm 5mo agoPeople are getting real stuff done with the internet. But there were also a whole lot of overhyped companies that rightly crashed back in '99. Once we've gone through the AI equivalent of the dot.com crash, will Anthropic still be scrambling for more capacity, or will they have more than they can profitably use, like the dark fiber we were left with last time?
- Frannky 5mo agoI shared a couple of days ago why they were not doing like Google and offering oss models, but damn, offering Anthropic models after all the badmouthing. Next news: OpenAI models live on Colossus 2
- deleted 5mo ago[deleted]
- Aeolun 5mo agoThey say usage limits on the 5h increased, but I don’t see a significant difference on the x20 plan.
- espeed 5mo agoHow do you select your data center like you can for AWS and Google Cloud?
- freakynit 5mo ago[dead]
- zelon88 5mo ago> We’re very intentional about where we’ll add capacity—partnering with democratic countries whose legal and regulatory frameworks support investments of this scale, and where the supply chain on which our compute depends—hardware, networking, and facilities—will be secure. *Buys compute from actual fascist Elon Musk in a failing democracy during the death throes of late state capitalism.
- chillfox 5mo agoWell, that's super disappointing :( I have got xAI blocked in OpenRouter as I do not want to support any business controlled by Musk.
- docmars 5mo agoInsert gaping soyjak face.
- danaw 5mo agothis feels more like something designed to bolster the spacex ipo than anything else 300MW is peanuts compared to their multiple 50GW+ deals to the point you start to wonder why just 300MW is making the difference in their capacity that they can increase limits this much... also, why couldn't their many existing multi billion dollar deals not allow them to expand capacity? when you take this into account, then you read their statement about orbital compute it starts to smell quite fishy
- sailingparrot 5mo agoThe difference is the 300 MW are real, the 50GW are printed on some paper and don’t exist. There aren’t that many 300MW+ datacenter in the world, relative to the capacity Anthropic has online, it’s a lot, probably in the 20% range.
- Jackson__ 5mo agoOh what is that, the most "ethical" AI company on the planet making deals with literal democracy undermining fascists? I'm starting to think the problem with "ethical" AI was always that no company could ever act ethically in the long term. They are and always will be a cancer to society and AI will only serve to amplify this further.
- crazyate 5mo ago[flagged]
- grim_io 5mo agoNot surprising, considering the recent news that xai only utilized 11% of their GPU's.
- deafpolygon 5mo agoWhat does SpaceX get out of this deal?
- frumiousirc 5mo agoIt's been a day or so and I don't see any change at https://usage.report/ https://usage.report/