4 ms·
I don't think you can buy a single "MI300X" unit, right? Only the box with x8 of these at a cost of ~250K EUR.
by majke 2mo ago
I don't think you can buy a single "MI300X" unit, right? Only the box with x8 of these at a cost of ~250K EUR.
- baalimago 2mo agoGive it an AI-bubble pop and these will be flooding the market.
- _factor 2mo agoThey will be instantly bought out by companies, not individuals. The consumer bubble won’t pop for quite a while yet. Production also won’t ramp up while lack of real competition keeps the demand high.
- aurareturn 2mo agoWhen is it popping? Is the AI bubble in the room with us now?
- baalimago 2mo agoNext month perpetually
- amrit3128 2mo agoTomorrow? Next year? In 5 years? Nobody can say. But we do know that AI is overvalued, so it WILL pop.
- aurareturn 2mo agoWell, if no body can say when it will pop, then can we really say it's a bubble and it's overvalued? I can tell you it will pop in 10 years and when it pops, it will still be 20x bigger than in 2026. Does that even make any sense? People said AI bubble will pop soon in 2024 and that it was overvalued. Turns out, many AI stocks 10x, 20x since 2024. Actual usage has gone exponential as well. Anthropic revenue went from $100m ARR at start of 2024 to $80b ARR today.
- Joel_Mckay 2mo agoMany are saying July 2027, as in the past these market corrections have correlated with Shrek movie releases. Debt-backed investors have to pay up eventually. =3
- ekidd 2mo ago> Well, if no body can say when it will pop, then can we really say it's a bubble and it's overvalued? Well, given the literal trillions being spent, the only ways this pays off are: 1. AI replaces a non-trivial fraction of human employees. 2. Someone builds a Culture Mind, and humans become (hopefully) pampered pets of AIs we don't understand. Seems unlikely, but it would arguably count as a payoff even if it made money meaningless. Or maybe the AIs don't want pets, and you get SkyNet. Which definitely doesn't care about paying off anyone's investments. When you look at various news articles about investors, yeah, there are definitely a lot of rich people who think that they're going to automate all human labor or just bring about the Singularity. Possibly with them in charge of the rest of us. If you don't make these kinds of wild assumptions, then yeah, this is looking like one of the biggest bubbles ever.
- aurareturn 2mo agoCan we see some actual numbers, projections, models instead of vibes?
- atwrk 2mo agoDebt is at $3 trillion right now: https://fortune.com/2026/07/31/ai-debt-hypescalers-capex-capital-spending-hidden-borrowing-bond-issuance/ https://fortune.com/2026/07/31/ai-debt-hypescalers-capex-cap... Interest alone, at assumed 5%, amounts to about $150 billions per year. That's probably higher than the combined AI revenue of the top 3 providers.
- aurareturn 2mo agoDid you read your own article? It's well less than $3 trillion. Now let's build the model out more. What is the projected revenue, backlog, improvements in existing big tech businesses such as AI helping Meta's ad business?
- slaw 2mo agoThe AI bubble will pop when China gets access to EUV, so the earliest it could happen is 2030
- dghlsakjg 2mo agoNvidia has ever so slightly underperformed the SP500 YTD (at the exact time this comment is being typed), so its basically the apocalypse already.
- tamimio 2mo agoThing is, GPUs will always be on demand, look at their history, initially for gaming, then for hash cracking, then 3D rendering, then for crypto mining, and now AI training and fine tuning. When AI bubble bursts, there will be another bubble taking over. The only solution is more companies making high end units, only competition will make it better for consumers.
- segmondy 2mo agono they won't , the bubble is a financial thing. the demand is real and not going away.
- atwrk 2mo agoThe big question is whether the demand will stay if the subsidized pricing ends. That's what the bubble talk is about. Right now all the players compete for market share and don't care about the losses (hence the debt). But what happens if no one wants to lend them anymore?
- jack_pp 2mo agoI don't think inference is subsidized, it's the training. So what happens is, there's no new models anymore or are released slower.
- FeepingCreature 2mo agoAPI inference is probably not subsidized. Coding plans absolutely are.
- vehemenz 2mo agoGood distinction. The demand is partially driven by the low costs, which are only low because the major providers are losing money.
- qwytw 2mo agoThere is no evidence they are losing money on inference, though? Also if they are keeping the price low because they want to gain market share and reduce the competitiveness of Chinese models they won't be able to raise prices without providers serving open models (at cost + low margin) severely undercutting them.
- vehemenz 2mo agoThere's no evidence they're making money, and we already know from the projected datacenter capacity in a few years that there will be for more supply than demand, so the major providers will have to repay that debt. Even if they are making money on inference, it's nowhere near enough to cover the bill. It's a losing proposition either way, especially with Chinese models now in play.
- zhoutong 2mo agoIt’s available on demand from a few cloud providers. Seems like the cheapest is AMD Developer Cloud (https://www.amd.com/en/developer/resources/cloud-access/amd-developer-cloud.html https://www.amd.com/en/developer/resources/cloud-access/amd-...) powered by Digital Ocean at $1.99/hour. Edit: Now I think about it, this might be the cheapest way to run the DeepSeek V4 Flash 0731 on a dedicated inference server at original weights. I haven’t run mixed load benchmarks but I guess it’s possible to generate $3-$4 worth of tokens per hour and still maintain a usable per-user throughput.
- WASDx 2mo agoAt 830tok/s * 1 hour that's almost 3M tokens which is just $0.54 worth of tokens at Deepseeks current output price.
- krisknez 2mo agoHow is that economically viable? They are selling at a loss?
- dietr1ch 2mo agoThey claim their advantage is knowing how to serve their models efficiently, which is quite possible since they design for it.
- gpugreg 2mo agoAgentic workloads are somewhere around 1%/0.5%/98.5% input/output/cached tokens. Cached tokens are pretty much free for inference providers (if they implement sparse and compressed attention properly) and throughput for input tokens is much higher. Lets assume that you've got 2 million input tokens, 1 million output tokens and 98.5 million cached tokens to process. That would cost 2 * $0.14 + 1 * $0.28 + 98.5 * $0.0028 = $0.8358 with DeepSeek API pricing. For comparison, it would take 2M / 8000 + 1M / 800 = 1500 seconds to process this amount of tokens with the linked framework, which is about $0.83 when we assume $2/hr for one MI300X. However, other inference providers have 10 times higher prices for cached tokens, which results in a comfortable margin. And we should not discount that DeepSeek also gets paid in data, which is probably more valuable to them. And I believe that this framework still has some room for optimization for generation with high batch sizes.
- Lwerewolf 2mo agoThe MI350p exists and should run a decent quant (say, the ~96GB antirez mix) well, but you can get two rtx pro 6000s for one of these, or 8x (actually more) r9700 + probably the gear to run them, etc. Otherwise, you can probably buy one of these second hand from somewhere (SXM A100s are available that way) and run it in an adapter board.
- touisteur 2mo agoI thought MI350P wasn't available yet, curious where to source it right now.
- FuriouslyAdrift 2mo agoThere's at least one systems integrator selling a rack server with 2x MI350Ps. I haven't seen the cards all by themselves yet.
- _joel 2mo agoI thought it was a consumer grade GPU until I saw the 192GB of HBM and 256GB or RAM.
- varispeed 2mo agoTo be fair the development of GPUs have stalled over the years. If they kept up with the progress instead of focusing on enterprise market, likely 256GB consumer GPU would be a norm today.
- FuriouslyAdrift 2mo agoIt's a chopped down MI350X (roughly half the performance)
- FuriouslyAdrift 2mo agoBasically right from Lisa Su's speech: "AMD is essentially taking one of its MI350X accelerators and cutting it in half, resulting in a card with half as many compute resources, half as much memory, and perhaps most importantly, a bit over half of the power consumption"
- throwawayffffas 2mo agoYou can get one on ebay for like 20k, but it comes without the backplane and i dont think there is a pcie card adaptor from china like the ones for h200.