25 ms·
The RAM shortage could last years
- ochre-ogre 6mo agocan't read the article due to a paywall.
- fouc 6mo agoI'm a bit surprised the article makes no mention of Google's TurboQuant[0] introduced 26 days prior. Given that TurboQuant results in a 6x reduction in memory usage for KV caches and up to 8x boost in speed, this optimization is already showing up in llama.cpp, enabling significantly bigger contexts without having to run a smaller model to fit it all in memory. Some people thought it might significantly improve the RAM situation, though I remain a bit skeptical - the demand is probably still larger than the reduction turboquant brings. [0] https://news.ycombinator.com/item?id=47513475 https://news.ycombinator.com/item?id=47513475
- WesolyKubeczek 6mo agoYou can still use as much memory, but fit more things into it, so I don’t think the current market hogs will let go easily.
- Bombthecat 6mo agoYou still need to hold the model in memory. If you have for example 16 GB ram, the gains aren't that much
- anon373839 6mo agoThat's not what consumes the most memory at scale. The KV caches are per-user.
- WingEdge777 6mo ago[dead]
- tuetuopay 6mo agoThe net effect won’t be a memory use reduction to achieve the same thing. We’ll do more with the same amount of memory. Companies will increase the context windows of their offerings and people will use it. That is the sad reality of the future of memory.
- ehnto 6mo agoI am not convinced that more context will be useful, practical use of current models at 1mil context window shows they get less effective as the window grows. Given model progress is slowing as well, perhaps we end up reaching a balance of context size and competency sooner than expected.
- tuetuopay 6mo agoStuff in more code. Stuff in more system prompt. Stuff in raw utf8 characters instead of tokens to fix strawberries. Stuff in WAY more reasoning steps. Given the current tech, I also doubt there will be practical uses and I hope we’ll see the opposite of what I wrote. But given the current industry, I fully trust them so somehow fill their hardware. Market history shows us than when the cost of something goes down, we do more with the same amount, not the same thing with less. But I deeply hope to be wrong here and the memory market will relax.
- lhl 6mo agoBTW, a number of corrections. The TurboQuant paper was submitted to Arxiv back in April 2025: https://arxiv.org/abs/2504.19874 https://arxiv.org/abs/2504.19874 Current "TurboQuant" implementations are about 3.8X-4.9X on compression (w/ the higher end taking some significant hits of GSM8K performance) and with about 80-100% baseline speed (no improvement, regression): https://github.com/vllm-project/vllm/pull/38479 https://github.com/vllm-project/vllm/pull/38479 For those not paying attention, it's probably worth sending this and ongoing discussion for vLLM https://github.com/vllm-project/vllm/issues/38171 https://github.com/vllm-project/vllm/issues/38171 and llama.cpp through your summarizer of choice - TurboQuant is fine, but not a magic bullet. Personally, I've been experimenting with DMS and I think it has a lot more promise and can be stacked with various quantization schemes. The biggest savings in kvcache though is in improved model architecture. Gemma 4's SWA/global hybrid saves up to 10X kvcache, MLA/DSA (the latter that helps solve global attention compute) does as well, and using linear, SSM layers saves even more. None of these reduce memory demand (Jevon's paradox, etc), though. Looking at my coding tools, I'm using about 10-15B cached tokens/mo currently (was 5-8B a couple months ago) and while I think I'm probably above average on the curve, I don't consider myself doing anything especially crazy and this year, between mainstream developers, and more and more agents, I don't think there's really any limit to the number of tokens that people will want to consume.
- fy20 6mo agoThe work going into local models seems to be targeting lower RAM/VRAM which will definately help. For example Gemma 4 32B, which you can run on an off-the-shelf laptop, is around the same or even higher intelligence level as the SOTA models from 2 years ago (e.g. gpt-4o). Probably by the time memory prices come down we will have something as smart as Opus 4.7 that can be run locally. Bigger models of course have more embedded knowledge, but just knowing that they should make a tool call to do a web search can bypass a lot of that.
- muyuu 6mo agothat will only increase the demand for RAM as models will now be usable in scenarios that weren't feasible prior, and the ceiling for model and context size is not even visible at this point I hate to mention Jevons paradox as it has become cliche by now, but this is a textbook such scenario
- gajjanag 6mo agoTurboQuant is known across the industry to not be state of the art. There are superior schemes for KV quant at every bitrate. Eg, SpectralQuant: https://github.com/Dynamis-Labs/spectralquant https://github.com/Dynamis-Labs/spectralquant among many, many papers. > Given that TurboQuant results in a 6x reduction in memory usage for KV caches All depends on baseline. The "6x" is by stylistic comparison to a BF16 KV cache; not a state of the art 8 or 4 bit KV cache scheme.
- throwaway613746 6mo ago[dead]
- NooneAtAll3 6mo ago> 6x reduction in memory usage for KV caches and up to 8x boost in speed mind that you're quoting marketing material that's largely based on unfair baseline testing (like comparing 4 bit vs 32 bit to get "8x speed") https://www.youtube.com/watch?v=haoAI2lIZ74 https://www.youtube.com/watch?v=haoAI2lIZ74
- ethan_smith 6mo agoYour skepticism is well placed. Every time a new quantization or compression technique drops, the immediate response is to just scale up context length or run a bigger model to fill whatever headroom was freed up. It's Jevons paradox applied to VRAM - efficiency gains get eaten by increased usage almost immediately.
- tim-projects 6mo agoThe era of optimisation is finally here. I'm excited.
- jeff_vader 6mo agoWait until China invades Taiwan.. (ok, it's not too likely, but what if?)
- Renaud 6mo agoI think RAM shortages would be the least of our problems… Assuming China takes TSMC in one piece (unlikely without internal sabotage in the best case scenario), it would still probably take years before it produces another high end GPU or CPU. We would probably be stuck with the existing inventory of equipment for a long time…
- necovek 6mo agoI am surprised we consider TSMC like a natural resource: isn't it really a combination of know-how and build-out according to that know-how? If smarts leave the country, perhaps this moves with them. The risk with China taking over Taiwan is that they mostly expedite their own production research by a couple of years.
- ndepoel 6mo agoIt kinda does resemble a natural resource though. The machines and technology in use at TSMC are so insanely complex, that there isn't a single person on earth who knows everything about how it works. TSMC functions only because of all of the pieces of the puzzle being together in the right place and arranged in just the right way. It's a very fragile balance that keeps it all running, and a major disruption could mean we get thrown back by a decade in chip-making technology.
- danaris 6mo agoWhat you say is absolutely true, and is a serious problem—but the way our system operates does not allow us to correct for it. Anyone trying to spin up a competitor to TSMC would have to first overcome a significant financial hurdle: the capital investment to build all the industrial equipment needed for fabrication. Then they'd have to convince institutions to choose them over TSMC when they're unproven, and likely objectively worse than TSMC, given that they would not have its decades of experience and process optimization. This would be mitigated somewhat if our institutions had common-sense rules in place requiring multiple vendors for every part of their supply chain—note, not just "multiple bids, leading to picking a single vendor" but "multiple vendors actively supplying them at all times". But our system prioritizes efficiency over resiliency. A wealthy nation-state with a sufficiently motivated voter base could certainly build up a meaningful competitor to TSMC over the course of, say, a decade or two (or three...). But it would require sustained investment at all levels—and not just investment in the simple financial sense; it requires people investing their time in education and research. Dedicating their lives to making the best chips in the world. And the only reason that would work is that it defies our system, and chooses to invest in plants that won't be finished for years, and then pay for chips that they know are inferior in quality, because they're our chips, and paying for them when they're lower quality is the only way to get them to be the best chips in the world.
- WesolyKubeczek 6mo agoI fear that the real reason we do have a shortage, I mean, the real reason for the demand, is AI companies scooping what they can so that their competitors, whether existing or incumbent, can’t get to it.
- tuetuopay 6mo agoThis was one of the theories behind the wafer buyout by OpenAI indeed. Pretty efficient way to make everyone panic and cut off of new hardware.
- WesolyKubeczek 6mo agoWas it debubked in any way (e.g. by OpenAI actually showing what they do with the wafers?)
- stuxnet79 6mo agoOk so Samsung, SK Hynix and Micron do not have the capacity to meet demand. Also, what little capacity they do have they are allocating to HBM over DRAM. Based on my limited knowledge HBM can not be easily repurposed for consumer electronics. Translation: main street is cooked for the next 3-4 years. It doesn't stop there though. OpenAI is currently mired in a capital crunch. Their last round just about sucked all the dry powder out of the private markets. Folks are now starting to ask difficult questions about their burn rate and revenue. It is increasingly looking like they might not commit to the purchase order they made which kick-started this whole panic over RAM. Soo ... how sure are we that the memory makers themselves are not going to be the ones holding the bag?
- mschuster91 6mo ago> Soo ... how sure are we that the memory makers themselves are not going to be the ones holding the bag? We aren't. The remaining memory manufacturers fear getting caught in a "pork cycle" yet again - that is why there's only the three large ones left anyway.
- ahartmetz 6mo agoIf they don't expand capacity much, the only negative consequences I foresee happening for them is that they might lose spending discipline, and that systems will be set up to make do with a little less memory. Apart from that, it's just very high profits followed by more or less regular profits.
- XorNot 6mo agoThey could wind up losing all their business to China though. China has memory makers who are creeping up through the stages of production maturity, and once they hit then there's no going back. If the existing makers can't meet supply such that Chinese exports get their foot in the door, they may find they never get ahead again due to volume - that domestic market is huge so they have scale, and the gaming market isn't going to care because they get anything at the moment, which is all you'll need for enterprise to say "are we really afraid of memory in this business?"
- rzmmm 6mo agoIt seems that RAM manufacturers are still reluctant to increase production. They know something that investors don't about long term RAM demands?
- tomaskafka 6mo agoThey are not happy to invest tens of billions for new capacities for Sam Altman’s “I will surely buy once I have money, pinky promise”.
- danaris 6mo agoThe same thing everyone who's paying attention to the real world (and not the financial fantasy world) does: that OpenAI's purchase commitments are wildly unrealistic and unsustainable.
- hsbauauvhabzb 6mo agoWhat’s the lose scenario for them? They’re basically a cartel, and you need ram irregardless. If they make less it’s still a cost:demand, just not the most optimal for them. They’ve done that math, and figure this is the best risk and reward for them. Your goodwill or opinion doesn’t matter to them, because you need them more than they need you.
- consp 6mo ago> They’re basically a cartel, The lawsuits in the past prove that statement to not be basically but actually.
- Legend2440 6mo agoThey've been burned before. The DRAM industry has a long history of booms and busts. Demand increased, everyone built new fabs, then prices dropped and they couldn't pay off their investments. Many went out of business. It happened in the 80s, it happened in the 90s, it happened in the 2000s. Now there's only three manufacturers left, and they know very well that demand for their product tends to be cyclical.
- lizknope 6mo ago
- tomaytotomato 6mo agoI just checked my gaming PC I built a few years ago with 64GB of DDR5 RAM, its actually gone up in value, that is unheard of generally. Think I will scrap my PC and sell its parts. I wonder if there are any niche companies building decent rigs with DDR3 and 5/6th generation Intel CPUs out there, it is cheap and might be a business opportunity?
- HerbManic 6mo agoI am still running a DDR3 2nd Gen i7. 32GB RAM, it is surprisingly comfortable but I also dont push it too hard.
- edb_123 6mo agoSame. Core i7 2600K clocked to 4.4GHz with 32GB DDR3. It still does its job as my stationary DAW, and basically handles anything DAW-related I throw at it with ease. The only issue is its lack of AVX2 support, and since this is required by Ableton Live 12, I'll be stuck at Ableton 11 forever.
- theandrewbailey 6mo agoI work at an e-waste recycling company. I have several dozen trays of RAM in my inventory, ~90% of it DDR3. DDR3 was selling as of a month ago, but I haven't tried to sell any RAM since. I'm looking forward to doing a huge one this week.
- 0xDEFACED 6mo agodo you have an online storefront?
- theandrewbailey 6mo agoWe have an Ebay store: https://www.ebay.com/str/evolutionecycling https://www.ebay.com/str/evolutionecycling
- Gud 6mo agoThank god they shut down 3D XPoint.
- chintech2 6mo agoI'm a bit surprised the article makes no mention of China's new memory companies. [0] https://techwireasia.com/2026/04/chinese-memory-chips-ymtc-cxmt-nand-dram-expansion/ https://techwireasia.com/2026/04/chinese-memory-chips-ymtc-c...
- jeroenhd 6mo agoAs the article states: >CXMT still trails Samsung, SK Hynix, and Micron by approximately three years in advanced DRAM node development, and yield rates on new production lines remain the variable that determines whether capacity targets translate into reliable supply. Liu notes that lines launched in the second half of 2026 are unlikely to change the global supply-demand balance until 2027. The Verge article talks about demand exceeding supply in 2028. Your article suggests it'll take until 2029 before Chinese production catches up to current technology. It'll help drive prices down in five yearss, but the Chinese memory production won't be ready and efficient enough to prevent the shortages from continuing to grow.
- ghighi7878 6mo agoMost people don't need current tech. Ddr4 is good enough
- u8080 6mo agoYou could buy CXMT DDR5 modules like, right now.
- Hamuko 6mo agoI'm personally hoping that one of the AI or data center companies is suddenly unable to pay for their bills and deflate the entire industry. Probably the only hope of things getting better before the 2030s.
- tuetuopay 6mo agoThat’s likely to happen if all the talks about OpenAI pulling out of their wafer deals are true.
- lousken 6mo agoIf only we have not allowed oligopolies to exist. Meanwhile, EU is not in the race at all and US has very few fabs.
- wmf 6mo agoPeople don't want to pay more every day so they can pay less in an emergency.
- sph 6mo agoI fear the author and most commenters are not aware of the law of demand and supply. If there is demand for consumer RAM, there will be supply for consumer RAM. It just takes time and risk-assessment to scale up operations. We have RAM shortage now, we will have very cheap RAM tomorrow. It’s not like production is bottlenecked by raw materials. Chip companies just need to assess if the demand by AI companies will last so it’s better to scale up, or perhaps they should wait it out instead of oversupplying and cutting into their profits.
- rt56a 6mo agoWe're talking about advanced semiconductor manufacture. It takes years and 100s millions to billions of dollars to scale up operations. That's something you don't do unless you know there's demand to sustain it in future.
- eulgro 6mo agoThe law of supply and demand works in a perfect competition market. There are two RAM suppliers...
- eatsyourtacos 6mo ago>I fear the author and most commenters are not aware of the law of demand and supply I cannot stand how you and people like you try to justify everything by supply and demand. Also you act like it's some natural law of nature. It's not a law of nature- if you took an economics class you would realize it's try to maximize PROFIT. It's not for the good of the people. All of these things are a CHOICE that people are making to now completely screw the average person for, again, the needs of big corporations and the top 0.01%.
- shevy-java 6mo agoI want those AI companies that drove the prices up, to pay an immediate back-tax to all of us. I don't want to pay more because of AI companies driving the price up. That is milking.
- WhereIsTheTruth 6mo agoFabricated shortage to fasten US Chip Act and US Chip Security Act
- jmyeet 6mo agoThis is simple extrapolation from current demand, nothing more. And that's a borderline silly analysis because it assumes the AI bubble won't burst. The great misadventure in the Persian Gulf probably accelerates that because we're almost certainly going to be facing a recession. Another thing I've been thinking about is what happens when the next generation of NVidia chips comes out? I suspect NVidia is going to delay this to milk the current demand but at some point you'll be able to buy something that's better than the H100 or B200 or whatever the current state-of-the-art for half the price. And what's that going to do to the trillions in AI DC investment? I'm interested when the next bump in DRAM chip density is coming. That's going to change things although it seems like much of production has moved from consumer DRAM chips to HBM chips. So maybe that won't help at all. I do think that companies will start seeing little ot no return from billions spent on AI and that's going to be aproblem. I also think that the hudnreds of billions of capital expenditure of OpenAI is going to come crashing down as there just isn't any even theoretical future revenue that can pay for all that.
- wmf 6mo agosomething that's better than the current state-of-the-art for half the price. And what's that going to do to the trillions in AI DC investment? They'll just spend whatever they were planning to spend and get more performance.
- cbdevidal 6mo agoI’m a bit of an optimist. I think this will smack the hands of developers who don’t manage RAM well and future apps will necessarily be more memory-efficient.
- pron 6mo agoUsing a lot less RAM often implies using more CPU, so even with inflated RAM prices, it's not a good tradeoff (at least not in general).
- IsTom 6mo agoOr just using less electron and writing less shit code.
- DimmieMan 6mo agoOnly if the software is optimised for either in the first place. Ton of software out there where optimisation of both memory and cpu has been pushed to the side because development hours is more costly than a bit of extra resource usage.
- codebje 6mo agoThat'll stay true for consumer software, because the cost for extra resource usage is not borne by the development house.
- zamadatix 6mo agoThe tradeoff has almost exclusively been development time vs resource efficiency. Very few devs are graced with enough time to optimize something to the point of dealing with theoretical tradeoff balances of near optimal implementations.
- pron 6mo agoThat's fine, but I was responding to a comment that said that RAM prices would put pressure to optimise footprint. Optimising footprint could often lead to wasting more CPU, even if your starting point was optimising for neither.
- coldtea 6mo agoExpect shortages across the board. RAM? That's the tip of the iceberg, think food and gas.
- black_13 6mo ago[dead]
- 1o1o1o1o1 6mo agoHilarious, The ram in my PC i built 5 years ago is will soon be worth more than i spent on building the whole PC.
- aidenn0 6mo agoI was about to give away my old PC, but I think it could be worth my hassle to sell it for the RAM now (64GB DDR4).
- cozzyd 6mo agoI bought a workstation with 3 TB of ram for FDTD simulations last year. Glad I got it then ...
- senfiaj 6mo agoI wonder if this might motivate to write more memory efficient software. I mean we have so much memory, but even some trivial programs eat hundreds of megabytes of ram.
- phamilton 6mo agoI've definitely done some vibe-coding with the explicit intent to reduce memory usage.
- senfiaj 6mo agoHow efficient is AI at reducing RAM consuption?
- dankwizard 6mo agoThis feels like an oxymoron
- fsckboy 6mo agoif a shortage lasts years, it's not a shortage. "The market clearing price of RAM in the face of expected sustained healthy demand should lead to a stable market for years." even if gaming is and will remain very popular for years, it and the desire to upgrade gaming rigs is still a discretionary activity with more price elasticity of demand than corporate uses for RAM in the dawn of the AI age. gamers live on the margin of this market, where low prices will stimulate upgrades and high prices will lead to holding out. The complaints about price are real, but that segment of the market is some combination of less large and less important.
- ls612 6mo agoThe issue is supply is inelastic so even as prices soar they can only make more so fast and that won’t get fixed until 2028.
- philistine 6mo agoWhy are you only talking about gamers? Apple, the most cautious planners in the whole industry have straight up cancelled their 512gb RAM Mac Studio. Don’t ask; they won’t sell you one. Everybody’s getting pinched, not just the gamers.
- shash 6mo agoIt’s not merely a “gaming vs data center“. There’s so many other places DRAM and NVM are needed - mobile, automotive, other consumer electronics,… the current situation is that _all_ of that is deprived of the memory that it needs. And much of this is critical to the real economy.
- fsckboy 6mo agoi could have used lemons and lemonade stands to explain supply and demand, the lesson is still the same. letting the market set prices ensures that the chips go to the critical markets and uses. less critical uses will not allocate funds for purchases.
- mongrelion 6mo ago
- BirAdam 6mo agoOf course, alternatively, the AI companies could go bust before finding profitability. Then, there’d be a ton of supply, prices would crash, and one or two of the current memory suppliers would go out of business. After that, the new Chinese memory companies might be producing at volume, and Renesas could be up and running. At the moment, nothing is certain. Could this last? Sure. Could it not last? Yup.
- bschwindHN 6mo agoBut thank god we were all able to generate some SVGs of pelicans, right guys?
- marcus_holmes 6mo agoThis could be great. There's a future where RAM makers tool up for this massively increased demand, then the AI companies go broke as the bubble bursts, so RAM is cheap as. So laptop manufacturers get on that and start making laptops with 1TB+ memory so we can run decent LLMs on the local machine. Everyone happy :)
- wao0uuno 6mo agoRAM makers are not increasing their capacity. If AI bubble bursts we might see a momentary drop in RAM prices but it won't be dramatic. Return to "normal" is the best scenario I can imagine but my gut tells me we're probably never going back to early 2025 memory prices.
- thijson 6mo agoI've read that the chip manufacturers are looking into high bandwidth flash for on package storage of ai models. That would solve some of the cost issue, flash is significantly cheaper than dram.
- nilkn 6mo agoAs an aside, recently I wanted to refresh my gaming PC, but the price shock and general lack of availability of buying components individually made it seem hardly worth it, so I just kept deferring the project. Then, mostly by chance, I saw that my local Microcenter had some pre-builts for sale, and I ended up picking one up for <$5k that had "best in slot" components across the board, including a 5090 and even a high-end power supply. The last time I built a gaming PC was upwards of a decade ago, and at that time the prevailing wisdom was to never buy a pre-built unless you had a massive amount of disposable income and couldn't spare even just one weekend to dedicate to a hobby project that could benefit you for years. Now, it was absolutely a no-brainer.
- Mawr 6mo ago> and at that time the prevailing wisdom was to never buy a pre-built That's still the case, and always will be — with a pre-built you're at the very least paying for someone to assemble it for you, so it's always going to be more expensive as a baseline. Beyond that, the chance they've chosen good components and haven't tried to screw you over on less flashy ones like the motherboard and power supply is low. That's not to say it's literally impossible to ever find a good deal. You very well might have. Doesn't change anything though.
- irishloop 6mo ago> with a pre-built you're at the very least paying for someone to assemble it for you, so it's always going to be more expensive as a baseline Except isn't it possible that pre-built companies actually get better deals on hardware bought in bulk, and therefore could offset the labor costs with cheaper materials?
- nilkn 6mo agoI believe this is exactly what's going on -- they're buying parts in bulk, often months in advance, and locking in deals that a single consumer can't easily go get on the open market right now. Hardware pricing and availabilty pre-COVID was pretty predictable and stable, which meant the consumer could extract a meaningful cost advantage if they were willing to do the relatively modest amount of work of sourcing components individually and personally assembling the build. Right now, though, some places like Microcenter appear to have a cost advantage that fundamentally relies on market and pricing instability and can only be achieved through deeper integration with the supply chain and bulk purchasing in advance -- something a retailer like Microcenter can do, but I personally cannot.
- vectorhacker 6mo agoIt sounds to me like an incentive for new companies to make RAM.
- alprado50 6mo agoIm thankfull for buying 16gb of RAM, but what is gonna happen in 5 years when users PCs start to fail?
- fuzzy2 6mo agoWhat do you mean, in 5 years? It's not like everyone just bought a new computer. My gut says it's exactly the other way around: most computers are old. They may fail as soon as today. All computers in my household are 8+ years old.
- wao0uuno 6mo agoThere is enough older hardware floating around to last us for decades. You don't need a gaming rig to do 99% of your computing (excluding gaming obviously). Also computers don't really just break. It's mostly the disks that wear out and PSUs that age.
- LastTrain 6mo agoSomething I haven’t been able to reconcile: If AI makes software easier to create, that will drive the price down. How are software companies going to make enough revenue to pay for AI, when the amount of money being spent on AI is already multiples of the current total global expenditure on software? This demand for RAM is built on a foundation of sand, there will be a glut of capacity when it all shakes out.
- locknitpicker 6mo ago> If AI makes software easier to create, that will drive the price down. Supposedly AI drives down the cost of producing software,not the "price". > How are software companies going to make enough revenue to pay for AI, when the amount of money being spent on AI is already multiples of the current total global expenditure on software? Currently, the cost of AI is between $20/month and around $200/month per developer. I think the huge billions you're seeing in the news are the investment cost on AI companies, who are burning through cash to invest in compute infrastructure to allow both training and serving users. > This demand for RAM is built on a foundation of sand, there will be a glut of capacity when it all shakes out. Who knows? What I know is that I need >64GB of RAM to run local models, and that means most people will need to upgrade from their 8Gb/16GB setup to do the same. Graphics cards follow mostly the same pattern.
- zozbot234 6mo ago> Who knows? What I know is that I need >64GB of RAM to run local models, and that means most people will need to upgrade from their 8Gb/16GB setup to do the same. Graphics cards follow mostly the same pattern. Depends how big the models are, how fast you want them to run and how much context you need for your usage. If you're okay with running only smaller models (which are still very capable in general, their main limitation is world knowledge) making very simple inferences at low overall throughput, you can just repurpose the RAM, CPUs/iGPUs and storage in the average setup.
- adrian_b 6mo agoYou need >64 GB of DRAM to run local models fast. You can run huge local models slowly with the weights stored on SSDs. Nowadays there are many computers that can have e.g. 2 PCIe 5.0 SSDs, which allow a reading throughput of 20 to 30 gigabyte per second, depending on the SSDs (or 1 PCIe 5.0 + 1 PCIe 4.0, for a throughput in the range 15-20 GB/s). There are still a lot of improvements that can be done to inference back-ends like llama.cpp to reach the inference speed limit determined by the SSD throughput. It seems that it is possible to reach inference speed in the range from a few seconds per token to a few tokens per second. That may be too slow for a chat, but it should be good enough for an AI coding assistant, especially if many tasks are batched, so that they can progress simultaneously during a single read pass over the SSD data.
- zizheruan 6mo agoSad news, I didn't buy enough RAM before....
- ares623 6mo agoAre we entering the Reverse-Moore's Law era.
- 5255652 6mo agoCan we stop advertising paid Blog/News websites we can't read without a subscription.
- librasteve 6mo agoThe RAM market is a square wave
- onchainintel 6mo agoYour instincts are likely right on this one OP. Memory prices surged 80–90% in Q1 2026 compared to Q4 2025, DRAM, NAND, and HBM all at record highs. 3 suppliers for the entire planet?
- thelastgallon 6mo agoIt will last forever. After covid, all manufacturers understood the value of limiting supply and extracting profits. Cars used to super cheap before covid, they will never go back to the same levels. From now on, RAM will always be super costly for consumers, because they can't make massive deals like Apple/OpenAI/etc. We are the bagholders.
- ymolodtsov 6mo agoOf course, clearly businesses were just stupid before and only learned how to make revenue in 2021.
- energy123 6mo agoWhen lithium prices decreased over 80% from 2022 to 2025, it was because lithium miners felt altruistic. Car manufacturers were feeling greedy. This is how bad the thinking has gotten.
- ponector 6mo ago>> Cars used to super cheap before covid Have they really ever been cheap? Also Tesla 3 is cheaper now, Yaris is still cheap as well.
- energy123 6mo agoCovid inflation was because of supply chain disruptions, loose fiscal policy (like Biden's ARP which a Central Bank analysis said added a few % to annual inflation), and money supply expansions. There was less goods and more money. When you go and trade money for goods, it should be obvious what happens.
- Legend2440 6mo agoIt won't. DRAM prices are cyclical. They were super high in 2020, then demand crashed in 2022 to the point that manufacturers couldn't sell all their inventory. Now it's high again, but give it a couple years and it'll once again crash.
- Chrisszz 6mo ago[dead]
- p0w3n3d 6mo agoWe have the saying in my country: The days of things being cheap are over.
- cylemons 6mo agoWhen I personally use chatgpt and friends, I am not seeing any slowdowns or anything, meaning that their servers can handle the loads just fine. So then, why are these companies spending so much building new capacity if the current capacity is enough?
- goldenarm 6mo agoFrontier labs flagship models are ~2T params at the moment, but they intend to ship 10T models like Claude Mythos, which would require substantial datacenter expansion. Same thing for training.
- rldjbpin 6mo agolet the analyst and news say what they want - the entire situation is artificial and is up to the manufacturers. the current relative spike in the prices misses the medium-term trend of the vast decrease in memory price post-covid that led to the recent surge. the cartel got another opportunity to make bank and they will use that lever to the max. funnily enough i've been personally stuck with 16 gigs since 2015, across three memory generations! but i am used to the past when you would spend 80-100 on an 8gb stick (jdec timings, nothing fancy but from a major brand) without accounting for inflation.
- cicko 6mo agoUsually, right after articles like this, things come crashing down.