6 ms·
I think the OpenAI deal to lock wafers was a wonderful coup. OpenAI is more and more losing ground against the regularity[0] of the improvements coming from Ant
by Loic 10mo ago
I think the OpenAI deal to lock wafers was a wonderful coup. OpenAI is more and more losing ground against the regularity[0] of the improvements coming from Anthropic, Google and even the open weights models. By creating a chock point at the hardware level, OpenAI can prevent the competition from increasing their reach because of the lack of hardware.
[0]: For me this is really an important part of working with Claude, the model improves with the time but stay consistent, its "personality" or whatever you want to call it, has been really stable over the past versions, this allows a very smooth transition from version N to N+1.
- codybontecou 10mo agoThis became very clear with the outrage, rather than excitement, of forcing users to upgrade to ChatGPT-5 over 4o.
- Grosvenor 10mo agoCould this generate pressure to produce less memory hungry models?
- hodgehog11 10mo agoThere has always been pressure to do so, but there are fundamental bottlenecks in performance when it comes to model size. What I can think of is that there may be a push toward training for exclusively search-based rewards so that the model isn't required to compress a large proportion of the internet into their weights. But this is likely to be much slower and come with initial performance costs that frontier model developers will not want to incur.
- UncleOxidant 10mo agoOr maybe models that are much more task-focused? Like models that are trained on just math & coding?
- agoodusername63 10mo agoisn't that what the mixture of experts trick that all the big players do is? Bunch of smaller, tightly focused models
- irthomasthomas 10mo agoNot exactly. MoE uses a router model to select a subset of layers per token. This makes them faster but still requires the same amount of RAM.
- Grosvenor 10mo agoYeah that was my unspoken assumption. The pressure here results in an entirely different approach or model architecture. If openAI is spending $500B then someone can get ahead by spending $1B which improves the model by >0.2% I bet there's a group or three that could improve results a lot more than 0.2% with $1B.
- parineum 10mo ago> so that the model isn't required to compress a large proportion of the internet into their weights. The knowledge compressed into an LLM is a byproduct of training, not a goal. Training on internet data teaches the model to talk at all. The knowledge and ability to speak are intertwined.
- thisrobot 10mo agoI wonder if this maintains the natural language capabilities which are what LLM's magic to me. There is a probably some middle ground, but not having to know what expressions, or idiomatic speech an LLM will understand is really powerful from a user experience point of view.
- jiggawatts 10mo ago> exclusively search-based rewards so that the model isn't required to compress a large proportion of the internet into their weights. That just gave me an idea! I wonder how useful (and for what) a model would be if it was trained using a two-phase approach: 1) Put the training data through an embedding model to create a giant vector index of the entire Internet. 2) Train a transformer LLM but instead only utilising its weights, it can also do lookups against the index. Its like a MoE where one (or more) of the experts is a fuzzy google search. The best thing is that adding up-to-date knowledge won’t require retraining the entire model!
- lofaszvanitt 10mo agoOf course and then watch those companies reined in.
- hodgehog11 10mo agoI don't see this working for Google though, since they make their own custom hardware in the form of the TPUs. Unless those designs include components that are also susceptible?
- frankchn 10mo agoTPUs use HBM, which are impacted.
- bri3d 10mo agoStill susceptible, TPUs need DRAM dies just as much as anything else that needs to process data. I think they use some form of HBM, so they basically have to compete alongside the DDR supply chain.
- UncleOxidant 10mo agoEven their TPU based systems need RAM.
- jandrese 10mo agoThat was why OpenAI went after the wafers, not the finished products. By buying up the supply of the raw materials they bottleneck everybody, even unrelated fields. It's the kind of move that requires a true asshole to pull off, knowing it will give your company an advantage but screw up life for literally billions of people at the same time.
- mbesto 10mo ago> By buying up the supply We actually don't know for certain whether these agreements are binding. If OpenAI gets in a credit crunch we'll soon find out.
- agoodusername63 10mo agoWent after the right component too. RAM manufacturers love an opportunity to create as much scarcity as possible.
- hnuser123456 10mo agoSure, but if the price is being inflated by inflated demand, then the suppliers will just build more factories until they hit a new, higher optimal production level, and prices will come back down, and eventually process improvements will lead to price-per-GB resuming its overall downtrend.
- malfist 10mo agoMicron has said they're not scaling up production. Presumably they're afraid of being left holding the bag when the bubble does pop
- Analemma_ 10mo agoNot just Micron, SK Hynix has made similar statements (unfortunately I can only find sources in Korean). DRAM manufacturers got burned multiple times in the past scaling up production during a price bubble, and it appears they've learned their lesson (to the detriment of the rest of us).
- fullstop 10mo agoWhy are they building a foundry in Idaho? https://www.micron.com/us-expansion/id https://www.micron.com/us-expansion/id
- delfinom 10mo agoFuture demand aka DDR6. The 2027 timeline for the fab is when DDR6 is due to hit market.
- roboror 10mo agoI mean it says on the page >help ensure U.S. leadership in memory development and manufacturing, underpinning a national supply chain and R&D ecosystem. It's more political than supply based
- mindslight 10mo agoHedging is understandable. But what I don't understand is why they didn't hedge by keeping Crucial around but more dormant (higher prices, less SKUs, etc)
- lysace 10mo agoPlease explain to me like I am five: Why does OpenAI need so much RAM? 2024 production was (according to openai/chatgpt) 120 billion gigabytes. With 8 billion humans that's about 15 GB per person.
- mebassett 10mo agolarge language models are large and must be loaded into memory to train or to use for inference if we want to keep them fast. older models like gpt3 have around 175 billion parameters. at float32s that comes out to something like 700GB of memory. newer models are even larger. and openai wants to run them as consumer web services.
- lysace 10mo agoI mean, I know that much. The numbers still don't make sense to me. How is my internal model this wrong? For one, if this was about inference, wouldn't the bottleneck be the GPU computation part?
- ssl-3 10mo agoConcurrency? Suppose some some parallelized, distributed task requires 700GB of memory (I don't know if it does or does not) per node to accomplish, and that speed is a concern. A singular pile of memory that is 700GB is insufficient not because it lacks capacity, but instead because it lacks scalability. That pile is only enough for 1 node. If more nodes were added to increase speed but they all used that same single 700GB pile, then RAM bandwidth (and latency) gets in the way.
- Chiron1991 10mo agoThis "memory shortage" is not about AI companies needing main memory (which you plug into mainboards), but manufacturers are shifting their production capacities to other types of memory that will go onto GPUs. That brings supply for other memory products down, increasing their market price.
- daemonologist 10mo ago
- Phelinofist 10mo ago> By creating a chock point at the hardware level, OpenAI can prevent the competition from increasing their reach because of the lack of hardware I already hate OpenAI, you don't have to convince me
- bakugo 10mo agoIs anyone else deeply perturbed by the realization that a single unprofitable corporation can basically buy out the entire world's supply of computing hardware so nobody else can have it? How did we get here? What went so wrong?
- zozbot234 10mo agoThey're simply making a bet that they can put the DRAM dies to more valuable use than any of the existing alternatives, including e.g. average folks playing the latest videogames on their gaming rig. At this kind of scale, they had better be right or they are toast: they have essentially gone all-in on their bet that this whole AI thing is not going to 'pop' anytime soon.
- bakugo 10mo ago> They're simply making a bet that they can put the DRAM dies to more valuable use than any of the existing alternatives They can't. They know they can't. We all know they can't. But they can just keep abusing the infinite money glitch to price everyone else out, so it doesn't matter.
- fullstop 10mo agoPerhaps ChatGPT has given them instructions.
- sophrosyne42 10mo agoWhen they find out that it is not, in fact, an infinite money glitch, they're going to have to eat that cost. It will work out great for everyone as long as they aren't bailed out.
- zozbot234 10mo agoIt's more like a waste-infinite-money glitch, if that's what they're trying. There's no way that a simple speculative attack actually makes DRAM more valuable in the long term on its own, and that's the only win condition for that kind of play. People have tried to hoard all sorts of commodities as a mere speculative play on the market, and it never works.
- beAbU 10mo agoI'm not too keyed into the economics of this supposed AI bubble, but is this not an unfathomably risky move on OpenAI's part? If this thing actually pops, or a competitor like Google actually pulls ahead and comes out victorious, then OpenAI will sit holding a very expensive bag of expensive but unusable raw materials that they'll have to sell of at a discount?