7 ms·
NVIDIA sells shovels to the gold rush. One miner (Liang Wenfeng), who has previously purchased at least 10,000 A100 shovels... has a "side project" where they f
by ozten 2y ago
NVIDIA sells shovels to the gold rush. One miner (Liang Wenfeng), who has previously purchased at least 10,000 A100 shovels... has a "side project" where they figured out how to dig really well with a shovel and shared their secrets.
The gold rush, wether real or a bubble is still there! NVIDA will still sell every shovel they can manufacture, as soon as it is available in inventory.
Fortune 100 companies will still want the biggest toolshed to invent the next paradigm or to be the first to get to AGI.
- culi 2y agoYeah but NVIDIA's amazing digging technique that could only be accomplished with NVIDIA shovels is now irrelevant. Meaning there are more people selling shovels for the gold rush
- ozten 2y agoCUDA begs to differ.
- dragonwriter 2y agoWhat about DeepSeek negates NVidia’s advantages over other GPU vendors?
- koolhead17 2y agoWhat is stopping huawei or other Chinese vendors to make chips on deepseek specification and 1/10th NVIDIA cost and mass market it?
- shawabawa3 2y ago> What is stopping huawei or other Chinese vendors to make chips on deepseek specification What is "deepseek specification"? Deepseek was trained on NVDA chips. If chinese vendors could build chips as good as NVDA it wouldn't have such a dominant position already, that hasn't changed
- culi 2y agoYou can train, or at least run, llms on intel and less powerful chips
- dragonwriter 2y ago> You can train, or at least run, llms on intel and less powerful chips The claimed training breakthrough is an optimization targeting NVidia chip, not something that reduces NVidia's relative advantage. Even if it is easily generalizable to other vendors hardware, it doesn't reduce NVidia's advantage over other vendors, it just proportionately scales down the training requirements for a model of a given capacity. Which, maybe, very short term reduces demands from the big existing incumbents, but it also increases the number of players for which investing in GPUs for model training at all is worthwhile, increasing aggregate demand.
- culi 2y agoIt's not an optimization targeting Nvidia chips. It's an optimization of the technique through and through regardless of chip But your point is well taken and perhaps both mine and GP's metaphors break down. Either way, we saw massive spikes in demand for Nvidia when crypto mining became huge followed by a massive drop when we hit the crypto winter. We saw another massive spike when LLMs blew up and this may just be the analogous drop in demand for LLMs
- Tostino 2y agoYou both seem to be talking past each other. There were a number of optimizations that made this possible. Some were with the model itself and are transferable, others are with the training pipeline and specific to the Nvidia hardware they trained on.
- sho_hn 2y agoDeepSeek's stuff is actually more dependent on nVidia shovels. They implemented a bunch of assembly-level optimizations below the CUDA stack that allowed them to efficiently use the H800s they have, which are memory-bandwidth-gimped vs. the H100s they can't easily buy on the open market. That's cool, but doesn't run on any other GPUs. Cue all of China rushing to Jensen to buy all the H800s they can before the embargo gets tightened, now that their peers have demonstrated that they're useful for something. At least briefly, Jensen's customer audience increased.
- yieldcrv 2y agoI was thinking about that, but don’t those same optimizations work on H100s? and the concepts work on every other chip from Nvidia and every other manufacturer’s chip I still think this is bullish: more people will be buying chips once cheaper and more accessible, and the things the will be training with be 1,000% to 10,000% larger
- jjk166 2y agoProbably possible is nothing compared to already implemented. How long will it take to apply those concepts to other chips? Will they also be made available to the degree DeepSeek has been? By the time those alternatives are implemented how much further improvement will be made on Nvidia chips? Worst case scenario someone implements and open sources these optimizations for a competitor's chip basically immediately in which case the competitive landscape remains unchanged, for all other scenarios this is a first mover advantage for Nvidia.
- elihu 2y agoJevon's paradox would imply that there's good reason to think that demand for shovels will increase. AI doesn't seem to be one of those things where society as a whole will say, "we have enough of that; we don't need any more". (Many individual people are already saying that, but they aren't the people buying the GPUs for this in the first place. Steam engines weren't universally popular either when they were introduced to society.)
- int_19h 2y agoThe other thing is that if this pushes the envelope further on what AI models can do given a certain hardware budget, this might actually change minds. The pushback against generative AI today is that much of it is deployed in ways that are ultimately useless and annoying at best, and that in turn is because the capabilities of those models are vastly oversold (including internally in companies that ship products with them). But if a smarter model can actually e.g. reliably organize my email, that's a very different story.
- jcgrillo 2y ago> AI doesn't seem to be one of those things where society as a whole will say, "we have enough of that; we don't need any more". Really? Has anyone made a useful, commercially successful product with it yet?
- mrbungie 2y ago
- addicted 2y agoWhat you are missing is that it turns out the gold isn’t actually gold. It’s bronze. So earliest, the shovelers were willing to spend thousands of dollars for a single shovel because they were expecting to get much more valuable gold out the other end. But now that it’s only bronze, they can’t spend that much money on their tools anymore to make their venture profitable. A lot of shovelers are gonna drop out of the race. And the ones that remain will not be willing to spend as much. The fact that there isn’t that much money to be made in AI anymore means that whatever percentage of money would have gone to NVIDIA from the total money to be made in AI will now shrink dramatically.
- deeviant 2y agoWait, so AI might become 25x cheaper to train and run, and your thesis is... no one will make money on AI now?!?!
- h0l0cube 2y agoPerhaps they mean there's less wealth to be extracted from the closed-source training side of the equation, which requires huge capital investment, and promises even bigger returns by gatekeeping the technology.
- Tostino 2y agoLet's shed a tear for the investor class who had their wealth extraction dreams dashed a bit today. Anyways, where were we...
- hattmall 2y agoSort of, it means there's less of a chance for massive market domination and a monopoly.
- Ekaros 2y agoMany discussed aspects are disconnected. Cost of training, cost of hardware(and margin there), cost of operation, possible use cases, and then finally demand. Cheaper training still expect there is some use case for those trained models. There might or might not be. It can very well be that cost of training did not really limit the number of usable models.
- UncleOxidant 2y agoAGI would be El Dorado in this analogy?
- buryat 2y agoEl Dorado will be AGI with Humanoid robots and actual real pick-axes
- deleted 2y ago[deleted]
- avs733 2y agoThe thing with a gold rush is you often end up selling shovels after the gold has run out, but no one knows that until hindsight. There will probably be a couple scares that the gold has run out first to. And again the difference is only visible in hindsight.
- medion 2y agoCan anyone comment on why Wenfeng shared his secret sauce? Other than publicity, there only seems to be downsides for him, as now everyone else with larger compute will just copy and improve?
- seydor 2y agoIs AI expanding horizontally or vertically? My understanding is that smarter models dominate over hordes of dumber ones
- bodegajed 2y agoThe gold rush is over because pre-trained models don't improve as much anymore. The application layer has massive gains in cost-to-value performance. We also gain more trust from the consumer as models don't hallucinate as much. This is what DeepSeek R1 has shown us. As Ilya Sutskever said, pre-training is now over. We now have very expensive Nvidia shovels that use a lot of power but do very little improvement to the models.
- onlyrealcuzzo 2y agoThis can all be true, and Nvidia's market cap can still go down A TON. Nvidia's market cap is based on extreme margins and absurd growth for 10 years. If either of those nobs get turned down a little, there can be a MASSIVE hit to the valuation - which is what happened.