29 ms·
The impact of competition and DeepSeek on Nvidia
- jamalms 2y ago[flagged]
- jamalms 2y ago[flagged]
- eigenvalue 2y agoYesterday I wrote up all my thoughts on whether NVDA stock is finally a decent short (or at least not a good thing to own at this point). I’m a huge bull when it comes to the power and potential of AI, but there are just too many forces arrayed against them to sustain supernormal profits. Anyway, I hope people here find it interesting to read, and I welcome any debate or discussion about my arguments.
- patrickhogan1 2y agoGood article. Maybe I missed it, but I see lots of analysis without a clear concluding opinion.
- scsilver 2y agoWanted to add a preface: Thank you for your time on this article, I appreciate your perspective and experience, hoping you can help refine and reign in my bull case. Where do you expect NVDA's forward and current eps to land? What revenue drop off are you expecting in late 2025/2026. Part of my bull case for NVDA, continuing, is it's very reasonable multiple on insane revenue. An leveling off can be expected, but I still feel bullish on it hitting $200+ (5 Trillion market cap? on ~195B revenue for Fiscal year 2026 (calendar 2025) at 33 EPS) based on this years revenue according to their guidance and the guidance of the hyperscalers spending. Finding a sell point is a whole different matter to being actively short. I can see the case to take some profits, hard for me to go short, especially in an inflationary environment (tariffs, electric energy, bullying for lower US interest rates). The scale of production of Grace Hopper and Blackwell amaze me, 800k units of Blackwell coming out this quarter, is there even production room for AMD to get their chips made? (Looking at the new chip factories in Arizona) R1 might be nice for reducing llm inferencing costs, unsure about the local llama one's accuracy (couldnt get it to correctly spit out the NFL teams and their associated conferences, kept mixing NFL with Euro Football) but I still want to train YOLO vision models on faster chips like A100's vs T4 (4-5x multiples in speed for me). Lastly, if the Robot/Autonomous vehicle ML wave hits within the next year, (First drones and cars -> factories -> humanoids) I think this compute demand can sustain NVDA compute demand. The real mystery is how we power all this within 2 years... * This is not financial advice and some of my numbers might be a little off, still refining my model and verifying sources and numbers
- zippyman55 2y agoSo at some point we will have too many cannon ball polishing factories and it will become apparent the cannon ball trajectory is not easily improved on.
- j7ake 2y agoThis was an amazing summary of the landscape of ML currently. I think the title does the article injustice, or maybe it’s too long for people to read to appreciate it (eg the deepseek stuff can be an article within itself). Whatever the ones with longer attention span will benefit from this read. Thanks for summarising this up!
- eigenvalue 2y agoThanks! I was a bit disappointed that no one saw it on HN because I think they’d like it a lot.
- j7ake 2y agoI think they would like it a lot, but I think the title doesn’t match the content, and it takes too much reading before one realises it goes beyond the title. Keep it up!
- dang 2y agoWe've changed the title to a different one suggested by the author.
- metadat 2y agoThe site is currently offline, here's a snapshot: https://archive.today/y4utp https://archive.today/y4utp
- diesel4 2y agoLink isn't working. Is there another or a cached version?
- eigenvalue 2y agoTry again! Just rebooted the server since it’s going viral now.
- deleted 2y ago[deleted]
- OutOfHere 2y agoIt seems like a pointless discussion since DeepSeek uses Nvidia GPUs after all.
- jjeaff 2y agoit uses a fractional amount of GPUs though.
- breadwinner 2y agoAs it says in the article, you are talking about a mere constant of proportionality, a single multiple. When you're dealing with an exponential growth curve, that stuff gets washed out so quickly that it doesn't end up matter all that much. Keep in mind that the goal everyone is driving towards is AGI, not simply an incremental improvement over the latest model from Open AI.
- high_na_euv 2y agoWhy do you assume that exponential growth curve is real?
- ithkuil 2y agoWhich due to the Jevons Paradox may ultimately cause more shovels to be sold
- cma 2y agoTheir loss curve with the RL didn't level off much though, could be taken a lot further and scaled up to more parameters on the big nvidia mega clusters out there. And the architecture is heavily tuned to nvidia optimizations.
- UltraSane 2y agoJevons Paradox states that increasing efficiency can cause an even larger increase in demand.
- dutchbookmaker 2y ago
- arcanus 2y ago> Amazon gets a lot of flak for totally bungling their internal AI model development, squandering massive amounts of internal compute resources on models that ultimately are not competitive, but the custom silicon is another matter Juicy. Anyone have a link or context to this? I'd not heard of this reception to NOVA and related.
- simonw 2y agoI think Nova may have changed things here. Prior to Nova their LLMs were pretty rubbish - Nova only came out in December but seems a whole lot better, at least from initial impressions: https://simonwillison.net/2024/Dec/4/amazon-nova/ https://simonwillison.net/2024/Dec/4/amazon-nova/
- arcanus 2y agoThanks! That's consistent with my impression.
- snowmaker 2y agoThis is an excellent article, basically a patio11 / matt levine level breakdown of what's happening with the GPU market.
- lxgr 2y agoCouldn't agree more! If this is the byproduct, these must be some optimized Youtube transcripts :)
- eprparadox 2y agolink seems to be dead... is this article still up somewhere?
- jazzyjackson 2y agoIt's back up, but just in case: https://archive.is/y4utp https://archive.is/y4utp
- deleted 2y ago[deleted]
- eigenvalue 2y agoSorry, my blog crashed! Had a stupid bug where it was calling GitHub too frequently to pull in updated markdown for the posts and kept getting rate limits. Had to rewrite it but it should be much better now.
- breadwinner 2y agoGreat article but it seems to have a fatal flaw. As pointed out in the article, Nvidia has several advantages including: - Better Linux drivers than AMD - CUDA - pytorch is optimized for Nvidia - High-speed interconnect Each of the advantages is under attack: - George Hotz is making better drivers for AMD - MLX, Triton, JAX: Higher level abstractions that compile down to CUDA - Cerbras and Groq solve the interconnect problem The article concludes that NVIDIA faces an unprecedented convergence of competitive threats. The flaw in the analysis is that these threats are not unified. Any serious competitor must address ALL of Nvidia's advantages. Instead Nvidia is being attacked by multiple disconnected competitors, and each of those competitors is only attacking one Nvidia advantage at a time. Even if each of those attacks are individually successful, Nvidia will remain the only company that has ALL of the advantages.
- toisanji 2y agoI want the NVIDIA monopoly to end, but there is no real competition still. * George Hotz has basically given up on AMD: https://x.com/__tinygrad__/status/1770151484363354195 https://x.com/__tinygrad__/status/1770151484363354195 * Groq can't produce more hardware past their "demo". It seems like they haven't grown capacity in the years since they announced, and they switched to a complete SaaS model and don't even sell hardware anymore. * I dont know enough about MLX, Triton, and JAX,
- simonw 2y agoThat George Hotz tweet is from March last year. He's gone back and forth on AMD a bunch more times since then.
- gnlrtntv 2y ago> While Apple's focus seems somewhat orthogonal to these other players in terms of its mobile-first, consumer oriented, "edge compute" focus, if it ends up spending enough money on its new contract with OpenAI to provide AI services to iPhone users, you have to imagine that they have teams looking into making their own custom silicon for inference/training This is already happening today. Most of the new LLM features announced this year are primarily on-device, using the Neural Engine, and the rest is in Private Cloud Compute, which is also using Apple-trained models, on Apple hardware. The only features using OpenAI for inference are the ones that announce the content came from ChatGPT.
- simonw 2y ago"if it ends up spending enough money on its new contract with OpenAI to provide AI services to iPhone users" John Gruber says neither Apple nor OpenAI are paying for that deal: https://daringfireball.net/linked/2024/06/13/gurman-openai-apple https://daringfireball.net/linked/2024/06/13/gurman-openai-a...
- lxgr 2y agoMark Gurman (from Bloomberg) is saying that.
- uncletaco 2y agoWhen he says better linux drivers than AMD he's strictly talking about for AI, right? Because for video the opposite has been the case for as far back as I can remember.
- eigenvalue 2y agoYes, AMD drivers work fine for games and things like that. Their problem is they basically only focused on games and other consumer applications and, as a result, ceded this massive growth market to Nvidia. I guess you can sort of give them a pass because they did manage to kill their archival Intel in data center CPUs but it’s a massive strategic failure if you look at how much it has cost them.
- simonw 2y agoThis is excellent writing. Even if you have no interest at all in stock market shorting strategies there is plenty of meaty technical content in here, including some of the clearest summaries I've seen anywhere of the interesting ideas from the DeepSeek v3 and R1 papers.
- eigenvalue 2y agoThanks Simon! I’m a big fan of your writing (and tools) so it means a lot coming from you.
- punkspider 2y agoI was excited as soon as I saw the domain name. Even after a few months, this article[1] is still at the top of my mind. You have a certain way of writing. I remember being surprised at first because I thought it would feel like a wall of text. But it was such a good read and I felt I gained so much. 1: https://youtubetranscriptoptimizer.com/blog/02_what_i_learned_making_the_python_backend_for_yto https://youtubetranscriptoptimizer.com/blog/02_what_i_learne...
- eigenvalue 2y agoI really appreciate that, thanks so much!
- nejsjsjsbsb 2y agoI was put off by the domain by bias against something that sounds like a company blog. Especially a "YouTube something". You may get more milage from excellent writing on a yourname.com. This is a piece that sells you not this product, plus it feels more timeless. In 2050 someone my point to this post. Better if it were on your own name.
- eigenvalue 2y agoI had no idea this would get so much traction. I wanted to enhance my organic search ranking of my niche web app, not crash the global stock market!
- andrewgross 2y ago> The beauty of the MOE model approach is that you can decompose the big model into a collection of smaller models that each know different, non-overlapping (at least fully) pieces of knowledge. I was under the impression that this was not how MoE models work. They are not a collection of independent models, but instead a way of routing to a subset of active parameters at each layer. There is no "expert" that is loaded or unloaded per question. All of the weights are loaded in VRAM, its just a matter of which are actually loaded to the registers for calculation. As far as I could tell from the Deepseek v3/v2 papers, their MoE approach follows this instead of being an explicit collection of experts. If thats the case, theres no VRAM saving to be had using an MOE nor an ability to extract the weights of the expert to run locally (aside from distillation or similar). If there is someone more versed on the construction of MoE architectures I would love some help understanding what I missed here.
- Kubuxu 2y agoNot sure about DeepSeek R1, but you are right in regards to previous MoE architectures. It doesn’t reduce memory usage, as each subsequent token might require different expert buy it reduces per token compute/bandwidth usage. If you place experts in different GPUs, and run batched inference you would see these benefits.
- andrewgross 2y agoIs there a concept of an expert that persists across layers? I thought each layer was essentially independent in terms of the "experts". I suppose you could look at what part of each layer was most likely to trigger together and segregate those by GPU though. I could be very wrong on how experts work across layers though, I have only done a naive reading on it so far.
- rahimnathwani 2y agoI suppose you could look at what part of each layer was most likely to trigger together and segregate those by GPU though Yes, I think that's what they describe in section 3.4 of the V3 paper. Section 2.1.2 talks about "token-to-expert affinity". I think there's a layer which calculates these affinities (between a token and an expert) and then sends the computation to the GPUs with the right experts. This doesn't sound like it would work if you're running just one chat, as you need all the experts loaded at once if you want to avoid spending lots of time loading and unloading models. But at scale with batches of requests it should work. There's some discussion of this in 2.1.2 but it's beyond my current ability to comprehend!
- metadat 2y ago> Another very smart thing they did is to use what is known as a Mixture-of-Experts (MOE) Transformer architecture, but with key innovations around load balancing. As you might know, the size or capacity of an AI model is often measured in terms of the number of parameters the model contains. A parameter is just a number that stores some attribute of the model; either the "weight" or importance a particular artificial neuron has relative to another one, or the importance of a particular token depending on its context (in the "attention mechanism"). Has a wide-scale model analysis been performed inspecting the parameters and their weights for all popular open / available models yet? The impact and effects of disclosed inbound data and tuning parameters on individual vector tokens will prove highly informative and clarifying. Such analysis will undoubtedly help semi-literate AI folks level up and bridge any gaps.
- naveen99 2y agoDeepseek iOS app makes TikTok ban pointless.
- pavelstoev 2y agoInteresting take. They are now reading our minds vs looking at our kids and interiors.
- naveen99 2y agoyeah, what’s stopping zoom from integrating Deepseek and doing an end run around Microsoft teams.
- naveen99 2y agoI guess deepseek banned themselves from new signups from outside China…
- lxgr 2y agoMan, do I love myself a deep, well-researched long-form contrarian analysis published as a tangent of an already niche blog on a Sunday evening! The old web isn't dead yet :)
- eigenvalue 2y agoHah thanks, that’s my favorite piece of feedback yet on this.
- pavelstoev 2y agoEnglish economist William Stanley Jevons vs the author of the article. Will NVIDIA be in trouble because of DSR1 ? Interpreting Jevon’s effect, if LLMs are “steam engines” and DSR1 brings 90% efficiency improvement for the same performance, more of it will be deployed. This is not considering the increase due to <think> tokens. More NVIDIA GPUs will be sold to support growing use cases of more efficient LLMs.
- breadwinner 2y agoPart of the reason Musk, Zuckerberg, Ellison, Nadella and other CEOs are bragging about the number of GPUs they have (or plan to have) is to attract talent. Perplexity CEO says he tried to hire an AI researcher from Meta, and was told to ‘come back to me when you have 10,000 H100 GPUs’ See https://www.businessinsider.nl/ceo-says-he-tried-to-hire-an-ai-researcher-from-meta-and-was-told-to-come-back-to-me-when-you-have-10000-h100-gpus/ https://www.businessinsider.nl/ceo-says-he-tried-to-hire-an-...
- mrbungie 2y agoMaybe DeepSeek ain't it, but I expect a big "box of scraps"[1] moment soon. Constraint is mother of invention, and they are evading constraints with a promise of never-ending scale. [1] https://youtu.be/9foB2z_OVHc?si=eZSTMMGYEB3Nb4zI https://youtu.be/9foB2z_OVHc?si=eZSTMMGYEB3Nb4zI
- rat9988 2y agoThat's a weird way to read into it.
- TwoFerMaggie 2y agoThis reminds of the joke in physics, in which theoretical particle physicists told experimental physicists, over and over again, "trust me bro, standard model will be proved at 10x eV, we just need a bigger collider bro" after another world's biggest collider is built. Wondering if we are in a similar position with "trust me bro AGI will be achieved with 10x more GPUs".
- vonneumannstan 2y agoThe difference is the AI researchers have clear plots showing capabilities scaling with GPUs and there's not a sign that it is flattening so they actually have a case for saying that AGI is possible at N GPUs.
- segasaturn 2y agoSauce? How do you even measure "capabilities" in that regard, just writing answers to standard tests? Because being able to ace a test doesn't mean it's AGI, it means its good at taking standard tests.
- jms55 2y agoGreat article, thanks for writing it! Really great summary of the current state of the AI industry for someone like me who's outside of it (but tangential, given that I work with GPUs for graphics). The one thing from the article that sticks out to me is that the author/people are assuming that deepseek needing 1/45th the amount of hardware means that the other 44/45ths large tech companies have invested were wasteful. Does software not scale to meet hardware? I don't see this as 44/45ths wasted hardware, but as a free increase in the amount of hardware people have. Software needing less hardware means you can run even _more_ software without spending more money, not that you need less hardware, right? (for the top-end, non-embedded use cases). --- As an aside, the state of the "AI" industry really freaks me out sometimes. Ignoring any sort of short or long term effects on society, jobs, people, etc, just the sheer amount of money and time invested into this one thing is, insane? Tons of custom processing chips, interconnects, compilers, algorithms, _press releases!_, etc all for one specific field. It's like someone taking the last decade of advances in computers, software, etc, and shoving it in the space of a year. For comparison, Rust 1.0 is 10 years old - I vividly remember the release. And even then it took years to propagate out as a "thing" that people were interested in and invested significant time into. Meanwhile deepseek releases a new model (complete with a customer-facing product name and chat interface, instead of something boring and technical), and in 5 days it's being replicated (to at least some degree) and copied by competitors. Google, Apple, Microsoft, etc are all making custom chips and investing insane amounts of money into different compilers, programming languages, hardware, and research. It's just, kind of disquieting? Like everyone involved in AI lives in another world operating at breakneck speed, with billions of dollars involved, and the rest of us are just watching from the sidelines. Most of it (LLMs specifically) is no longer exciting to me. It's like, what's the point of spending time on a non-AI related project? We can spend some time writing a nice API and working on a cool feature or making a UI prettier and that's great, and maybe with a good amount of contributors and solid, sustained effort, we can make a cool project that's useful and people enjoy, and earns money to support people if it's commercial. But then for AI, github repos with shiny well-written readmes pop up overnight, tons of text is being written, thought, effort, and billions of dollars get burned or speculated on in an instant on new things, as soon as the next marketing release is posted. How can the next advancement in graphics, databases, cryptography, etc compete with the sheer amount of societal attention AI receives? Where does that leave writing software for the rest of us?
- mgraczyk 2y agoThe beginning of the article was good, but the analysis of DeepSeek and what it means for Nvidia is confused and clearly out of the loop. * People have been training models at <fp32 precision for many years, I did this in 2021 and it was already easy in all the major libraries. * GPU FLOPs are used for many things besides training the final released model. * Demand for AI is capacity limited, so it's possible and likely that increasing AI/FLOP would not substantially reduce the price of GPUs
- lysecret 2y agoWhere do you have this "capacity" limit from? I can get as many H100s from GCP or wherever as I wish, the only thing that is capacity limited are 100k clusters ala ELON+X, but what DeepSeek (and the recent evidence of a limit in pure base-model scaling) shows is that this might actually not be profitable, and we end up with much smaller base models scaled at inference time. The moat for Nvidia in this inference time scaling is much smaller, also you don't need the humongous clusters for that either you can just distribute the inference (and in the future run it locally too).
- aorloff 2y agoHis DeepSeek argument was essentially that experts who look at the economics of running these teams (eg. ha ha the engineers themselves might dabble) are looking over the hedge at DeepSeek's claims and they are really awestruck
- mkalygin 2y agoThis is such a comprehensive analysis, thank you. For someone just starting to learn about the field, it’s a great way to understand what’s going on in the industry.
- miraculixx 2y agoIf we are to get to AGI why do we need to train on all data? That's silly, and all we get is compression and probabliatic retrieval. Intelligence by definition is not compression, but ability to think and act according to new data, based on experience. Trully AGI models will work on the this principle, not on best compression of as much data as possible. We need a new approach.
- eigenvalue 2y agoActually, compression is an incredibly good way to think about intelligence. If you understand something really well then you can compress it a lot. If you can compress most of human knowledge effectively without much reconstruction error while shrinking it down by 99.5%, then you must have in the process arrived at a coherent and essentially correct world model, which is the basis of effective cognition.
- chpatrick 2y ago“If you can't explain it to a six year old, you don't understand it yourself.” -> "If you can compress knowledge, you understand it."
- AnotherGoodName 2y agoFwiw there's highly cited papers that literally map AGI to compression. As in they map to the same thing and people write papers on this fact that are widely respected. Basically a prediction engine can be used to make a compression tool and an AI equally. The tldr; if given inputs and a system that can accurately predict the next sequence you can either compress that data using that prediction (arithmetic coding) or you can take actions based on that prediction to achieve an end goal mapping predictions of new inputs to possible outcomes and then taking the path to a goal (AGI). They boil down to one and the same. So it's weird to have someone state they are not the same when it's widely accepted they absolutely are.
- jwan584 2y agoThe point about using FP32 for training is wrong. Mixed precision (FP16 multiplies, FP32 accumulates) has been use for years – the original paper came out in 2017.
- eigenvalue 2y agoFair enough, but that still uses a lot more memory during training than what DeepSeek is doing.
- suraci 2y agoDeepSeek is not the black swan NVDA was overpriced a lot already even without r1, the market is full of air GPUs hiding in the capex of tech giants like MSFT. If orders are canceled or delivery fails for any reason, NVDA’s EPS would be pulled back to its fundamentally justified level or if all those air GPUs are produced and delivered in recent years, and the demand keeps rising? well, that will be a crazy world then it's a finance game, not related with the real world
- naiv 2y agoI used to own several adult companies in the past. Incredible huge margins and then along came Pornhub and we could barely survive after it as we did not adapt. With Deepseek this is now the 'Pornhub of AI' moment. Adapt or die.
- logicchains 2y agoCurious what Pornhub did better, if you're able to say. Provide content at much lower cost, like DeepSeek?
- naiv 2y agoYes. close to free content They understood the Dmca brilliantly so they did bulk cheap content purchases and hid behind the Dmca for all non licensed content which was "uploaded by users". They did bulk purchases of cheap content from some studios but that was just a fraction Of course their risk of going advertise revenue only was high and in the beginning mostly only cam providers would advertise Our problem was that we had contracts and close relationships with all the big studios so going the Dmca route would have severed these ties for an unknown risk. In hindsight not creating a company which did abuse the Dmca was the right decision. I am very loyal and it would have felt like cheating Now it's a different story after the credit card shake down when they had to remove millions of videos and be able to provide 2257 documentation for each video
- nejsjsjsbsb 2y agoThat analogy would be if right if a startup could dredge beach sand and pump put trillions of AI chips. What actually happened was a better algorithm was created and people are betting against the main game in town for running said algorithm. If someone came up with a CPU-superior AI that'd be worrying for NVidia.
- naiv 2y agoGroq lpu interference chip?
- poopiokaka 2y ago[flagged]
- homarp 2y agosee also https://news.ycombinator.com/item?id=42839650 https://news.ycombinator.com/item?id=42839650
- chvid 2y agoFor sure NVIDIA is priced for perfection perhaps more than any of the other of similar market value. I think two threats are the biggest: First Apple. TSMC’s largest customer. They are already making their own GPUs for their data centers. If they were to sell these to others they would be a major competitor. You would have the same GPU stack on your on phone, laptop, pc, and data center. Already big developer mind share. Also useful in a world where LLMs run (in part) on the end user’s local machine (like Apple Intelligence). Second is China - Huawei, Deepseek etc. Yes - there will be no GPUs from Huawei in the US in this decade. And the Chinese won’t win in a big massive battle. Rather it is going to be death by a thousand cuts. Just as what happened with the Huawei Mate 60. It is only sold in China but today Apple is loosing business big time in China. In the same manner OpenAi and Microsoft will have their business hurt by Deepseek even if Deepseek was completely banned in the west. Likely we will see news on Chinese AI accelerators this year and I wouldn’t be surprised if we soon saw Chinese hyperscalars offering cheaper GPU cloud compute than the west due to a combination of cheaper energy, labor cost, and sheer scale. Lastly AMD is no threat to NVIDIA as they are far behind and follow the same path with little way of differentiating themselves.
- Giorgi 2y agoLooks like huge astroturfing effort from CCP. I am seeing these coordinated propaganda inside every AI related sub on reddit, on social media and now - here.
- dartos 2y agoThis just in. Competition lowers the value of monopolies.
- manojlds 2y ago>With the advent of the revolutionary Chain-of-Thought ("COT") models introduced in the past year, most noticeably in OpenAI's flagship O1 model (but very recently in DeepSeek's new R1 model, which we will talk about later in much more detail), all that changed. Instead of the amount of inference compute being directly proportional to the length of the output text generated by the model (scaling up for larger context windows, model size, etc.), these new COT models also generate intermediate "logic tokens"; think of this as a sort of scratchpad or "internal monologue" of the model while it's trying to solve your problem or complete its assigned task. Is this right? I thought CoT was a prompting method and are we calling the reasoning models as CoT models?
- veesahni 2y agoReasoning models are a result of the learnings from CoT prompting.
- s1mplicissimus 2y agoI'm curious what are the key differences between "a reasoning model" and good old CoT prompting. Is there any reason to believe that the fundamental limitations of prompting don't apply to "reasoning models"? (hallucinations, plainly wrong output, bias towards to training data mean etc.)
- itchyjunk 2y agoThe level of sophistication for CoT model varies. "good old CoT prompting" is you hoping the model generates some reasoning tokens prior to the final answer. When it did, the answers tended to be better for certain class of problems. But you had no control over what type of reasoning tokes it was generating. There were hypothesis that just having a <pause> tokens in between generated better answers as it allowed n+1 steps to generate an answer over n. I would consider Meta's "continuous chain of thought" to be on the other end of "good old CoT prompting" where they are passing back the next tokens from the latent space back to the model getting a "BHF" like effect. Who knows what's happening with O3 and Anthropics O3 like models.. The problems you mentioned is very broad and not limited to prompting. Reasoning models tend to outperform older models on math problems. So I'd assume it does reduce hallucination on certain class of problems.
- kimbler 2y agoNvidia seem to be one step ahead of this and you can see their platform efforts are pushing towards creating large volumes of compute that are easy to manage for whatever your compute requirements are, be that training, inference or whatever comes next and whatever form. People are maybe tackling some of these areas in isolation but you do not want to build datacenters where everything is ringfenced per task or usage.
- colinnordin 2y agoGreat article. >Now, you still want to train the best model you can by cleverly leveraging as much compute as you can and as many trillion tokens of high quality training data as possible, but that's just the beginning of the story in this new world; now, you could easily use incredibly huge amounts of compute just to do inference from these models at a very high level of confidence or when trying to solve extremely tough problems that require "genius level" reasoning to avoid all the potential pitfalls that would lead a regular LLM astray. I think this is the most interesting part. We always knew a huge fraction of the compute would be on inference rather than training, but it feels like the newest developments is pushing this even further towards inference. Combine that with the fact that you can run the full R1 (680B) distributed on 3 consumer computers [1]. If most of NVIDIAs moat is in being able to efficiently interconnect thousands of GPUs, what happens when that is only important to a small fraction of the overall AI compute? [1]: https://x.com/awnihannun/status/1883276535643455790 https://x.com/awnihannun/status/1883276535643455790
- tomrod 2y agoConversely, how much larger can you scale if frontier models only currently need 3 consumer computers? Imagine having 300. Could you build even better models? Is DeepSeek the right team to deliver that, or can OpenAI, Meta, HF, etc. adapt? Going to be an interesting few months on the market. I think OpenAI lost a LOT in the board fiasco. I am bullish on HF. I anticipate Meta will lose folks to brain drain in response to management equivocation around company values. I don't put much stock into Google or Microsoft's AI capabilities, they are the new IBMs and are no longer innovating except at obvious margins.
- danaris 2y agoThis assumes no (or very small) diminishing returns effect. I don't pretend to know much about the minutiae of LLM training, but it wouldn't surprise me at all if throwing massively more GPUs at this particular training paradigm only produces marginal increases in output quality.
- 2y ago
- brandonpelfrey 2y agoGreat article. I still feel like very few people are viewing the Deepseek effects in the right light. If we are 10x more efficient it's not that we use 1/10th the resources we did before, we expand to have 10x the usage we did before. All technology products have moved this direction. Where there is capacity, we will use it. This argument would not work if we were close to AGI or something and didn't need more, but I don't think we're actually close to that at all.
- VHRanger 2y agoCorrect. This effect is known in economics since forever - new technology has - An "income effect". You use the thing more because it's cheaper - new usecases come up - A "substitution effect." You use other things more because of the savings. I got into this on labor economics here [1] - you have counterintuitive examples with ATMs actually increasing the number of bank branches for several decades. [1]: https://singlelunch.com/2019/10/21/the-economic-effects-of-automation-arent-what-you-think-they-are https://singlelunch.com/2019/10/21/the-economic-effects-of-a...
- neuronic 2y agoWould this not mean we need much much more training data to fully utilize the now "free" capacities?
- vonneumannstan 2y agoIt's pretty clear that the reasoning models are using mass amounts of synthetic data so it's not a bottleneck.
- jnwatson 2y agoThis is called Jevons Paradox. https://en.wikipedia.org/wiki/Jevons_paradox https://en.wikipedia.org/wiki/Jevons_paradox.
- BlackSwanMan 2y ago[dead]
- 2y ago
- p0w3n3d 2y ago> which require low-latency responses, such as content moderation, fraud detection, dynamic pricing, etc. Is it even legal to give different prices to different customers?
- jnwatson 2y agoOf course it is. That how the airlines stay in business.
- p0w3n3d 2y agoHowever imagine entering a store where the camera looks up your face in shared database and profiles you as a person who will pay higher prices - and the prices are displayed near you according to your profile...
- esafak 2y agoIt depends on what basis. You can't discriminate based on protected classes.
- typeofhuman 2y agoI'm rooting for DeepSeek (or any competitor) against OpenAI because I don't like Sam Altman. I'm confident in admitting it.
- 1970-01-01 2y agoThe enemy of your enemy is only temporarily your friend.
- typeofhuman 2y agoWise words from the epoch of time.
- brianbest101 2y ago[dead]
- TypingOutBugs 2y agoAs a European I really don’t see the difference between US and Chinese tech right now - the last week from Trump has made me feel more threatened from the US than I ever have been by China (Greenland, living in a Nordic country with treaties to defend it). I appreciate China has censorship, but the US is going that way too (recent “issues” for search terms). Might be different scales now, but I think it’ll happen. I don’t care as much if a Chinese company wins the LLM space than I did last year.
- rwoerz 2y agoIndeed! Just ask DeepSeek something about Tiananmen or Taiwan. Answering seems to be an absolute "no-brainer" for it.
- eigenvalue 2y agoI really don't think he's a bad guy. He helped accelerate timelines and backed this tech when it was still a dream. Maybe he's not the brains behind it but he's been the brawn, and I think people should try to be more charitable and gracious about him rather than constantly vilify him.
- liuliu 2y agoThis is a humble and informed acrticle (comparing to others written by financial analysts the past a few days). But still have the flaw of over-estimating efficiency of deploying a 687B MoE model on commodity hardware (to use locally, cloud providers will do efficient batching and it is different): you cannot do that on any single Apple hardware (need to at least hook up 2 M2 Ultra). You can barely deploy that on desktop computers just because non-register DDR5 can have 64GiB per stick (so you are safe with 512 RAM). Now coming to PCIe bandwidth: 37B per token activation means exactly that, each activation requires new set of 37B weights, so you need to transfer 18GiB per token into VRAM (assuming 4-bit quant). PCIe 5 (5090) have 64GB/s transfer speed so your upper bound is limited to 4 tok/s with a well balanced propose built PC (and custom software). For programming tasks that usually requires ~3000 tokens for thinking, we are looking at 12 mins per interaction.
- lvass 2y agoIs it really 37B different parameters for each token? Even with the "multi-token prediction system" that the article mentions?
- liuliu 2y agoI don't think anyone uses MTP for inference right now. Even if you use MTP for drafting, you need to batching in the next round to "verify" it is the right token, if that happens you need to activate more experts. DELETED: If you don't use MTP for drafting, and use MTP to skip generations, sure. But you also need to evaluate your use case to make sure you don't get penalized for doing that. Their evaluation in the paper don't use MTP for generation. EDIT: Actually, you cannot use MTP other than drafting because you need to fill in these KV caches. So, during generation, you cannot save your compute with MTP (you save memory bandwidth, but this is more complicated for MoE model due to more activated experts).
- pjdesno 2y agoThe description of DeepSeek reminds me of my experience in networking in the late 80s - early 90s. Back then a really big motivator for Asynchronous Transfer Mode (ATM) and fiber-to-the-home was the promise of video on demand, which was a huge market in comparison to the Internet of the day. Just about all the work in this area ignored the potential of advanced video coding algorithms, and assumed that broadcast TV-quality video would require about 50x more bandwidth than today's SD Netflix videos, and 6x more than 4K. What made video on the Internet possible wasn't a faster Internet, although the 10-20x increase every decade certainly helped - it was smarter algorithms that used orders of magnitude less bandwidth. In the case of AI, GPUs keep getting faster, but it's going to take a hell of a long time to achieve a 10x improvement in performance per cm^2 of silicon. Vastly improved training/inference algorithms may or may not be possible (DeepSeek seems to indicate the answer is "may") but there's no physical limit preventing them from being discovered, and the disruption when someone invents a new algorithm can be nearly immediate.
- TMWNN 2y ago>but there's no physical limit preventing them from being discovered, and the disruption when someone invents a new algorithm can be nearly immediate. The rise of the net is Jevons paradox fulfilled. The orders of magnitude less bandwidth needed per cat video drove much more than that in overall growth in demand for said videos. During the dotcom bubble's collapse, bandwidth use kept going up. Even if there is a near-term bear case for NVDA (dotcom bubble/bust), history indicates a bull case for the sector overall and related investments such as utilities (the entire history of the tech sector from 1995 to today).
- accra4rx 2y agoLove those analogies . This is one of main reason I love hacker news / reddit . Honest golden experiences
- AlanYx 2y agoAnother aspect that reinforces your point is that the ATM push (and subsequent downfall) was not just bandwidth-motivated but also motivated by a belief that ATM's QoS guarantees were necessary. But it turned out that software improvements, notably MPLS to handle QoS, were all that was needed.
- brianbest101 2y ago[dead]
- aurareturn 2y agoPerhaps most devastating is DeepSeek's recent efficiency breakthrough, achieving comparable model performance at approximately 1/45th the compute cost. This suggests the entire industry has been massively over-provisioning compute resources. I wrote in another thread why DeepSeek should increase demand for chips, not lower. 1. More efficient LLMs should lead to more usage, which means more AI chip demand. Jevon's Paradox. 2. Even if DeepSeek is 45x more efficient (it is not), models will just become 45x+ bigger. It won’t stay small. 3. To build a moat, OpenAI and American AI companies need to up their datacenter spending even more. 4. DeepSeek's breakthrough is in distilling models. You still need a ton of compute to train the foundational model to distill. 5. DeepSeek's conclusion in their paper says more compute is needed for next break through. 6. DeepSeek's model is trained on GPT4o/Sonnet outputs. Again, this reaffirms the fact that in order to take the next step, you need to continue to train better models. Better models will generate better data for next-gen models. I think DeepSeek hurts OpenAI/Anthropic/Google/Microsoft. I think DeepSeek helps TSMC/Nvidia. Combined with the emergence of more efficient inference architectures through chain-of-thought models, the aggregate demand for compute could be significantly lower than current projections assume. This is misguided. Let's think logically about this. More thinking = smarter models Faster hardware = more thinking More/newer Nvidia GPUs, better TSMC nodes = faster hardware Therefore, you can conclude that Nvidia and TSMC demand should go up because of CoT models. In 2025, CoT models are clearly bottlenecked by not having enough compute. The economics here are compelling: when DeepSeek can match GPT-4 level performance while charging 95% less for API calls, it suggests either NVIDIA's customers are burning cash unnecessarily or margins must come down dramatically. Or that in order to build a moat, OpenAI/Anthropic/Google and other laps need to double down on even more compute.
- outside1234 2y agoBut Microsoft hosts 3rd party models too, and cheaper models means more usage, which means more $$$ to scaled cloud providers right?
- clvx 2y agoit means they can serve more with what they have if they implement models with deepseek's optimizations. More usage doesn't mean Nvidia will get the same margins when cloud providers scale out with this innovation.
- blurbleblurble 2y agoThis is exactly where project digits comes in. Nvidia needs to pivot toward being a local inference platform if they want to survive the next shift.
- skizm 2y agoI'm wondering if there's a (probably illegal) strategy in the making here: - Wait till NVDA rebounds in price. - Create an OpenAI "competitor" that is powered by Llama or a similar open weights model. - Obscure the fact that the company runs on this open tech and make it seem like you've developed your own models, but don't outright lie. - Release an app and whitepaper (whitepaper looks and sounds technical, but is incredibly light on details, you only need to fool some new-grad stock analysts). - Pay some shady click farms to get your app to the top of Apples charts (you only need it to be there for like 24 hours tops). - Collect profits from your NVDA short positions.
- tw1984 2y agothis is exactly what DeepSeek is doing, the only difference is they built the real model, not a fake one.
- startupsfail 2y ago- Fail at the above. I don’t think this is what happened with DeepSeek. It seems that they’ve genuinely optimized their model for efficiency and used GPUs properly (tiled FP8 trick and FP8 training). And came out on top. The impact on the NVIDIA stock is ridiculous. DeepSeek took the advantage of flexible GPU architecture (unlike inflexible hardware acceleration).
- deleted 2y ago[deleted]
- mmiliauskas 2y agoThis is what I still don't understand, how much of what they claim has been actually replicated? From what I understand the "50x cheaper" inference is coming from their pricing page, but is it actually 50x cheaper than the best open source models?
- zamadatix 2y ago50x cheaper than OpenAI's pricing on an open source model which doesn't require giving that quality level up. The best open source models were much closer in pricing but V3/R1 are that way while being a results topper.
- fairity 2y agoDeepSeek just further reinforces the idea that there is a first-move disadvantage in developing AI models. When someone can replicate your model for 5% of the cost in 2 years, I can only see 2 rational decisions: 1) Start focusing on cost efficiency today to reduce the advantage of the second mover (i.e. trade growth for profitability) 2) Figure out how to build a real competitive moat through one or more of the following: economies of scale, network effects, regulatory capture On the second point, it seems to me like the only realistic strategy for companies like OpenAI is to turn themselves into a platform that benefits from direct network effects. Whether that's actually feasible is another question.
- Mistletoe 2y agoI feel like AI tech just reverse scales and reverse flywheels, unlike the tech giant walls and moats now, and I think that is wonderful. OpenAI has really never made sense from a financial standpoint and that is healthier for humans. There’s no network effect because there’s no social aspect to AI chatbots. I can hop on DeepSeek from Google Gemini or OpenAI at ease because I don’t have to have friends there and/or convince them to move. AI is going to be a race to the bottom that keeps prices low to zero. In fact I don’t know how they are going to monetize it at all.
- tw1984 2y ago> DeepSeek just further reinforces the idea that there is a first-move disadvantage in developing AI models. you are assuming that what DeepSeek achieved can be reasonably easily replicated by other companies. then the question is when all big techs and tons of startups in China and the US are involved, how come none of those companies succeeded? deepseek is unique.
- 11101010001100 2y agoDeepseek is unique, but the US has consistently underestimated Chinese R&D, which is not a winning strategy in iterated games.
- 2y ago
- 11101010001100 2y agoI think this is just a(nother) canary for many other markets in the US v China game of monopoly. One weird effect in all this is that US Tech may go on to be over valued (i.e., disconnect from fundamentals) for quite some time.
- btbuildem 2y agoI always appreciate reading a take from someone who's well versed in the domains they have opinions about. I think longer-term we'll eat up any slack in efficiency by throwing more inference demands at it -- but the shift is tectonic. It's a cultural thing. People got acclimated to shlepping around morbidly obese node packages and stringing together enormous python libraries - meanwhile the deepseek guys out here carving bits and bytes into bare metal. Back to FP!
- vonneumannstan 2y agoThis is a bizarre take. First Deepseek no doubt is still using the same bloated Python ML packages as everyone else. Second since this is "open source" it's pretty clear that the big labs are just going to replicate this basically immediately and with their already massive compute advantages put models out that are extra OOM larger/better/etc/ than what Deepseek can possibly put out. Theres just no reason to think that e.g. a 10x increase in training efficiency does anything but increase the size of the next model generation by 10x.
- christkv 2y agoAll this is good news for all of us. Bad news probably for Nvidia's margins long term but who cares. If we can train and inference in less cycles and watts that is awesome.
- qwertox 2y agoConsidering the fact that current models were trained on top-notch books, those read and studied by the most brilliant engineers, the models are pretty dumb. They are more like the thing which enabled computers to work with and digest text instead of just code. The fact that they can parrot pretty interesting relationships from the texts they've consumed kind of proofs that they are capable of statistically "understanding" what we're trying to talk with them about, so it's a pretty good interface. But going back to the really valuable content of the books they've been trained on, they just don't understand it. There's other AI which needs to get created which can really learn the concepts taught in those books instead of just the words and the value of the proximities between them. To learn that other missing part will require hardware just as uniquely powerful and flexible as what Nvidia has to offer. Those companies now optimizing for inference and LLM training will be good at it and have their market share, but they need to ensure that their entire stack is as capable of Nvidia's stack, if they also want to be part of future developments. I don't know if Tenstorrent or Groq are capable of doing this, but I doubt it.
- lenerdenator 2y agoI think it's more than just the market effect on "established" AI players like Nvidia. I don't think it's necessarily a coincidence that DeepSeek dropped within a short time frame of the announcement of the AI investment initiative by the Trump administration. The idea is to get the money from investors who want to earn a return. Lower capex is attractive to investors, and DS drops capex dramatically. It makes Chinese AI talent look like the smart, safe bet. Nothing like DS could happen in China unless the powers-that-be knew about it and got some level of control. I'm also willing to bet that this isn't the best they've got. They're saying "we can deliver the same capabilities for far less, and we're not going to threaten you with a tariff for not complying".
- robomartin 2y agoDespite the fact that this article is very well written and certainly contains high quality information, I choose to remain skeptical as it pertains to Nvidia's position in the market. I'll come right out and say that my experience likely makes me see this from a biased position. The premise is simple: Business is warfare. Anything you can do to damage or slow down the market leader gives you more time to get caught up. FUD is a powerful force. My bias comes from having been the subject of such attacks in my prior tech startup. Our technology was destroying the offerings of the market leading multi-billion-dollar global company that pretty much owned the sector. The natural processes of such a beast caused them not to be able to design their way out of a paper bag. We clearly had an advantage. The problem was that we did not have the deep pockets necessary to flood the market with it and take them out. What did they do? The started a FUD campaign. They went to every single large customer and our resellers (this was a hardware/software product) a month or two before the two main industry tradeshows, and lied to them. They promised that they would show market-leading technology "in just a couple of months" and would add comments like "you might want to put your orders on hold until you see this". We had multi-million dollar orders held for months in anticipation of these product unveilings. And, sure enough, they would announce the new products with a great marketing push at the next tradeshow. All demos were engineered and manipulated to deceive, all of them. Yet, the incredible power of throwing millions of dollars at this effort delivered what they needed, FUD. The problem with new products is that it takes months for them to be properly validated. So, if the company that had frozen a $5MM order for our products decides to verify the claims of our competitor, it typically took around four months. In four months, they would discover that the new shiny object was shit and less stellar than what they were told. I other words, we won. Right? No! The mega-corp would then reassure them that they iterated vast improvements into the design and those would be presented --I kid you not-- at the next tradeshow. Spending millions of dollars they, at this point, denied us of millions of dollars of revenue for approximately one year. FUD, again. The next tradeshow came and went and the same cycle repeats...it would take months for customers to realize the emperor had no clothes. It was brutal to be on the receiving end of this without the financial horsepower to be able to break through the FUD. It was a marketing arms race and we were unprepared to win it. In this context, the idea that a better mouse trap always wins is just laughable. This did not end well. They were not going to survive another FUD cycle. Reality eventually comes into play. Except that, in this case, 2008 happened. The economic implosion caught us in serious financial peril due to the damage done by the FUD campaign. Ultimately, it was not survivable and I had to shut down the company. It took this mega-corp another five years to finally deliver a product that approximated what we had and another five years after that to match and exceed it. I don't even want to imagine how many hundreds of millions they spent on this. So, long way of saying: China wants to win. No company in China is independent from government forces. This is, without a doubt, a war for supremacy in the AI world. It is my opinion that, while the technology, as described, seems to make sense, it is highly likely that this is yet another form of a FUD campaign to gain time. If they can deny Nvidia (and others) the orders needed to maintain the current pace, they gain time to execute on a strategy that could give them the advantage. Time will tell.
- mars009 2y ago[dead]
- samiv 2y agoI think the biggest threat for future NVIDIa right now is their own current success. Their software platforms and CUDA are a very strong moat against everyone else. I don't see any beating them on that front right now. The problem is that I'm afraid that all that money sloshing inside the company is rotting the culture and that will compromise future development. - Grifters are filling out positions in many orgs only trying to milk it as much as possible. - Old employees become complacent with their nice RSU packages Rest & Vest. NVIDIA used to be extremely nimble and was way fighting way above it's weight class. Prior to Mellanox acquisition only around 10k employees and after another 10k more. If there's a real threat to their position at the top of the AI offerings will they be able to roll up the sleeves and get back to work or will the organizations be unable to move ahead. Long term I think it's inevitable that China will take over the technology leadership. They have the population and they have the education programs and the skill to do this. At the same time in the old western democracies things are becoming stagnant and I even dare to say that the younger generations are declining. In my native country the educational system has collapsed, over 20% kids that finish elementary school cannot read or write. They can mouth-breath and scroll TikTok though but just barely since their attention span is about the same as gold fish.
- _DeadFred_ 2y agoLOL. This isn't rot, it is reaching the end goal, the people doing the work reach the rewards they were working towards. Rot would imply somehow management should prevent rest and vest but that is the exact model that they acquired their talent on. You would have to remove capitalism from companies when companies win at capitalism making it all just a giant rug pull for employees.
- scudsworth 2y agowhat a compelling domain name. it compels me not to click on it
- indymike 2y agoThis story could be applied to every tech breakthrough. We start where the breakthrough is moated by hardware, access to knowledge, and IP. Over time: - Competition gets crucial features into cheaper hardware - Work-arounds for most IP are discovered - Knowledge finds a way out of the castle This leads to a "Cambrian explosion" of new devices and software that usually gives rise to some game-changing new ways to use the new technology. I'm not sure where we all thought this somehow wouldn't apply to AI. We've seen the pattern with almost every new technology you can think of. It's just how it works. Only the time it takes for patents to expire changes this... so long as everyone respects the patent.
- eigenvalue 2y agoYes this is exactly right. All you need is the right incentives and enough capital and markets will find away to breech any moat that’s not enforced via regulations.
- indymike 2y agoThey'll even solve the regulations part most of the time as well. See Uber.
- _DeadFred_ 2y agoIt's still wild to me that toasters have always been $20 but extremely expensive lasers, digital chips, amps, motors, LCD screens worked their way down to $20 CD players.
- indymike 2y agoSo... Electric toasters came to market in the 1920s, priced from $15, eventually getting as low as $5. Adjusting for inflation, that $15 toaster cost $236.70 in 2025 USD. Today's $15 toaster would be about 90¢ in 1920s dollars... so it follows the story.
- _DeadFred_ 2y ago
- lxgr 2y agoThe most important part for me is: > DeepSeek is a tiny Chinese company that reportedly has under 200 employees. The story goes that they started out as a quant trading hedge fund similar to TwoSigma or RenTec, but after Xi Jinping cracked down on that space, they used their math and engineering chops to pivot into AI research. I guess now we have the answer to the question that countless people have already asked: Where could we be if we figured out how to get most math and physics PhDs to work on things other than picking up pennies in front of steamrollers (a.k.a. HFT) again?
- rfoo 2y agoThis is completely fake though. It was more like their founder decided to start a branch to do AI research. It was well planned, they bought significantly more GPUs than they can use for quant research even before they start to do anything AI. There was a crack down on algorithmic trading, but it didn't had much impact and IMO someone higher up definitely does not want to kill these trading firms.
- lxgr 2y agoThe optimal amount of algorithmic trading is definitely more than none (I appreciate liquidity and price quality as much as the next guy), but arguably there's a case here that we've overshot a bit.
- rightbyte 2y agoThe price data I (we?) get is 15 minute delayed. I would guess most of the profiteering is from consumers not knowing the last transaction prices? I.e. an artificially created edge by the broker who then sells the API to clean their hands of the scam.
- lxgr 2y agoReal-time price data is indeed not free, but widely available even in retail brokerages. I've never seen a 15 minute delay in any US based trade, and I think I can even access level 2 data a limited number of times on most exchanges (not that it does me much good as a retail investor). > I would guess most of the profiteering is from consumers not knowing the last transaction prices? No, not at all. And I wouldn't even necessarily call it profiteering. Ironically, as a retail investor you even benefit from hedge funds and HFTs being a counterpart to your trades: You get on average better (and worst case as good) execution from PFOF. Institutional investors (which include pension funds, insurances etc.) are a different story.
- hn_throwaway_99 2y agoI'm curious if someone more informed than me can comment on this part: > Besides things like the rise of humanoid robots, which I suspect is going to take most people by surprise when they are rapidly able to perform a huge number of tasks that currently require an unskilled (or even skilled) human worker (e.g., doing laundry ... I've always said that the real test for humanoid AI is folding laundry, because it's an incredibly difficult problem. And I'm not talking about giving a machine clothing piece-by-piece flattened so it just has to fold, I'm talking about saying to a robot "There's a dryer full of clothes. Go fold it into separate piles (e.g. underwear, tops, bottoms) and don't mix the husband's clothes with the wife's". That is, something most humans in the developed world have to do a couple times a week. I've been following some of the big advances in humanoid robot AI, but the above task still seems miles away given current tech. So is the author's quote just more unsubstantiated hype that I'm constantly bombarded with in the AI space, or have there been advancements recently in robot AI that I'm unaware of?
- ieee2 2y agoI saw such robot's demos doing exactly that on youtube/x - not very precisely yet, but almost sufficiently enough. And it is just a beginning. Considering that majority of the laundry is very similar (shirts, t-shirts, trousers, etc..) I think this will be solved soon with enough training.
- hn_throwaway_99 2y agoCan you share what you've seen? Because from what I've seen, I'm far from convinced. E.g. there is this, https://youtube.com/shorts/CICq5klTomY https://youtube.com/shorts/CICq5klTomY , which nominally does what I've described. Still, as impressive as that is, I think the distance from what that robot does to what a human can do is a lot farther than it seems. Besides noticing that the folded clothes are more like a neatly arranged pile, what about all the edge cases? What about static cling? Can it match socks? What if something gets stuck in the dryer? I'm just very wary of looking at that video and saying "Look! It's 90% of the way there! And think how fast AI advances!", because that critical last 10% can often be harder than the first 90% and then some.
- rashidae 2y agoWhile Nvidia’s valuation may feel bloated due to AI hype, AMD might be the smarter play.
- UncleOxidant 2y agoEven if DeepSeek has figured out how to do more (or at least as much) with less, doesn't the Jevons Paradox come into play? GPU sales would actually increase because even smaller companies would get the idea that they can compete in a space that only 6 months ago we assumed would be the realm of the large mega tech companies (the Metas, Googles, OpenAIs) since the small players couldn't afford to compete. Now that story is in question since DeepSeek only has ~200 employees and claims to be able to train a competitive model for about 20X less than the big boys spend.
- samvher 2y agoMy interpretation is that yes in the long haul, lower energy/hardware requirements might increase demand rather than decrease it. But right now, DeepSeek has demonstrated that the current bottleneck to progress is _not_ compute, which decreases the near term pressure on buying GPUs at any cost, which decreases NVIDIA's stock price.
- kemiller 2y agoShort term, I 100% agree, but remains to be seen what "short" means. According to at least some benchmarks, Deepseek is two full orders of magnitude cheaper for comparable performance. Massive. But that opens the door for much more elaborate "architectures" (chain of thought, architect/editor, multiple choice) etc, since it's possible to run it over and over to get better results, so raw speed & latency will still matter.
- groby_b 2y agoI think it's worth carefully pulling apart _what_ DeepSeek is cheaper at. It's somewhat cheaper at inference (0.3 OOM), and about 1-1.5 OOM cheaper for training (Inference costs: https://www.latent.space/p/reasoning-price-war https://www.latent.space/p/reasoning-price-war) It's also worth keeping in mind that depending on benchmark, these values change (and can shrink quite a bit) And it's also worth keeping in mind that the drastic drop in training cost(if reproducible) will mean that training is suddenly affordable for a much larger number of organizations. I'm not sure the impact on GPU demand will be as big as people assume.
- mackid 2y agoMicrosoft did a bunch of research into low-bit weights for models. I guess OAI didn’t look at this work. https://proceedings.neurips.cc/paper/2020/file/747e32ab0fea7fbd2ad9ec03daa3f840-Paper.pdf https://proceedings.neurips.cc/paper/2020/file/747e32ab0fea7...
- highfrequency 2y agoThe R1 paper (https://arxiv.org/pdf/2501.12948 https://arxiv.org/pdf/2501.12948) emphasizes their success with reinforcement learning without requiring any supervised data (unlike RLHF for example). They note that this works well for math and programming questions with verifiable answers. What's totally unclear is what data they used for this reinforcement learning step. How many math problems of the right difficulty with well-defined labeled answers are available on the internet? (I see about 1,000 historical AIME questions, maybe another factor of 10 from other similar contests). Similarly, they mention LeetCode - it looks like there are around 3000 LeetCode questions online. Curious what others think - maybe the reinforcement learning step requires far less data than I would guess?
- mrinterweb 2y agoThe vast majority of Nvidia's current value is tied to their dominance in AI hardware. That value could be threatened if LLMs could be trained and or ran efficiently using a CPU or a quantum chip. I don't understand enough about the capabilities of quantum computing to know if running or training a LLM would be possible using a quantum chip, but if it becomes possible, NVDA stock is unlikely to fair well (unless they are making the new chip).
- tempeler 2y agoFirst of all, I don't invest in Nvidia and like Olygopols. But it is too early to talk about Nvidia's future. People are just betting and wishing about Nvidia's future. No one knows people's what people will do in the future. what they will think? It's just guessing and betting. Their real competitor is not Deepseek. Did AMD or others release something new and compete with Nvidia's products? If NVDIA will be the market leader, this means they will lead the price. Being Olygopol is something like that. They don't need to compete for the price of competitors.
- wtcactus 2y agoTo me, this seems like we are back again in 1953 and a company just announced they are now capable of building one of IBM's 5 computers for 10% of the price. I really don't understand the rationale of "We can now train GPT 4o for 10% the price, so that will bring demand for GPUs down.". If I can train GPT 4o for 10% the price, and I have a budget of 1B USD, that means I'm now going to use the same budget and train my model for 10x as long (or 10x bigger). At the same time, a lot of small players that couldn't properly train a model before, because the starting point was simply out of their reach, will now be able to purchase equipment that's capable of something of note, and they will buy even more GPUs. P.S. Yes, I know that the original quote "I think there is a world market for maybe five computers", was taken out of context. P.S.S. In this rationale, I'm also operating under the assumption that Deepseek numbers are real. Which, given the track record of Chinese companies, is probably not true.
- ozten 2y agoNVIDIA sells shovels to the gold rush. One miner (Liang Wenfeng), who has previously purchased at least 10,000 A100 shovels... has a "side project" where they figured out how to dig really well with a shovel and shared their secrets. The gold rush, wether real or a bubble is still there! NVIDA will still sell every shovel they can manufacture, as soon as it is available in inventory. Fortune 100 companies will still want the biggest toolshed to invent the next paradigm or to be the first to get to AGI.
- 0n0n0m0uz 2y agoPlease tell me if I am wrong. I know very little details and heard a few headlines and my hasty conclusion is that this development clearly shows the exponential nature of AI development in terms of how people are able to piggyback from the resources, time and money of the previous iteration. They used the output from chatgpt as the input to their model. Is this true, more or less accurate or off base?
- coolThingsFirst 2y agoAs a bystander it's so refreshing to see this, global tech competition is great for the market and it gives hope that LLMs aren't locked behind Bs of investments and smaller players can compete well as well. Exciting times to be living in .
- plaidfuji 2y agoThis is such a great read. The only missing facet of discussion here is that there is a valuation level of NVDA such that it would tip the balance of military action by China against Taiwan. TSMC can only drive so much global value before the incentive to invade becomes irresistible. Unclear where that threshold is; if we’re being honest, could be any day.
- lauriewired 2y agoDoes no one realize this is a thinly-veiled ad? The URL is bizarre
- eigenvalue 2y agoA thinly veiled ad? You must be joking.
- dang 2y agoRelated ongoing thread: Nvidia’s $589B DeepSeek rout - https://news.ycombinator.com/item?id=42839650 https://news.ycombinator.com/item?id=42839650 - Jan 2025 (574 comments)
- nokun7 2y agoVery interesting and it seems like there is more room for optimizations for WASM using SIMD, boosting performance by a lot! It's cool to see how AI can now run even faster on web browsers.
- greenie_beans 2y agoreading this gave me a great idea for https://bookhead.net https://bookhead.net. thanks!! also thank you for the incredibly informative article.
- greenie_beans 2y agoactually idk if the idea is "great" but it felt like it at the time