7 ms·
Great article. >Now, you still want to train the best model you can by cleverly leveraging as much compute as you can and as many trillion tokens of high quali
by colinnordin 2y ago
Great article.
>Now, you still want to train the best model you can by cleverly leveraging as much compute as you can and as many trillion tokens of high quality training data as possible, but that's just the beginning of the story in this new world; now, you could easily use incredibly huge amounts of compute just to do inference from these models at a very high level of confidence or when trying to solve extremely tough problems that require "genius level" reasoning to avoid all the potential pitfalls that would lead a regular LLM astray.
I think this is the most interesting part. We always knew a huge fraction of the compute would be on inference rather than training, but it feels like the newest developments is pushing this even further towards inference.
Combine that with the fact that you can run the full R1 (680B) distributed on 3 consumer computers [1].
If most of NVIDIAs moat is in being able to efficiently interconnect thousands of GPUs, what happens when that is only important to a small fraction of the overall AI compute?
[1]: https://x.com/awnihannun/status/1883276535643455790 https://x.com/awnihannun/status/1883276535643455790
- tomrod 2y agoConversely, how much larger can you scale if frontier models only currently need 3 consumer computers? Imagine having 300. Could you build even better models? Is DeepSeek the right team to deliver that, or can OpenAI, Meta, HF, etc. adapt? Going to be an interesting few months on the market. I think OpenAI lost a LOT in the board fiasco. I am bullish on HF. I anticipate Meta will lose folks to brain drain in response to management equivocation around company values. I don't put much stock into Google or Microsoft's AI capabilities, they are the new IBMs and are no longer innovating except at obvious margins.
- danaris 2y agoThis assumes no (or very small) diminishing returns effect. I don't pretend to know much about the minutiae of LLM training, but it wouldn't surprise me at all if throwing massively more GPUs at this particular training paradigm only produces marginal increases in output quality.
- tomrod 2y agoI believe the margin to expand is on CoT, where tokens can grow dramatically. If there is value in putting more compute towards it, there may still be returns to be captured on that margin.
- stormfather 2y agoGoogle is silently catching up fast with Gemini. They're also pursuing next gen architectures like Titan. But most importantly, the frontier of AI capabilities is shifting towards using RL at inference (thinking) time to perform tasks. Who has more data than Google there? They have a gargantuan database of queries paired with subsequent web nav, actions, follow up queries etc. Nobody can recreate this, Bing failed to get enough marketshare. Also, when you think of RL talent, which company comes to mind? I think Google has everyone checkmated already.
- moffkalast 2y agoNever underestimate Google's ability to fall flat on their face when it comes to shipping products.
- _DeadFred_ 2y agoHow quickly the narrative went from 'Google silently has the most advanced AI but they are afraid to release it' to 'Google is silently catching up' all using the same 'core Google competencies' to infer Google's position of strength. Wonder what the next lower level of Google silently leveraging their strength will be?
- stormfather 2y agoGoogle is clearly catching up. Have you tried the recent Gemini models? Have you tried deep research? Google is like a ship that is hard to turn around but also hard to stop once in motion.
- shwaj 2y agoCan you say more about using RL at inference time, ideally with a pointer to read more about it? This doesn’t fit into my mental model, in a couple of ways. The main way is right in the name: “learning” isn’t something that happens at inference time; inference is generating results from already-trained models. Perhaps you’re conflating RL with multistage (e.g. “chain of thought”) inference? Or maybe you’re talking about feeding the result of inference-time interactions with the user back into subsequent rounds of training? I’m curious to hear more.
- onlyrealcuzzo 2y agoIf you watch this video, it explains well what the major difference is between DeepSeek and existing LLMs: https://www.youtube.com/watch?v=DCqqCLlsIBU https://www.youtube.com/watch?v=DCqqCLlsIBU It seems like there is MUCH to gain by migrating to this approach - and it theoretically should not cost more to switch to that approach than vs the rewards to reap. I expect all the major players are already working full-steam to incorporate this into their stacks as quickly as possible. IMO, this seems incredibly bad to Nvidia, and incredibly good to everyone else. I don't think this seems particularly bad for ChatGPT. They've built a strong brand. This should just help them reduce - by far - one of their largest expenses. They'll have a slight disadvantage to say Google - who can much more easily switch from GPU to CPU. ChatGPT could have some growing pains there. Google would not.
- wolfhumble 2y ago> I don't think this seems particularly bad for ChatGPT. They've built a strong brand. This should just help them reduce - by far - one of their largest expenses. Often expenses like that are keeping your competitors away.
- onlyrealcuzzo 2y agoYes, but it typically doesn't matter if someone can reach parity or even surpass you - they have to surpass you by a step function to take a significant number of your users. This is a step function in terms of efficiency (which presumably will be incorporated into ChatGPT within months), but not in terms of end user experience. It's only slightly better there.
- ReptileMan 2y agoOne data point but my subscription for chatgpt is cancelled every time. So I made every month decision to resub. And because the cost of switching is essentially zero - the moment a better service is up there I will switch in an instant.
- 2y ago
- simpaticoder 2y ago>Imagine having 300. Would it not be useful to have multiple independent AIs observing and interacting to build a model of the world? I'm thinking something roughly like the "councelors" in the Civilization games, giving defense/economic/cultural advice, but generalized over any goal-oriented scenario (and including one to take the "user" role). A group of AIs with specific roles interacting with each other seems like a good area to explore, especially now given the downward scalability of LLMs.
- tomrod 2y agoYes; to my understanding that is MoE.
- JoshTko 2y agoThis is exactly where Deepseeks enhancements come into play. Essentially deepseek lets the model think out loud via chain of thought (o1 and Claude also do this) but DS also does not supervise the chain of thought, and simply rewards CoT that get the answer correctly. This is just one of the half dozen training optimization that Deepseek has come up with.
- neuronic 2y ago> NVIDIAs moat Offtopic, but your comment finally pushed me over the edge to semantic satiation [1] regarding the word "moat". It is incredible how this word turned up a short while ago and now it seems to be a key ingredient of every second comment. [1] https://en.wikipedia.org/wiki/Semantic_satiation https://en.wikipedia.org/wiki/Semantic_satiation
- mikestew 2y agoIt is incredible how this word turned up a short while ago… I’m sure if I looked, I could find quotes from Warren Buffet (the recognized originator of the term) going back a few decades. But your point stands.
- mikeyouse 2y agoYeah, he's been talking about "economic moats" since at least the 1990s. At least since 1995; https://www.berkshirehathaway.com/letters/1995.html https://www.berkshirehathaway.com/letters/1995.html
- pillefitz 2y agoNobody claimed it's a new word. Still, the frequency increased 100x over the last days, subjectively speaking.
- kccqzy 2y agoThe earliest occurrence of the word "moat" that I could find online from Buffett is from 1986: https://www.berkshirehathaway.com/letters/1986.html https://www.berkshirehathaway.com/letters/1986.html That shareholder letter is charmingly old-school. Unfortunately letters before 1977 weren't available online so I wasn't able to search. It also helps that I've been to several cities with an actual moat so this word is familiar to me.
- fastasucan 2y agoThe word moat was first used in english in the 15th century https://www.merriam-webster.com/dictionary/moat https://www.merriam-webster.com/dictionary/moat
- arresin 2y agoLink has all the params but running at 4 bit quant.
- qingcharles 2y ago4-bit quant is generally kinda low, right? I wonder how badly this quant affects the output on DeepSeek?
- tw1984 2y ago> If most of NVIDIAs moat is in being able to efficiently interconnect thousands of GPUs nah. it moat is CUDA and millions of devs using CUDA aka the ecosystem
- mupuff1234 2y agoBut if it's not combined with super high end chips with massive margins that moat is not worth anywhere close to 3T USD.
- ReptileMan 2y agoAnd then some chineese startup create an amazing compiler that takes cuda and moves it to X (AMD, Intel, Asic) and we are back at square one. So far it seems that the best investment is in RAM producers. Unlike compute the ram requirements seem to be stubborn.
- 01100011 2y agoDon't forget that "CUDA" involves more than language constructs and programming paradigms. With NVDA, you get tools to deploy at scale, maximize utilization, debug errors and perf issues, share HW between workflows, etc. These things are not cheap to develop.
- Symmetry 2y agoIt might not be cheap to develop them but if you can save $10B in hardware costs by doing so you're probably looking at positive ROI.
- 01100011 2y agoYeah, I mean, 9 women can make a baby in a month so why not? Oh wait, it takes years to do all that and in the meantime you're wasting energy on not staying at the forefront of a hot tech trend.
- a_wild_dandan 2y agoRunning a 680-billion parameter frontier model on a few Macs (at 13 tok/s!) is nuts. That'a two years after ChatGPT was released. That rate of progress just blows my mind.
- qingcharles 2y agoAnd those are M2 Ultras. M4 Ultra is about to drop in the next few weeks/months, and I'm guessing it might have higher RAM configs, so you can probably run the same 680b on two of those beasts. The higher performing chips, with one less interconnect, is going to give you significantly higher t/s.