4 ms·
I've seen this claim a few time across the last couple years and I have a pet theory why this isn't explored a lot: Nvidia funds most research around LLMs, and
by ranguna 2y ago
I've seen this claim a few time across the last couple years and I have a pet theory why this isn't explored a lot:
Nvidia funds most research around LLMs, and they also fund other companies that fund other research. If transformers were to use addition and remova all usage of floating point multiplication, there's a good chance the gpu would no longer be needed, or in the least, cheaper ones would be good enough. If that were to happen, no one would need nvidia anymore and their trillion dollar empire would start to crumble.
University labs get free gpus from nvidia -> University labs don't want to do research that would make said gpus obsolete because nvidia won't like that.
If this were to be true, it would mean that we are stuck on an inificient research path due to corporate greed. Imagine if this really was the next best thing, and we just don't explore it more because the ruling corporation doesn't want to lose their market cap.
Hopefully I'm wrong.
- yieldcrv 2y agoAlternatively, other people fund LLM research
- chpatrick 2y agoIt's still a massively parallel problem suited to GPUs, whether it's float or int, or addition or multiplication doesn't really matter.
- teaearlgraycold 2y agoNVidia GPUs support integer operations specifically for use with deep learning models.
- londons_explore 2y agoIf an addition-only LLM performed better, nvidia would probably still be the market leader. Next gen nvidia chips would have more adders and fewer multipliers.
- cpldcpu 2y agoI have to disagree. Nvidia spent a lot of effort on researching improved numerical representations. You can see a summary in this talk: https://www.youtube.com/watch?v=gofI47kfD28 https://www.youtube.com/watch?v=gofI47kfD28 A lot of their work was published but went by unnoticed. But in fact the majority of their performance increase in new architecture is resulting from this work. Reading between the lines, it seems that they came to the conclusion that a 4 bit representation with a group exponent ("FP4") is the most efficient representation of weights for inference. Reducing the number of bits in weights has the biggest impact on LLMs inference, since they are mostly memory bound. At these low bit numbers, the impact of using multiplication or other approaches is not really significiant anymore. (multiplying a 4 bit wight with a larger activation is effectively 4 additions, barely more than what the paper proposes)
- nayroclade 2y ago"Good enough" for what? We're in the middle of an AI arms race. Why do you believe people would choose to run the same LLMs on cheaper equipment instead of using the greater efficiency to train and run even larger LLMs? Given LLM performance seems to scale with their size, this would result in more powerful models, which would grow the applicability, use and importance of AI, which would in turn grow the use and importance of Nvidia's hardware. So this theory doesn't really stack up for me.
- raincole 2y ago> I have a pet theory You mean you have a conspiracy theory. Why wouldn't other companies that buy Nvidia GPU fund these researches? It would greatly cut their cost.
- yunohn 2y agoGoogle & Apple already run custom chips, Meta and MS are deploying their own soon too. Your theory is that none of them have researched non-matrix-multiplication solutions before investing billions?
- miohtama 2y agoThere are several patents on this topic so they have
- twoodfin 2y agoI’d estimate that fraction of Nvidia’s dominance that’s dependent on their distinctive advantages in kernel primitives (add vs. multiply) would be a rounding error in FP8. The CUDA tooling and ecosystem, VLSI architecture, organizational prowess… all matter at multiple orders of magnitude more.
- iamgopal 2y agono matter how fast cpu, network and browser has become, websites are still slow. we will run out of data to train much earlier than people will stop inventing even larger models.
- WrongAssumption 2y agoSo let me get this straight. Universities don’t want to show that Nvidia gpus are obsolete, so they can receive a steady stream of obsolete gpus? For what possible reason, that doesn’t make sense.