4 ms·
They mean 25 x RTX 4090 GPUs. 4090 is a model number
by fnands 3y ago
They mean 25 x RTX 4090 GPUs.
4090 is a model number
- jasonjmcghee 3y agoAnd to give a bit more context, it's one of the top consumer grade cards available (and has 24GB of RAM). It costs on the order of $1.6k instead of $15-25k of H100.
- jiggawatts 3y agoThat just goes to show how huge the markup is on those H100 cards! Ten 4090 cards have more compute and more memory (240GB!) than a H100 card. The cost of the 4090s is the same or lower for about 4x the compute.
- mrtranscendence 3y agoI think the question stands. You can't fit 25 4090s in a robot (unless we're talking about something massive with an equally massive battery), and even if you could an LLM wouldn't be appropriate for driving a robot. Given the pace of improvements I don't see how you compress 25 4090s into a single GPU in 7 years. A 4090 isn't 25 times the power of a GTX 980, it's closer to maybe three times.
- machiaweliczny 3y agoLight or quantum could possibly deliver by that time
- LarsDu88 3y agoNeither of these are likely nor necessary to deliver the necessary results.
- LarsDu88 3y agoTeslas are already consumer items that rock massive batteries. My 25 count is off. It's probably closer to 75 GPUs right now. Let's say today's models are running vanilla transformers via pytorch without any of Deepmind's Flamingo QKV optimizations. In 10 years algorithmic optimizations, and ML platform improvements push that efficiency up 3-5 fold. We're down in the ballpark of 25 GPUs (again) Now, we ditch the general purpose GPUs entirely and go for specially built AI inference chip. The year is 2033 and specialty inference chips are better and more widespread. Jettison those ray tracing cores, and computer rendering stuff. Another 3x improvement and we're at ~8 GPUs Now you said the 4090 is about 3.5x faster than the 780 from a decade ago. We are now on the order of 2 chips to run inference. These models won't just be running in Teslas, they will be running in agricultural vehicles, military vehicles, and eventually robot baristas.
- deleted 3y ago[deleted]
- ChatGTP 3y agoI guess you’re onto something, that combined with reductive, concave, inference architecture gains and new flux capacitor designs we are on 10x performance maximum optimisation curve. I can taste that robot barista cappuccino now.