3 ms·
Shall we bet on when the hardware needed for this (without quantizing and at good speed) will reach < 10k USD? I'm betting 2040. I can download it now, and then
by edg5000 2mo ago
Shall we bet on when the hardware needed for this (without quantizing and at good speed) will reach < 10k USD? I'm betting 2040. I can download it now, and then get the hardware later. Eventually we can all have these things running 24/7 in our home if we wanted to. I currently would not have any task for it that would really utilize the hardware 24/7, but maybe in 20 years I will.
- arthurcolle 2mo agoapproximately $20 million for 750TB unified memory custom interconnect right now
- edg5000 2mo agoWow, that's way higher than I assumed. I looked into it and remember ending up with something like 200k, but I must have been off. That's crazy.
- arthurcolle 2mo ago[dead]
- ak_t 2mo agoI think it is more likely that a smaller model (<400B) with similar intelligence gets developed long before the hardware to serve a 2.4T model gets cheaper than 10k.
- edg5000 2mo agoAh, I hadn't thought of that. To what degree have we seen this already? What would you consider as the biggest jump in intelligence per weight?
- ak_t 2mo agoIt's been pretty consistent, the smaller models (hundreds of billions of params) usually catch up in 6 months or so to their frontier counterparts, at least on benchmarks.
- Manfrednotfunny 2mo agoI would say 5 years. The whole industry is now pushing through memory. In 5 years you have either some type of explosion which willjust make all the hardware from today affordable or you have such an AI explosion, that the today hardware is written off and not efficient enough anymore that you can buy it for cheap. In parallel, its clear that we need more memory. In parallel models in hardware will become a thing on mass market. In parallel everything gets more efficient. The 30B parameter model will be for sure more intelligent in 5 years than it is today.
- edg5000 2mo agoArguably no one was really prepared for AI, and NVidia kind of happended to coincidentally have suitable hardware. I have the feeling that something simpler and more efficient is possible when designed for the ground up purely for AI. But I could be wrong. That would mean there is a lot of room left for improvement.
- cmrdporcupine 2mo agoIt is not at all just quantities of memory (or speed of computation) It's in large part a problem of bandwidth, too. Mostly really. HBM memory can do up to 3TB/s vs DDR5 like 250GB/S. The latter is just too slow to process 2.5B parameter models, it simply can't move the values back and forth fast enough. It would drag to a crawl. Much smaller dense models at that speed on the NVIDIA Spark can't do more than 15tok/sec. Real serving systems for these models involve large numbers of parallel GPUs with massive memory bandwidth, hooked up via NVlink. It will take a long time for that level of tech to get down to consumer level. (An ideal computing architecture built for LLMs would in fact offer some way of colocating computation with memory. If you can put matmul etc right in the DRAM and avoid going back and forth over the bus...)