5 ms·
If anything performance should get better with time
by pdksam 3y ago
If anything performance should get better with time
- marginalia_nu 3y agoWhy do you think that would be the case?
- Damogran6 3y agoMoore's law-ish like optimization. You Z80 computer cost $700 in the lat 70's...they're now in sub-$1 embedded controllers.
- marginalia_nu 3y agoBut what is being optimized? Hardware sure isn't getting faster in a hurry, and I don't see anything on the horizon that will aid in optimizing software.
- ben_w 3y agoThe various open source LLMs are doing things like reducing bits-per-parameter to reduce hardware requirements; if they're using COTS hardware it almost certainly isn't optimised for their specific models; Moore's Law is pretty heavily reinterpreted, so although we normally care about "operations per second at a fixed number of monies" what matters here is "joules per operation" which can improve a by a huge margin even before human level, which itself appears to be a long way from the limits of the laws of physics; and even if we were near the end of Moore's Law and there was only a 10% total improvement available, that's 10% of a big number.
- marginalia_nu 3y agoMoore's law was an effect that stemmed from the locally exponential efficiency increase from designing computers using computers, each iteration growing more powerful and capable of designing still more powerful hardware. 10% here and there is very small compared to the literal orders magnitude improvements during the reign of Moore's Law. I don't really see anything like that here.
- reitanqild 3y ago> 10% here and there is very small compared to the literal orders magnitude improvements during the reign of Moore's Law. I can't confirm it, but I noticed this comment says "gpu tech has beat Moore’s law for DNNs the last several years": https://news.ycombinator.com/item?id=35653231 https://news.ycombinator.com/item?id=35653231
- marginalia_nu 3y agoWe're actually at an inflection point where this isn't the case anymore. For a long time, GPU hardware basically became more powerful with each generation, but prices stayed roughly the same plus minus inflation. Last couple of years, this trend has broken. You pay double or even quadruple the price for a relatively tenuous increase in performance.
- Damogran6 3y agoWe said that in 1982, and 1987, and 1993, and 1995, and 2001, 2003, 2003.5 You get the point. There's always local optimization that leads to improvements. Look at the Apple M1 chip rollout as a prime example of that. Big/Little processors, on die RAM, shared memory with the GPU and Neural Engine, power integration with the OS. LOTS of things that led to a big leap forward.
- marginalia_nu 3y agoBig difference now is that we have a clear inflection point. Die processes aren't getting much smaller than they are. A sub-nanometer process would involve arranging single digit counts of atoms into a transistor. A sub-Å process would involve single atom transistors. A sub 0.5Å process would mean making them out of subatomic particles. This isn't even possible in sci-fi. You can re-arrange them for minor boosts, double the performance a few times sure, but that's not a sustained improvement month upon month like we have in the past. As anyone who has ever optimized code will attest, optimization within fixed constraints typically hits diminishing returns very quickly. You have to work harder and harder for every win, and the wins get smaller and smaller.
- CuriouslyC 3y agoThere are a few avenues. Further specialization of hardware around LLMs, better quantization (3 bits/p seems promising), improved attention mechanisms, use of distilled models for common prompts, etc.
- marginalia_nu 3y agoThis would be optimizations, which is not really the same thing as moore's law-like growth which was absolutely mind-boggling, like it's hard to even wrap your head around how fast tech was moving in that period since humans don't really grok exponentials too well, we just think they look like second degree polynomials.
- CuriouslyC 3y agoProbabilistic computing offers the potential of a return to that pace of progress. We spend a lot of silicon on squashing things to 0/1 with error correction, but using analog voltages to carry information and relying on parameter redundancy for error correction could lead to much greater efficiency both in terms of OPS/mm^2 and OPS/watt.
- kolinko 3y agoI am wondering about this as well - wondering how difficult it would be to build an analog circuit for a small LLM (7B?). And wondering if anyone's working on that yet. Seems like an obvious avenue to huge efficiency gains.
- marginalia_nu 3y agoSeems very unrealistic when considering how electromagnetic interference works. Clamping the voltages to high and low goes some way to mitigate that problem.
- CuriouslyC 3y agoThat's only an issue if the interference is correlated.
- EMM_386 3y ago> Hardware sure isn't getting faster in a hurry How is it not? These LLMs were recently trained using NVidia A100 GPUs. Now NVidia has H100 GPUs. The H100 is up to nine times faster for AI training and 30 times faster for inference than the A100.
- bobsmooth 3y agoNot soon but all the major players are making even more AI specialized silicon.
- roflyear 3y agoWhat I mean is resources will be limited or models that are slightly worse will be released that will be much more cost effective but not quite as good. This is often the case with these types of technologies.