12 ms·
Addition is all you need for energy-efficient language models
- md_rumpf 2y agoThe return of the CPU?!
- anticensor 2y agoThe reign of Threadripper!
- visarga 2y ago> can potentially reduce 95% energy cost by elementwise floating point tensor multiplications and 80% energy cost of dot products It this were about convolutional nets then optimizing compute would be a much bigger deal. Transformers are lightweight on compute and heavy on memory. The weakest link in the chain is fetching the model weights into the cores. The 95% and 80% energy reductions cited are for the multiplication operations in isolation, not for the entire inference process.
- lifthrasiir 2y agoI'm also sure that fp8 is small enough that multiplication can really be done in a much simpler circuit than larger fp formats. Even smaller formats like fp4 would be able to just use a lookup table, and that makes them more like sort-of-standardized quantization schemes.
- tankenmate 2y agoi suspect that you could do fp8 with log tables and interpolation if you really wanted to (compared to the memory required for the model it's peanuts), it just turns into a LUT (log table look up) and bit shift (interpolation). so again, memory bandwidth is the limiting factor for transformers (as far as energy is concerned).
- lifthrasiir 2y agoThis time though LUT exists in a circuit, which is much more efficient than typical memory lookup. Such LUT would have to exist per each ALU though, so it can't be too large.
- brilee 2y agofp4/fp8 for neural networks don't work the way you think they do - they are merely compression formats - a set of, say, 256 fp32 weights from 1 neuron are lossily turned into 1 max value (stored in fp32 precision) and 256 fp4/fp8 numbers. Those compressed numbers are multiplied by the fp32 number at runtime to restore the original weights and full fp32 multiplication + additions are executed.
- imjonse 2y agoWith w8a8 quantization the hw (>= hopper) can do the heavy math in fp8 twice as fast as fp16.
- SuchAnonMuchWow 2y agoThe goal of this type of quantization is to move the multiplication by the fp32 rescale factor outside of the dot-product accumulation. So the multiplications+additions are done on fp8/int8/int4/whatever (when the hardware support those operators of course) and accumulated in a fp32 or similar, and only the final accumulator is multiplied by the rescale factor in fp32.
- lifthrasiir 2y agoYou are correct that the accumulation (i.e. additions in dot products) has to be done in a higher precision, however the multiplication can still be done via LUT. (Source: I currently work at a hardware-accelerated ML hardware startup.)
- rajnathani 2y agoThat's how Nvidia's mixed precision training worked with FP32-FP16, but it isn't the case for Bfloat16 on TPUs and maybe (I'm not sure) FP8 training on Nvidia Hopper GPUs.
- bee_rider 2y agoWhat is fp4? 3 bits of exponent and one of mantissa?
- wruza 2y agoSEEM (sign, exp, mantissa)
- bee_rider 2y agoInteresting… I guess it must be biased, m*2^ee would leave like half of the limited space wasted, so 1.m*2^ee? I always wonder with these tiny formats if 0 should even be represented…
- wruza 2y agoI’m not a binary guy that much, but iirc all floats are 1.m*2^e — “1.” is always there except for subnormals. There’s also SEEE FP4 which is basically +-2^([u?]int3). https://medium.com/@harrietfiagbor/floating-points-and-deep-learning-dl-supports-8ee35053ea01 https://medium.com/@harrietfiagbor/floating-points-and-deep-...
- kendalf89 2y agoMaybe this technique can be used for training then since that is a lot more compute intensive?
- SuchAnonMuchWow 2y agoIts worse than that: the energy gains are when comparing computations made with fp32, but for fp8 the multipliers are really tiny and the adder/shifters represent a largest part of the operators (energy-wise and area-wise) and this paper will only have small gains. On fp8, the estimated gate count of fp8 multipliers is 296 vs. 157 with their technique, so the power gain on the multipliers will be much lower (50% would be a more reasonable estimation), but again for fp8 the additions in the dot products are a large part of the operations. Overall, its really disingenuous to claim 80% power gain and small drop in accuracy, when the power gain is only for fp32 operations and the small drop in accuracy is only for fp8 operators. They don't analyze the accuracy drop in fp32, and don't present the power saved for fp8 dot product.
- bobsyourbuncle 2y agoI’m new to neural nets, when should one use fp8 vs fp16 vs fp32?
- deleted 2y ago[deleted]
- reissbaker 2y agoBasically no one uses FP32 at inference time. BF16/FP16 is typically considered unquantized, whereas FP8 is lightly quantized. That being said there's pretty minimal quality loss at FP8 compared to 16-bit typically; Llama 3.1 405b, for example, only benchmarks around ~1% worse when run at FP8: https://blog.vllm.ai/2024/07/23/llama31.html https://blog.vllm.ai/2024/07/23/llama31.html Every major inference provider other than Hyperbolic Labs runs Llama 3.1 405b at FP8, FWIW (e.g. Together, Fireworks, Lepton), so to compare against FP32 is misleading to say the least. Even Hyperbolic runs it at BF16. Pretraining is typically done in FP32, although some labs (e.g. Character AI, RIP) apparently train in INT8: https://research.character.ai/optimizing-inference/ https://research.character.ai/optimizing-inference/
- tarasglek 2y agoSambaNova does bf16
- imjonse 2y agoThat is true for single user/light inference only. For training and batch inference you can get compute bound fast enough.
- saagarjha 2y agoThat really depends on what you're doing. Trying to feed a tensor core is pretty hard–they're really fast.
- woadwarrior01 2y agoPre-fill (even in the single batch case) and multi-batch decoding are still compute dominated. The oft repeated trope of "decoder only transformer inference is bottle-necked on memory bandwidth" is only strictly true in the single batch decoding case, because you're mostly doing vector matrix mults when the batch size is one.
- ein0p 2y agoNot even single batch. If you want reasonable latency per token (TPOT) even larger batches do not give you high compute utilization during extend. It’s only when you don’t care about TPOT at all, and your model is small enough to leave space for a large batch on an 8 GPU host, that’s when you could get decent utilization. That’s extend only - it’s easy to get high utilization in prefill.
- mikewarot 2y agoImagine if you had a systolic array large enough that all the weights would only have to be loaded once at startup. Eliminating the memory-compute bottleneck of the von Neumann architecture could make this quite a bit more efficient.
- api 2y agoSounds like the awesome architecture for transformers would be colocation of memory and compute.
- Joker_vD 2y agoYes, that's why we generally run them on GPUs.
- moffkalast 2y agoGPUs that pull a kilowatt when running yes. This might actually work on an FPGA if the addition doesn't take too many clock cycles compared to matmuls which were too slow.
- phkahler 2y agoThat's why we need a row of ALUs in RAM chips. Read a row of DRAM and use it in a vector operation. With the speed of row reading, the ALU could take many cycles per operation to limit area.
- namibj 2y agoThe big problem is that DRAM is extremely secretive about their processes, and they largely don't do that well for logic.
- api 2y agoGPUs are better but I'm thinking of even tighter coupling, like an integrated architecture.
- h_tbob 2y agoBro... they are NOT lightweight on compute!
- cpldcpu 2y agoIt puzzles me that there does not seem to be a proper derivation and discussion of the error term in the paper. It's all treated indirectly way inference results.
- Lerc 2y agoThe paper has an odd feel about it to me too. Doing a gate estimation as a text explanation without a diagram makes it too easy to miss some required part. It wouldn't need to be a full gate level explanation but blocks labeled 'adder'. Seeing the name de Vries in the first paragraph didn't help my sense of confidence either.
- brcmthrowaway 2y agoBecause of the twisted mentat?
- Lerc 2y agoNo more because of things like http://blog.zorinaq.com/bitcoin-electricity-consumption/ http://blog.zorinaq.com/bitcoin-electricity-consumption/ It's a long read to go over multiple years worth of posts and comments but gives you a measure of the man.
- CGamesPlay 2y agoI believe this reduces the compute required, but still uses 8 bits per value, so it does not reduce the memory requirements required to run inference, so it doesn’t particularly make the models more accessible for inference. Is this storage method suitable for training? That could potentially be an interesting application.
- Manabu-eo 2y agoIt actually is about 0.5 bits less efficient per weight in terms of precision/range, something the paper never highlights.
- scotty79 2y agoAll You Need is Considered Harmful.
- TaurenHunter 2y agoWe will need a paper titled '"Considered Harmful" Articles is All You Need' to complete that cycle.
- js8 2y agoHaven't read it, but isn't this just logarithmic tables in some form? I am asking not to dismiss it, I genuinely feel I don't understand logarithms on a fundamental level (of logic gates etc.). If multiplication can be replaced with table lookup and addition, then there has to be a circuit that gives you difficult addition and easy multiplication, or any combination of those tradeoffs.
- pclmulqdq 2y agoYes, this is logarithmic number systems at work.
- sabhiram 2y agoLog space is nice, multiplication can be replaced by addition. This part is easy and anyone can implement hardware to do this. The tricky bit is always the staying in log space while doing accumulations, especially ones across a large range.
- pjc50 2y ago"We recommend training and hosting L-Mul-based models on devices integrated with specialized architectural designs. Patent pending" (from footnote in method section)
- ranguna 2y agoI've seen this claim a few time across the last couple years and I have a pet theory why this isn't explored a lot: Nvidia funds most research around LLMs, and they also fund other companies that fund other research. If transformers were to use addition and remova all usage of floating point multiplication, there's a good chance the gpu would no longer be needed, or in the least, cheaper ones would be good enough. If that were to happen, no one would need nvidia anymore and their trillion dollar empire would start to crumble. University labs get free gpus from nvidia -> University labs don't want to do research that would make said gpus obsolete because nvidia won't like that. If this were to be true, it would mean that we are stuck on an inificient research path due to corporate greed. Imagine if this really was the next best thing, and we just don't explore it more because the ruling corporation doesn't want to lose their market cap. Hopefully I'm wrong.
- yieldcrv 2y agoAlternatively, other people fund LLM research
- chpatrick 2y agoIt's still a massively parallel problem suited to GPUs, whether it's float or int, or addition or multiplication doesn't really matter.
- teaearlgraycold 2y agoNVidia GPUs support integer operations specifically for use with deep learning models.
- londons_explore 2y agoIf an addition-only LLM performed better, nvidia would probably still be the market leader. Next gen nvidia chips would have more adders and fewer multipliers.
- cpldcpu 2y agoI have to disagree. Nvidia spent a lot of effort on researching improved numerical representations. You can see a summary in this talk: https://www.youtube.com/watch?v=gofI47kfD28 https://www.youtube.com/watch?v=gofI47kfD28 A lot of their work was published but went by unnoticed. But in fact the majority of their performance increase in new architecture is resulting from this work. Reading between the lines, it seems that they came to the conclusion that a 4 bit representation with a group exponent ("FP4") is the most efficient representation of weights for inference. Reducing the number of bits in weights has the biggest impact on LLMs inference, since they are mostly memory bound. At these low bit numbers, the impact of using multiplication or other approaches is not really significiant anymore. (multiplying a 4 bit wight with a larger activation is effectively 4 additions, barely more than what the paper proposes)
- concrete_head 2y agoJust too add an alternative addition based architecture into the mix. https://www.youtube.com/watch?v=VqXwmVpCyL0 https://www.youtube.com/watch?v=VqXwmVpCyL0
- cpldcpu 2y agoBill Dally from nvidia introduced a log representation that basically allows to replace a multiplication with an add, without loss of accuracy (in contract to proposal above) https://youtu.be/gofI47kfD28?t=2248 https://youtu.be/gofI47kfD28?t=2248
- nickpsecurity 2y agoPaper? https://research.nvidia.com/publication/2022-12_lns-madam-low-precision-training-logarithmic-number-system-using-multiplicative https://research.nvidia.com/publication/2022-12_lns-madam-lo...
- Buttons840 2y agoWould using this neural network based on integer addition be faster? The paper does not claim it would be faster, so I'm assuming not? What about over time? If this L-Mul (the matrix operation based on integer addition) operation proved to be much more energy efficient and became popular, would new hardware be created that was faster?
- tantalor 2y ago[2023] GradIEEEnt half decent: The hidden power of imprecise lines http://tom7.org/grad/murphy2023grad.pdf http://tom7.org/grad/murphy2023grad.pdf Also in video form: https://www.youtube.com/watch?v=Ae9EKCyI1xU https://www.youtube.com/watch?v=Ae9EKCyI1xU
- indrora 2y agoI had hoped that they would reference this in their paper as some kind of "supporting previous exploration" but no, alas.
- dang 2y agoGradIEEEnt half decent: The hidden power of imprecise lines [video] - https://news.ycombinator.com/item?id=36806970 https://news.ycombinator.com/item?id=36806970 - July 2023 (9 comments) GradIEEEnt half decent - https://news.ycombinator.com/item?id=35780921 https://news.ycombinator.com/item?id=35780921 - May 2023 (32 comments)
- A4ET8a8uTh0 2y agoUhh.. I hate to be the one to ask this question, but shouldn't we be focused on making LLMs work well first and then focused on desired optimizations? Using everyone's car analogy, it is like making sure early cars are using lower amount of coal. It is a fool's errand.
- Maken 2y agoThe optimizations described could easily work on other models, not just transformers. Following your analogy, this is optimizing plumbing, pistons and valves on steam engines, it could be useful for whatever follows.
- lukev 2y agoAlso, making neural networks faster/cheaper is a big part of how they advance. We've known about neural architectures since the 70s, but we couldn't build them big enough to be actually useful until the advent of the GPU. Similarly, the LLM breakthrough was because someone decided it was worth spending millions of dollars to train one. Efficiency improvements lower that barrier for all future development (or alternatively, allow us to build even bigger models for the same cost.)
- itishappy 2y agoCoal (and even wood!) powered cars actually existed long before Ford, but didn't take off because they were too heavy and unwieldly. The Model T was the result of a century of optimization. https://en.wikipedia.org/wiki/Nicolas-Joseph_Cugnot https://en.wikipedia.org/wiki/Nicolas-Joseph_Cugnot
- spencerchubb 2y agoCheaper compute is basically a prerequisite to making better models. You can get some improvements on the margins by making algorithms better with current hardware, but not an order of magnitude improvement. When there is an order of magnitude improvement in hardware, the AI labs will figure out an algorithm to best take advantage of it.
- fennecfoxy 2y agoYou're also welcome to contribute. There are many people doing many things at once in this space, I don't think experiments like this are a problem at all.
- shrubble 2y agoI remember that many years ago, when floating point computation was expensive for Intel CPUs to do, there were multiple ways that programmers used integer trickery to work around this. Chuck Moore of Forth fame demonstrated taking the value, say 1.6 multiplied by 4.1 and doing all the intermediate calculations via integers (16 * 41) and then formatting the output by putting the decimal point back in the "right place"; this worked as long as the range of floating point values was within a range that multiplying by 10 didn't exceed 65536 (16 bit integers), for instance. For embedded chips where for instance, you have an analog reading with 10 bits precision to quickly compute multiple times per second, this worked well. I also recall talking many years ago with a Microsoft engineer who had worked with the Microsoft Streets and Trips program (https://archive.org/details/3135521376_qq_CD1 https://archive.org/details/3135521376_qq_CD1 for a screenshot) and that they too had managed to fit what would normally be floating point numbers and the needed calculations into some kind of packed integer format with only the precision that was actually needed, that was faster on the CPUs of the day as well as more easily compressed to fit on the CDROM.
- candiddevmike 2y agoAFAIK this is still the best way to handle money/financial numbers.
- amanda99 2y agoThat's got nothing to do with perf tho.
- Maxatar 2y agoNothing to do with perf is a strong claim. If you genuinely don't care about performance you can use an arbitrary-precision rational number representation. But performance often matters, so you trade off precision for performance. I think people are wrong to dismiss floating point numbers in favor of fixed point arithmetic, and I've seen plenty of fixed point arithmetic that has failed spectacularly because people think if you use it, it magically solves all your problems... Whatever approach you take other than going all in with arbitrary precision fractions, you will need to have a good fundamental understanding of your representation and its trade-offs. For me personally I use floating point binary and adjust the decimal point so I can exactly represent any value to 6 decimal places. It's a good trade-off between performance, flexibility, and precision. It's also what the main Bitcoin implementation uses.
- deleted 2y ago[deleted]
- ein0p 2y agoMore than 10x the amount of energy is spent moving bytes around. Compute efficiency is not as big of an issue as people think. It’s just that the compute is in the wrong place now - it needs to be right next to memory cells, bypassing the memory bus, at least in the initial aggregations that go into dot products.
- entropicdrifter 2y agoThis could still be useful for battery constrained devices, right?
- ein0p 2y agoIt’s even worse in battery constrained devices - they tend to also be memory constrained and run with batch size 1 during extend. IOW the entire model (or parts thereof, if the model is MoE), gets read for every generated token. Utilization of compute is truly abysmal in that case and almost all energy is spent pushing bytes through the memory bus, which on battery powered devices doesn’t have high throughput
- presspot 2y agoFrom my experience, the absolute magicians in fixed point math were the 8-bit and 16-bit video game designers. I was in awe of the optimizations they did. They made it possible to calculate 3D matrix maths in real time, for example, in order to make the first flight simulators and first person shooter games.
- hinkley 2y agoRedefining degrees to be 2pi = 256 was a pretty clever trick.
- m3kw9 2y agoSo instead of say 2x3 you go 2+2+2?
- dwrodri 2y ago7 years of the same title format is all you need.
- alvinadiaz2 2y ago[dead]
- maryfriese57 2y ago[dead]
- jenda23 2y agoHighly recommended!! Success achieved! Previously I had worked with another well regarded company to attempt recovering an Ethereum presale wallet passphrase that I had forgotten. After 14 months of trying there was no success, so then I looked into ReWallet. They were able to find the password solution in 6 weeks! Since I only remembered a few portions or clues, it seemed like a nearly impossible task. They worked diligently and very professionally. I fully recommend and trust these guys, the result speaks for itself. Contact email, rewalletshieldcoinrecovery@ aol.com or WhatsApp::+1 (757) 332-1885