4 ms·
> Tensordyne is expected to accelerate AI inference with a logarithmic number system that leans on a property of logarithms: The log of A times B equals the log
by swimwiththebeat 17d ago
> Tensordyne is expected to accelerate AI inference with a logarithmic number system that leans on a property of logarithms: The log of A times B equals the log of A plus the log of B. So, storing numbers as their exponents lets the chip add where it would otherwise multiply. That matters in silicon because multiplier circuits draw more power and use more die area than adders do. Tensordyne says its rack-scale hardware, called Napier, can produce up to 1,300 tokens per second per user, and can do so while using less than a tenth as much power as comparable Nvidia hardware.
Did not know about this cool trick about storing numbers as exponents! Is there a name for this technique? Wouldn’t there be overhead in converting back and forth between the exponent and the number?
- gcr 17d agohang on, isn't this the standard way to implement IEEE 754 floating-point multiply since forever? normalize the two numbers A and B to have the same exponent, add the mantissa, then convert back to IEEE 754?
- tasty_freeze 17d agono. IEEE splits the power of two exponent from the base 2 mantissa. Yes, the exponents are added during a multiply, but the mantissas do an ordinary multiply. The idea is rather than storing a number x as (exponent, mantissa), just store (log x) as a fixed precision number. Multiplying two such numbers is just addition, dividing is just subtraction. TBH I didn't read the article, but my reaction is that yes, that works, but one must sum all those products, and now summing becomes an expensive operation. Maybe the total cost saves area and power, but it beggars belief that it is 10x more efficient. They must be doing PR math: our low precision log scheme is 10x more efficient than a higher precision traditional approach. Another thing to keep in mind is a lot of inference is done using very low precision math and so the cost of doing multiplies isn't that bad. Yes, it is still (n bits) squared, but as n gets small, n^2 still isn't too bad.
- hgoel 17d agoAlso I wonder how this affects the distribution and precision needed for storing the log weights compared to regular ones.
- windenntw 17d agoThis is essentially the way many 8 bit games did 3d rendering ( for example the world famous Elite )... you just need two tables, one linear2log and another log2linear, with careful measurement of the ranges and number of elements needed in you application ( which is easy for inference ). ps. Also used in the original circuits of the Yamaha DX7 synthetiser ( https://www.righto.com/2021/11/reverse-engineering-yamaha-dx7.html https://www.righto.com/2021/11/reverse-engineering-yamaha-dx... ).
- vrighter 17d agobut aren't additions in exponents expensive? as in each operation is an fma