3 ms·
I'm about to head out the door, but here's what I see. He's got two fundamental changes to the way a modern CPU is designed: * Internal representation of float
by patrickyeon 15y ago
I'm about to head out the door, but here's what I see. He's got two fundamental changes to the way a modern CPU is designed:
* Internal representation of floats is logarithmic
* Small processing elements are connected in a mesh network that is made explicit to the programmer
The outcome of the second is that there is more work for a programmer, as details of the hardware are no longer hidden. The tradeoff is that none of the overhead to hide those details is required in hardware anymore. This is not new in other application areas, I'm sure DSP folk (and probably GPU folk) can tell us all about it.
The first gives us very fast arithmetic. What I think he's describing is, instead of working with x and y, you would work with log(x) and log(y). If `x * y = z`, then `log(x) + log(y) = log(z)`. (Additionally, the logarithm of exponentiation is multiplication, so square roots are just bit shifts!) Addition can be done in an approximate manner as well, using a bit more work. This is all from slide 5.
Releasing the hardware from the requirement that it be extremely accurate (he sets up his logarithmic representation to allow for ~1% error) allows him to make all this much smaller. He's shown that it can work "well enough" in some cases, and is working on finding more applications for the tech.
This isn't a free lunch (and isn't proposed as such), and not applicable across all of general purpose computing. But I would think there are lots of areas it could improve; whether it's huge datasets and customers that can afford expensive, custom hardware, or low power application areas that would benefit from approximate methods only a full-blown desktop can handle now (pattern recognition or noise reduction on phones and cameras, for example).
- rhino42 15y agoThe other issue will be the I/O bottleneck. A practical application would need to have on-chip memory to buffer data (not mentioned here, I think). Even so, the applications will be severely limited by having data with enough processing needs. I spend most of my time working on FPGAs, and I would imagine that in the end, a practical implementation of this chip would involve a good deal of similar work. Data would have to be piped from one core to adjacent cores, to keep the inter-chip bandwidth tractable. This could result in portions on the edges left unused because the data can't get there. In summary, the applications for such a chip are limited, and will require a different skillset from what we usually ascribe to a programmer, but in those constraints these chips could be incredibly high-performance. EDIT: I haven't played with log-scale arithmetic yet, but I think that adds / subs will be much more computationally complex--probably more so than multiplies in the current linear-scale numbers. Just a thought.
- omaranto 15y agoI don't think addition and subtraction are too bad in this model. The slides mention that there is a simple circuit computing F(t):=log(1+2^t). So if you have two numbers x and y, represented by their logarithms log(x), log(y), the log of their sum can be computed using F as follows: log(x+y) = log(x) + log(1 + y/x) = log(x) + F(log(y) - log(x)). Whether this is better or worse than multiplies in the current linear-scale representation depends on how hard it is to compute F, I guess.