13 ms·
I think you are spot on and it's IMO unfortunate that John overreaches wrt. the benefits of posit as it distracts from the actual major advantages: * A much mo
by FullyFunctional 5y ago
I think you are spot on and it's IMO unfortunate that John overreaches wrt. the benefits of posit as it distracts from the actual major advantages:
* A much more consistent and sane floating point (most arithmetic rules _do_ apply for posits unlike for IEEE Std 754 and you never round to infinity).
* Much greater range and precision for the same bits.
posit32 falls somewhere between floats and double in precision and the hardware implementation will reflect this (which isn't a bad thing given the much higher space efficiency).
I'm not sure about the quire as it looks expensive to me, but I haven't tried implementing it. However what it does provide is pretty remarkable: zero rounding errors for reasonable sized dot-products (IIRC < 2^32 elements).
- adgjlsfhk1 5y agoThe main problem I see with the quire idea is that John tries to use it as an argument that fma isn't necessary, and while a quire is strictly more useful, you won't be able to use multiple of them at the same time (due to hardware constraints). For applications taking dot products, this isn't a problem, but for fast evaluation of polynomials, interleaving multiple fmas is essential for good performance. As such, I think that quires are probably a really good idea for 8 and 16 bit posits, but for 32 and 64 bit, I think having an fma instruction is pretty much necessary.
- FullyFunctional 5y agoActually they have changed their position on this in the latest, now ratified spec; the quire is now just a data type. How you implement it is up to you, but you can certainly have more than one. A lot of the material about posits is out of date, including the Cult of Posits. The +/- Inf has been replace with NaR. As an aside: interestingly MININT maps to and from NaR when converting between posits and integers.
- adgjlsfhk1 5y agogood to know, but even if you can have more than 1 semantically, in hardware, they take 512 bits for 32 bit, or 2048 for 64 bit, and most CPUs only have 1 (occasionally 2) units of vector math per core, so I think it is unlikely that they would be able to efficiently work with multiple quires. 32 bit quires are possible, but 64 bit almost certainly aren't. Also, for vectorization, modern cpus are capable of doing 16x 32 bit fma at a time, but processing 16x 512 bits for a quire in a similar amount of time is totally out of the question.
- Dylan16807 5y ago> Also, for vectorization, modern cpus are capable of doing 16x 32 bit fma at a time, but processing 16x 512 bits for a quire in a similar amount of time is totally out of the question. I'd be worried about space, but I'm not sure why time would be an issue? I'd assume a big quire would have some kind of delayed carry mechanism.
- adgjlsfhk1 5y agoeven with a delayed carry, you still need to pump the data through the CPU. For a basic place where this becomes a problem, consider taking exp of each element in a vector. A vectorized version of this code will spend most of it's time computing a polynomial which is just 8 fmas (in 2 chains of 4). With a dedicated fma instruction, the cpu can do each of those fmas in 4 cycles, and overlap the chains, leading to a total time of 17 cycles (assuming 4 cycle fma which is fairly standard). Doing the same computation with quires would require moving around 16x as much data, which is impossible to do as quickly.
- Dylan16807 5y agoIf you leave the quires in place you don't have that issue, though. Do your eight multiplies, feed them to a single adder, collapse the result.
- adgjlsfhk1 5y agoYou can't keep that much memory in registers easily. Just storing all those quires would 16 512 bit registers (which on x86 is all you get)
- Dylan16807 5y agoHence why my initial comment was "I'd be worried about space, but I'm not sure why time would be an issue?"
- deleted 5y ago[deleted]
- jhj 5y agoThe energy requirement for a quire is very high unless you are talking less than 12 bit posits or so (in which case the strategy actually becomes superior from what I've seen due to the lack of needing to convert back to a float/posit via rounding). >500 flops to hold state, or a >500 bit RAM for a (32, 2) posit quire uses a crapload of power compared to just a 32 bit register. Pipelining becomes difficult too, as carries mean that potentially every bit held in the quire needs to be updated, so in the naive implementation you need to wait for the latency of a >500 bit adder. You can solve this partially by bucketing based on where in the quire a posit value expanded into fixed point should be added as fixed point addition is associative (ignoring overflow) so in a single cycle you don't need to wait for everything and can re-order as needed, but there's still the potential that a carry would need to propagate over the entire quire. Also, this means additional flop/RAM state and more bookkeeping which burns even more energy. Maybe you can get more exotic with asynchronous logic (like Nvidia did a paper on recently for handling logarithmic addition), but good luck verifying timing and everything for that. It's been quite a while since I've run numbers on this, but I wouldn't be surprised if you could get 4-10+ posit non-quire FMAs in the same power budget as a single posit FMA using a quire, and fit a lot more onto a chip. Don't get me wrong, I think the posit is a wonderful idea (make the bits you are storing in memory more meaningful), but I think a lot of Gustafson's proposals hinges upon the quire being available too ("the end of error"), which is a strong ask save for the few applications that actually need it.