14 ms·
Having investigated posits by running RTL implementations through synthesis on 7 and 28 nm nodes, I don't buy the claim that a posit FPU is smaller than a IEEE
by jhj 5y ago
Having investigated posits by running RTL implementations through synthesis on 7 and 28 nm nodes, I don't buy the claim that a posit FPU is smaller than a IEEE float FPU. The implication is probably more around that one could use a 32 bit word size posit than a 64 bit word size posit, or similar, for many applications which would make this true. This is still TBD for many classical HPC problems I think though.
On an equivalent word size basis, the maximum precision of a posit (assuming reasonable exponent scale) is much larger than an IEEE float at a given word size. Adders and multipliers must be sized to handle this (e.g., a multiplication between two posits with maximal precision (1 <= x < 2) involves requires the full multiplier to handle this case). Multipliers have a quadratic dependency on the significand size. Subnormal handling in IEEE adds a lot of complexity, but not as much as a significantly larger multiplier.
- phkahler 5y ago>> Having investigated posits by running RTL implementations through synthesis on 7 and 28 nm nodes... Is it possible to implement 32 and 64 bit posits in similar area to floating point? Can the calculations be had in the same number of cycles or fewer? I feel like these are the two most important questions. If the answers are yes, then I think it may be worth doing implementations for the increased precision and simplicity (lack on NaN and other IEEE quirks).
- Someone 5y agoI don’t understand why not having NaNs would increase simplicity. Isn’t that just moving the problem of detecting various forms of over/underflow from the CPU to the programmer using it?
- adgjlsfhk1 5y agoIt's not quite right to say that Posits don't have NaN. They have a NaR (not a real) value that's fairly similar. The simplification is that they don't have 2^52 of them, and don't have -0, Inf, -Inf, or the mess with subnormals (which are necessary but not always well supported). Also, NaR compares equal to itself, fixing one of the biggest bugs in floating point.
- DaiPlusPlus 5y ago> Also, NaR compares equal to itself, fixing one of the biggest bugs in floating point. I thought that was by-design? Just like how NULL != NULL in SQL?
- adgjlsfhk1 5y agoAs far as I can tell (see https://stackoverflow.com/questions/1565164/what-is-the-rationale-for-all-comparisons-returning-false-for-ieee754-nan-values https://stackoverflow.com/questions/1565164/what-is-the-rati...), the only reason for it was that early chips had no other single instruction that could be used to implement isnan, so they decided to use == for that purpose.
- adrian_b 5y agoNo, a NaN not being equal to itself is the correct behavior for an order relation that is a partial order, not a total order. The bug is neither in the FP standard nor in the implementations. The bug is in almost all programming languages, which annoyingly provide only the 6 relational operators needed for total order relations. To deal with partial order relations, which appear not only in operations with floating-point numbers, but also in many other applications, you need 14 relational operators. Only when they are provided by the programming language you no longer need to test whether a number is a NaN, because you can always use the appropriate relational operator, depending on what the test must really do. The 2 relational operators "is equal to" and "is either equal to or unordered", are distinct operators, exactly like "is equal to" and "is equal to or greater" are distinct. If you do not have distinct symbols for"is equal to" and "is either equal to or unordered", the compiler cannot know whether the test must be true or false when a NaN operand is encountered. If you do not know which of the 2 operators must be used, you must think more about it, because your program will be wrong in exceptional cases. Specifying a stupid rule like a NaN being equal to itself and claiming that the correct behavior is a bug shows a dangerous lack of understanding of mathematics, which is not acceptable for someone designing a new number representation method (even if a single NaN value is used, it can be generated as a result of different, completely unrelated computations, and there is no meaningful way in which those results can be considered equal). In general almost all programming languages are designed by people without experience in numerical computation, who completely neglect to include in the language the features required for the operations with floating-point numbers, resulting in an usually dreadful support for the IEEE floating-point arithmetic standard, even after almost 40 years since it became widespread. It is not the standard which must be corrected, but the programming languages, which do not provide good access to what the hardware can do. Also the education of the programmers is lacking because most of them are familiar only with the 6 relational operators for total orders, instead of being equally familiar with the 14 operators for partial orders. Partial orders are encountered extremely frequently. If only the 6 relational operators are available in such cases, then they must always be preceded by a test for unordered operands. Forgetting the additional test guarantees errors. To avoid the double tests that are needed in most programming languages, it may be possible to define macros corresponding to all the 14 relational operators, but they cannot be as convenient as proper support in the language.
- FullyFunctional 5y agoHave you read the NaN rules? Know when NaN is converted to the various forms? It's maddening, truly. A posit implementation is definitely simpler from the point of view of dealing with that, but outside the NaN and Inf numbers, it's no easier (and technically a wee bit more work to deal with the regiments).
- throwaway_posit 5y agoI will say one thing: I pity the person (grad student?) that has to do error propagation analysis on a research project using posits (I'm the original implementor)
- adgjlsfhk1 5y agoYeah, I pretty much think that posit only makes sense for 32 bit and smaller, and that you want your 64 bit numbers to be closer to float64 (although with the posit semantics for Inf/NaN/-0).
- throwaway_posit 5y agoOh I actually only think it's useful for machine learning. I have some unpublished, crudely done research showing that the extended accumulator is only necessary for the Kronecker delta stage of the back propagation (posits trivially convert to higher precision by zero-padding).. you can see what I'm talking about I'm the Stanford video. Fun fact: John sometimes claims he invented the name, but this is untrue; my old college website talks about building "positronic brains" and it's long been a goal of mine to somehow "retcon" Asimov/tng data, into being a real thing, and when this opportunity for some clever wordsmithing arrived I coined the term with the hopes that someone would make a posit-based perceptron, or "positron".
- FullyFunctional 5y agoI think you are spot on and it's IMO unfortunate that John overreaches wrt. the benefits of posit as it distracts from the actual major advantages: * A much more consistent and sane floating point (most arithmetic rules _do_ apply for posits unlike for IEEE Std 754 and you never round to infinity). * Much greater range and precision for the same bits. posit32 falls somewhere between floats and double in precision and the hardware implementation will reflect this (which isn't a bad thing given the much higher space efficiency). I'm not sure about the quire as it looks expensive to me, but I haven't tried implementing it. However what it does provide is pretty remarkable: zero rounding errors for reasonable sized dot-products (IIRC < 2^32 elements).
- adgjlsfhk1 5y agoThe main problem I see with the quire idea is that John tries to use it as an argument that fma isn't necessary, and while a quire is strictly more useful, you won't be able to use multiple of them at the same time (due to hardware constraints). For applications taking dot products, this isn't a problem, but for fast evaluation of polynomials, interleaving multiple fmas is essential for good performance. As such, I think that quires are probably a really good idea for 8 and 16 bit posits, but for 32 and 64 bit, I think having an fma instruction is pretty much necessary.
- FullyFunctional 5y agoActually they have changed their position on this in the latest, now ratified spec; the quire is now just a data type. How you implement it is up to you, but you can certainly have more than one. A lot of the material about posits is out of date, including the Cult of Posits. The +/- Inf has been replace with NaR. As an aside: interestingly MININT maps to and from NaR when converting between posits and integers.
- adgjlsfhk1 5y agogood to know, but even if you can have more than 1 semantically, in hardware, they take 512 bits for 32 bit, or 2048 for 64 bit, and most CPUs only have 1 (occasionally 2) units of vector math per core, so I think it is unlikely that they would be able to efficiently work with multiple quires. 32 bit quires are possible, but 64 bit almost certainly aren't. Also, for vectorization, modern cpus are capable of doing 16x 32 bit fma at a time, but processing 16x 512 bits for a quire in a similar amount of time is totally out of the question.
- scythe 5y ago>the maximum precision of a posit (assuming reasonable exponent scale) Having run into this problem in computing the partition function of a quantum system, 10^300 is not always enough for everyone. So I don't agree with Gustafson's attempt to mimic the dynamic range of IEEE floats: give us more, since you have it. If that makes the implementation cheaper, all the better.
- Dylan16807 5y ago> On an equivalent word size basis, the maximum precision of a posit (assuming reasonable exponent scale) is much larger than an IEEE float at a given word size. What's 'much' here? If I'm doing the math right for a 64 bit number, it's 53 bits of precision vs. 59?
- throwaway_posit 5y agoYep. This works out to be on the order of 60x6x2 adders which is honestly not that much.
- adrian_b 5y agoWhenever the precision of some number representation format is said to be higher than the precision of another representation format, it must be kept in mind that this claim can be true only in a limited interval. When a number representation uses a fixed number of bits, e.g. 64 bits, there are a fixed number of points on the number axis that can be represented, e.g. 2^64. Any other representation has the same number of points. If you increase the density of the points in some interval, to increase the precision, you must take them from another interval, where fewer points will remain, resulting in a lower precision. The traditional floating-point numbers are distributed so that the relative error is almost constant as long as there is neither overflow nor underflow. For most computations that belong to simulations of complex physical systems, the physical quantities may have values varying within many orders of magnitude and a constant relative error is what is desired for maximum accuracy in the final results. The posit representation increases the precision for numbers close to 1, with the price of decreasing the precision for large or small numbers. There are applications that can benefit from the increased precision around 1, but there are also others whose precision would be seriously affected if posits were used. So it is wrong to claim that posits have greater precision in general. They have greater precision only for the applications that use only numbers that are neither too large, nor too small. It is likely that the largest benefits from posits are available only for the applications that use only small number sizes, i.e. no more than 32 bits. Such small number formats have a too small exponent range to be used in applications where very large or very small numbers are common, so if 32-bit or smaller floating-point numbers are already used, it is likely that the application belongs to those that might benefit from posits, due to the limited range of the numbers handled by it. An application that really needs 64-bit numbers, e.g. the simulation of a semiconductor device, is more likely to have its accuracy worsened than improved by posits.
- throwaway_posit 5y agoIn your RTL synthesis did you use "hidden -2 bit" for negative posits? Assuming you are "cheating" IEEE by not implementing subnormals or NaN... This is one key insight that makes posit sizes much smaller, but the algebra that you have to do to get the correct circuits is a bit trickier!
- adgjlsfhk1 5y agoCan you expand on this? This sounds really interesting.
- throwaway_posit 5y agoIf you're familiar with how the hidden bit works for IEEE floating point, use a hidden '10' in front of the fraction for negative numbers for posits and suddenly a whole bunch of math falls out. This is equivalent to having the fraction be added to a -2 value which pins the 'overall value' of the fraction to be between -1 and -2, (like how for positive values the hidden bit is 1 and the 'overall value' of the fraction is pinned to between 1 and 2).