4 ms·
I am sorry for Intel. Perhaps John Gustafson’s 16 bit posits or unums would have made a better choice.
by beckerdo 7y ago
I am sorry for Intel. Perhaps John Gustafson’s 16 bit posits or unums would have made a better choice.
- H8crilA 7y agoWhy? Google has certainly researched their floats before committing an entire line of silicon chips. It's easy to just enumerate all possible float16 configurations in a simulator to see which one performs best on a wide range of neural network applications. Then pick the best one. Big data driven organizations do this all the time (brute force through an entire line of solutions, pick best results).
- deleted 7y ago[deleted]
- TazeTSchnitzel 7y agoAlso, a format that's just a slight tweak on IEEE-754 is going to be way easier to implement on existing hardware than a wholly new format.
- mochomocha 7y agoI know nothing about ASIC or CPU simulators but I suspect that it's not as easy as you make it sound: for machine-learning related tasks, performance doesn't only come from raw compute numbers: you'll also want to model the actual data movement costs across the caches hierarchy and registers. Because a lot of time training is not necessarily compute-bound: the relative cost of data transfer (VS compute) can be quite high, or even dominate.
- exsf0859 7y agobfloat16 is "best" primarily because it has the same exponent range as float32. That makes it easy to port models that have been developed using float32 to bfloat16. (As opposed to using int8 or float16, both of which have a smaller exponent range.) It's possible that some other custom format is better in absolute terms for models trained specifically for the custom format. But for the current ecosystem, where models are trained primarily using float32, bfloat16 is a very good choice.
- TomVDB 7y agoFor the same number of bits, posits are quite a bit more expensive to implement in terms of area than traditional floats.
- dnautics 7y agoThat's not true, as I have implemented both, in FPGA. The INRA implementation missed key optimizations in the adder and multiplier.
- TomVDB 7y agoThe INRIA paper was indeed my reference. It seems like a huge mistake if they missed key optimizations, but I'm happy to take your word for it. Are there write-ups that go in detail about these mistakes? It's the kind of somebody-is-wrong-on-the-Internet topic that would result in flaming blog posts. :-)
- jdsully 7y agoThat paper used High Level Synthesis which would be the equivalent of coding something in ruby and comparing it with another algorithm written in optimized assembly.
- dnautics 7y agoNot really. If you look at the inria pseudocode they check if the posit is negative or positive before doing addition, and convert, in the style of 754 one's complement encoding, but you shouldn't need to do that with posits since the encoding is two's complement. I mean, I helped design the posit spec and the twos complements treatment is something not even John Gustafson understands... The key insight is that the hidden bit is -2 for negative numbers (instead of 1 as it is for positive numbers). It's kind of nonobvious and I happened upon it by accident one night while fooling around with circuit diagrams. If people really get serious about it I'm sure though that it will get rediscovered by EDA folks smarter than I.
- josefx 7y agoWhy would you want to introduce a floating point type with completely different and incompatible behavior when you can just change the mantissa of an existing one and reuse everything already build around it?