13 ms·
Floating Point Visually Explained
- vxNsr 9y agoI was hoping for something akin to a xkcd or SMBC comic, this isn't really much better than what my asm prof said when he explained it for MIPS programming. Maybe it's because I don't get what he means by offset and window, but this wasn't really that helpful.
- stilist 9y agoThe window limits the value the offset can represent -- minimum and maximum bounds. The window doubles each time you increment it since it's binary / power-of-two (0, 1, 2, 4, 8, …). Within any given window there are 2^23 possible values. If you imagine those values as an array, the offset is the array index. So in the example of 3.14, the window is [2, 4] since that's the power-of-two pair that contains the value. The offset is (value - minimum) / (maximum - minimum) = 0.57, which you multiply by 2^23 to get the array index. Because the window doubles each time it's incremented, as the window gets bigger you lose precision -- in his example, [0, 1] gives 15 decimal places but [2048, 4096] only gives 4 decimal places.
- throwaway2016a 9y agoThank you. I am certain the article said that somewhere but that didn't click until your explanation.
- katastic 9y agoThe window is an index into a consecutive block of numbers (like a 4K page), and the offset is the offset into that block. [edit: the other poster highlighted the important distinction--that the window doubles in size each time so your fixed offsets into that window become further apart. If you had 1024 offsets into a 1024 window, you get index 0 for 0. index 1 for 1. 2 for 2. etc. But when you double that window, you're now one index for two steps. #1 for 2, #2 for 4, #3 for 6. and then you add that onto the top of the previous window's starting point, which in my example was 0-1023. So add 1023 for the second block. But a diagram would be really helpful at this point.] Now, I don't see the huge benefit of introducing a whole extra topic (pointer arithmetic) to explaining floating point. But if it works for someone, more power to them! [edit: Also, since pointers don't normally swell/scale/double like floating point, then even if you explain using the pointer analogy... you then have to NOTE the difference. So I still think it's a really bad analogy to use. Even worse when tons of modern day programmers know almost nothing about system programming and addressing modes.] Anyway, best of wishes. The guy certainly put a lot of work into writing an article he thought would be helpful.
- s17n 9y agoYeah it's really not a great explanation, IMO.
- tpeo 9y agoPersonally, I'm grateful that it isn't a comic. Between some stick figures plus a diagram and just the plain diagram, I'll take the latter. Comics might help writers order their thoughts because due to the constrained and sequential nature of speech bubbles, but generally they don't add much to the explanation itself.
- dragontamer 9y agoHere's everything you need to know about Floating Point in as shortly as I can write it. 1. Floating points are simply "Binary Scientific notation". The speed of light is 2.98E8... which in "normal form" is written 298,000,000. An IEEE 754 Single has 8-bits for the exponent (E8 in the speed of light), and 24-bits for the mantissa (the 2.98 part). There's some complicated stuff like offset shifting here, but this is the "core idea" of floating point. 2. "Rounding" is forced to happen in Floating Point whenever "information drops off" the far side of the mantissa. The mantissa is only 24-bits long, and many numbers (such as .1) require an infinite number of bits to represent! As such, this "rounding error" builds up exponentially the more operations you perform. 3. Subtraction (cancellation error) is the biggest single source of error and the one that needs to be most studied. "Subtraction" can occur when a positive and negative number is added together. 4. Because of this error (and all errors!), Floating point operations are NOT associative. (A + B) + C gives a different value than A + (B + C). The commutative property remains for multiplication and addition (A+B == B+A). If you require "bit-perfect" and consistent floating-point simulations, you MUST take into account the order of all operations, even simple addition and multiplication. For example: Try "0.1 + 0.7 + 1" vs "1 + 0.1 + .7" in Python, and you'll see that these to orderings lead to different results. --------------- Once you fully know and understand these 4 facts, then everything else is just icing on the cake. For example to prevent "cancellation error" (#3), you can sort the numbers by magnitude, and then add them up from smallest magnitude to largest magnitude.
- seattleeng 9y agoThis is an excellent short list of the main points I learned as an undergrad, and what I've retained today. One other point is that 64 bit floating points (aka doubles) really are about double the precision of 32 bit floating points, in terms of significant digits. An interesting implementation detail is that they spend a higher fraction of their "bit budget" on the mantissa to get this (52 out of 64 bits, vs 23 out of 32).
- deleted 9y ago[deleted]
- 9y ago
- kehrlann 9y agoBest simple explanation I've seen so far. I highly recommend Fabien's Game Engine Black Book. I'm halfway through it, and it's really fun. I've only been a software dev for 6 years, so looking at how things could be hacked around in the 90s to squeeze every drop of performance out of very constrained devices is fascinating.
- s17n 9y agoI dunno all this talk of "windows" and "buckets" etc doesn't seem particularly simple to me.
- kehrlann 9y agoI guess I'm not math-y enough so that I intuitively understand the simple math formula. On the other hand, I like his window image : e.g. I understand better how, the further you get from 0, the more your precision goes down, because "bigger windows, divided in the same number of buckets".
- conistonwater 9y agoSo the question is: if you feel like you don't get the math formula, but you get the windows-and-buckets explanation of it, is it still possible that your understanding doesn't match the true underlying concept? Because that is the pitfall with a lot of intuitive explanations, that unless you are sure that the explanations are equivalent to the true thing, you might end up understanding an idea that is close but slightly off. So a puzzle: if two positive numbers in exactly the same window are subtracted, what is the worst-case rounding error you can get in the result?
- sharpneli 9y agoError relative to the original numbers or to the result?
- deleted 9y ago[deleted]
- pmelendez 9y agoOff topic: The response of dragontamer is one example about why down votes alone are not enough. It was downvoted to the dead level and now nobody can reply to it. But also nobody gave the reason why what he was saying is incorrect.
- dragontamer 9y agoHuh, that's literally never happened to me before. Seems like an odd way of holding a discussion. Anyway, I'm still here to defend what I've stated down there. I admit that there's a lot of shortcuts in thinking in my post, but I'm trying to keep it short so that people's time isn't wasted reading my comment.
- deleted 9y ago[deleted]
- gumby 9y agoI agree. I read HN with dead comments not suppressed and I estimate about 50% of the time I consider them legit. Not always what I agree with, but sometimes I even want to reply to them.
- noxToken 9y agoIf you think a dead comment should be seen by the community, vouch for it. It's a way for the community to help moderate the discussion by allowing habitual offenders to still input relevant discussion.
- gumby 9y agoOnce it's marked dead (and not simply downvoted to 0) I don't see a way to upvote it back to life nor a way to reply to it. Perhaps other people do. Otherwise I would definitely do so!
- cesarb 9y agoTry clicking on the time ("6 hours ago") to open the comment by itself, for me that makes the "vouch" option appear.
- s17n 9y agoA much easier and better way to understand floating point is to just do it in base 10.
- kehrlann 9y agoOoooh I see. I like it, thanks !
- haberman 9y agoI like the "window/offset" concept. I wrote an extended blog article with yet different visual aids: http://blog.reverberate.org/2014/09/what-every-computer-programmer-should.html http://blog.reverberate.org/2014/09/what-every-computer-prog...
- xkpd3 9y agoExcellent post. I like it better then the one posted here.
- endorphone 9y agohttps://dennisforbes.ca/index.php/2017/04/11/floating-point-numbers-an-infinite-number-of-mathematicians-enter-a-bar/ https://dennisforbes.ca/index.php/2017/04/11/floating-point-...
- maurits 9y agoObligatory: What Every Computer Scientist Should Know About Floating-Point Arithmetic [1] (pdf) [1]: http://www.itu.dk/~sestoft/bachelor/IEEE754_article.pdf http://www.itu.dk/~sestoft/bachelor/IEEE754_article.pdf
- jordigh 9y agoOkay, fine, I agree that sometimes mathematical notation is bad and we are all computer people here, not math people, so we get really scared of mathematical notation. But is (-1)^S 1.M 2^(E-127) so bad that it required a whole blog post to explain it? Except for the "1.M" pseudo-notation to explain the mantissa with the implicit on bit, all of those symbols are found in most programming languages we use. I don't think the value of the blog post was explaining the notation. We all knew what operations to perform when we saw it. The value seems to lie more in thinking of the exponent as the offset on the real line and the mantissa as a certain window inside that offset. Personally, though, this still doesn't seem like a huge, deep insight to me, but maybe I'm just way too used to floating point and have forgotten how hard it was to learn this. I did learn about mantissa, exponents, and even learned how to use a log table in high school, but maybe I'm just old and had an unusual high school experience.
- kemerover 9y agoI don't really understand significance of this insight easier. Floating point number is just stored in scientific notation base 2. That's it. And kids learn scientific notation in 7th grade? 8th tops. I mean, swap base 2 to base 10 in this image, and the effect changes from "woah" to "duh, obviously".
- coldtea 9y ago>But is (-1)^S 1.M 2^(E-127) so bad that it required a whole blog post to explain it? Well, for one the expression doesn't tell us anything, could as well describe alien gravity -- even if we know math. You still need to explain what S, M, and E are used for to understand it.
- tpeo 9y agoHow is this issue unique to strings? Nobody can look at a random graph and know what it stands for either. Otherwise there would be no need for labeled axes. You still need to explain what is it that we're looking at This isn't to say that visualizations aren't useful, of course.
- javajosh 9y agoWhy did they fix the bit-width for the mantissa and exponent? It would be nice to have more bits for the mantissa when you are near 1, and then ignore the mantissa entirely when you're dealing with enormous exponents, and very far from one. Granted, there would be some overhead (e.g. a 3-bit field describing the exponent length, or something) but it would be a useful data-structure.
- coldtea 9y agoThat exists, it's called arbitrary precision floating point (as opposed to single/double precision etc).
- dnautics 9y agoYou may like: https://youtu.be/aP0Y1uAA-2Y https://youtu.be/aP0Y1uAA-2Y It turns out the math is slightly harder but it's faster than IEEE FP in hardware, probably because of less conditionals in the spec. There is better performance in terms of numerical accuracy at the cost of substantially more difficult error analysis. Disclaimer: I'm in the video.
- dsego 9y ago> Instead of Exponent, think of a Window between two consecutive power of two integers. I know what an exponent is, or if you want "order of magnitude". Sorry, but "A window between two consecutive power of two integers" doesn't make it easier to think about.
- conistonwater 9y agoI really wish he didn't make [0,1] one of the windows, because in floating point arithmetic the range [0,1] contains approximately as many floating point numbers (a billion or so in Float32) as the range [1,∞). There are "windows" [2^k,2^(k+1)] for positive as well as negative k. Just creates unnecessary scope for further confusion.
- azernik 9y agoEDIT: Post on wrong article. > But brain scientists say young people often lack the perspective and judgment, especially in the moment, to know what’s in their best interest. And yet, the actual changes seem to mostly have to do with language skills, comprehension, and familiarity with the US legal system. Compare: Original Miranda warning: "You have the right to consult an attorney before speaking to the police and to have an attorney present during questioning now or in the future. If you cannot afford an attorney, one will be appointed for you before any questioning if you wish." Revised "for kids" Miranda warning: "You have the right to talk to a free lawyer right now. That free lawyer works for you and is available at any time – even late at night. That lawyer does not tell anyone what you tell them." Given that the original Miranda case had to do with immigrants, I think the revised version is actually much more fitting to the purpose than the current standard.
- coldtea 9y agoI see how some people just get the math, but I don't see why programmers here say they find it difficult to understand the window / offset explanation the article gives. A "window" is a common programming term for a range between two values. An "offset" is a common term for where a value falls after a starting point. In simpler decimal and equidistant terms, the idea is to split a range of values in windows, divide each window in N values, and store an FP number by storing which window and which index inside the window (0 to N) it falls. The FP scheme actually uses powers of 2 instead of equal distant windows (so the granularity becomes coarser as the numbers become bigger) but the principle is the same.
- Michielvv 9y agoI'm guessing because if you put those words together, you automatically assume the offset means the position of the window. It's clear enough with the illustration, but using terms that are in very common use for other concepts can make it confusing at the first glance. Also the window explanation hides the fact that it's able to represent really small numbers. (graph implies that 0-1 is the smallest window)
- squeaky-clean 9y agoSmaller is an ambiguous term when it comes to numbers. Is it the lesser value? Or is it the value nearest to 0? It depends on the context. Either way, they cover the sign bit, and the brackets on the diagram show the graphs are not using the sign, just exponent and mantissa. So I'd assume the same principles stand but for 0 to -1, -1- to -2, and so on.
- Michielvv 9y agoI mean smaller as in closer to zero as in exponent lower than 127. That part is obvious in the formula, but not from the graphs. The interesting part of the explanation is that it shows very clearly the effect of the 1.M in the formula. Unfortunately it then does not answer the resulting question: how is zero represented. ( https://en.wikipedia.org/wiki/Signed_zero https://en.wikipedia.org/wiki/Signed_zero does a pretty good job at that though)
- agumonkey 9y agoThe last bits of trivia are very nice. The x87 coprocessor makes me wonder about days were each chip changed you system. It was such a different mindset that videogame consoles had parallel routes between the board and the cartridge themselves to allow hardware extension per game.
- Veedrac 9y agoLet's represent the number 42,643,192, or 10100010101010111011111000₂, in different "floating point" representations. Scientific notation with 5 significant figures: 4.2643 × 10⁷ Scientific notation in base 2 with 17 significant binary figures: 1.0100010101010111₂ × 2²⁵ Let's pack this in a fixed-length datatype. Note that 011001₂ is the binary encoding of 25. 1 0100010101010111 011001 1 mantissa exp. This doesn't suffice because a. We're wasting a bit on the leading 1. b. We want to support negative values. c. We want to support negative exponents. d. It would be nice if values of the same sign sorted by their representation. The leading 1 can be dropped and replaced with a sign bit (0 for "+", 1 for "-"). The exponent can have 100000₂ subtracted from it, so 011001₂ represents 25-32, or -7, and 111001₂ represents 25. Sorting can be handled by putting the exponent before the mantissa. Thus we get to a traditional floating point representation. 0 111001 0100010101010111 ± exp. mantissa Real floating point has a little more on top (infinities, standardised field sizes, etc.) but is fundamentally the same.
- kibwen 9y ago> Sorting can be handled by putting the exponent before the mantissa. Don't denormal numbers prevent simply sorting this way? I could use a diagram like the OP's to remind myself how they work...
- thethirdone 9y ago> Don't denormal numbers prevent simply sorting this way? I could use a diagram like the OP's to remind myself how they work... No, there is nothing really special about subnormal numbers other than that they have a smaller than normal precision. For example 0x00000001 < 0x00400000 in both floating point and normal integer arithmetic. The only weird bit about sorting floats as integers is that the order is reversed; eg -0f is represented by the smallest integer (0x80000000) and -Inf is represented by 0xFF800000.
- cesarb 9y agoUsing the article's "window" analogy: denormals are in the smallest "window", so the sort order between the "windows" (exponents) is kept. Their special property is that they don't have the implicit one to the left of the point (it's an implicit zero instead); the order within the mantissa is still kept. That is: 0 000000 0100010101010111 ± exp. mantissa Here, the value is 0.0100010101010111₂ × 2^-30 (I hope I calculated the exponent correctly) As a bonus, the zero value comes naturally in this approach: it's a denormal with a mantissa of zero. Without denormals, the implicit one would get in the way.
- makmanalp 9y ago> Since floating point units were so slow, why did the C language end up with float and double types ? After all, the machine used to invent the language (PDP-11) did not have a floating point unit! The manufacturer (DEC) had promised to Dennis Ritchie and Ken Thompson the next model would have one. Being astronomy enthusiasts they decided to add those two types to their language. Wait, what was the alternative? No floats? How the heck would people calculate things with only integers? edit: AFAIK bignums are even slower, and fixed-point accumulates error like crazy
- coldtea 9y ago>How the heck would people calculate things with only integers? Easily -- it happens all the time in industries which demand specific precision -- e.g. integers are used to calculate monetary values in languages that don't have a decimal/big number type. You just need to format them to the precision you want when you show them to the user. It would just be slow.
- makmanalp 9y agoSorry, by bignum I didn't mean "big numbers", I meant arbitrary-precision, i.e. 1/3 instead of 0.333333333, is this what you mean? I don't see how you get around stuff like many divisions or nth roots. You'd lose precision during the calculation operation, right? Whether you format it at the end has little bearing, since it'll be garbage by that time.
- pjc50 9y agoYou can just as easily lose precision in FP arithmetic if you're not careful. But for engineering purposes fixed point works quite well because having arithmetic much more precise than manufacturing tolerance is no use. Fixed point was good enough to go to the moon: https://www.netjeff.com/humor/item.cgi?file=ApolloComputer https://www.netjeff.com/humor/item.cgi?file=ApolloComputer
- makmanalp 9y ago
- triangleman 9y ago>I wanted to vividly demonstrate how much of a handicap it was to work without floating points. So, did he manage to demonstrate that in the book? Because the page linked here, while explaining how floating points are represented in memory, does not explain how computers perform operations on them, or what purpose does a FPU serve (how does it differ from an ALU).
- jokoon 9y agoImagine a ruler with all floating point values on it, each time the mantissa comes at its maximum, you increase the exponent, so the space between farther float values doubles. The number of mantissa values being constant for each exponent value, the exponent describes some kind of "zoom level". Float values on a ruler would sort of looks like this: ... x x x x x x x x x x... ^ exponent increases, spacings are doubled
- oxide 9y agoAs a complete layman with only a cursory knowledge of programming, as well as a complete lack of math skills above Algebra 2, (I didn't even complete that, tbh, once they threw graphing into the equation. I did get slope-intercept form down, but that's it.) I ended up finding this easier to understand than I expected, and a great read. I love explanations like these, with a visual breakdown. It really helps it "click." as long as I glazed over the math formulas and didn't let the numbers overwhelm me. This is what I took away: the exponent "reaches" out to the max value of the [0,1] [2,4] etc, and the number represented tends to be like 51-53% of the way down the line of the mantissa. It "clicked" a bit for me, see? Am I way off? This is the way I always learned math the best in school, an alternate explanation that helps it "click." Very good explanation, from my point of view, of how floating point numbers work and what they even are. That's a nice feeling for someone like me who is pretty bad at math and finds formulas like the one shown in the article to be, frankly, indecipherable. But now I (sort of) understand how floating point numbers work, (sort of) what they are, why they are important, and what role they play. Could I program anything using one? No. But, I could learn someday, and explanations like these give me some hope that I just might be able to learn a programming language if I put the effort in. That I could learn the math required of me, even!
- mdip 9y ago> People who really wanted an hardware floating point unit in 1991 could buy one. The only people who could possibly want one back then would have been scientists (as per Intel understanding of the market). They were marketed as "Math CoProcessor". Performance were average and price was outrageous (200 USD in 1993 equivalent to 350 USD in 2016.). As a result, sales were mediocre. Actually, that's only partly true. My father owned a company that outfitted large manufacturing shops (MI company, you can imagine who his customers were). As a result, he used AutoCAD. The version of AutoCAD he used had a hard requirement on the so-called "Math Co-processor", so he ended up having to purchase one and install it himself. That was my first taste of taking a computer apart and upgrading it and I credit that small move with my becoming interested in building PCs, which led to my dad and I starting a business in the 90s doing that for individuals and businesses. There were definitely more reasons for that kind of add-on than just scientific fields; anyone in the computer aided drafting world at that time needed one as well.
- snaky 9y agoThat's why there was math coprocessor software emulator. It worked fine on 386SX.
- BinaryBullet 9y agoSee also: An interactive floating point visualization: https://evanw.github.io/float-toy/ https://evanw.github.io/float-toy/
- userbinator 9y agoI think the example values at https://en.wikipedia.org/wiki/Minifloat https://en.wikipedia.org/wiki/Minifloat are most useful for intuitively understanding how floating point works --- especially the "all values" table, which shows how the numbers are spaced by 1s, then 2s, then 4s, etc. meaning the same number of values can represent a larger range of magnitudes, but sacrificing precision in the process.
- cesarb 9y agoIMO, the best way to explain floating point is to play with a tiny float. With an 8-bit float (1 bit sign, 4 bits exponent, 3 bits mantissa, exponent bias 7), there are only 256 possible values. One can write by hand a table with the corresponding value for each of the 256 possibilities, and get a feel to how it really works. (I got the 1+4+3 from http://www.toves.org/books/float/ http://www.toves.org/books/float/, I don't know if it's the best allocation for the bits; but for didactic purposes, it works.)
- deleted 9y ago[deleted]
- piyush_soni 9y agoDoes anyone have an alternate link, for some strange reason this link appears to be blocked at my work.