6 ms·
I'm pretty sure most of us DS know about significant digits and are usually calculating the maximum to enable future flexibility. For a single output, I can und
by lifeslogit 3y ago
I'm pretty sure most of us DS know about significant digits and are usually calculating the maximum to enable future flexibility. For a single output, I can understand how you'd be upset we don't round. But for the 20+ column tables we usually build, I've found most will calculate the maximum scale allowed by the database and call it a day. The best of us certainly find the right formatting for the presentation.
- cauch 3y agoI agree with that. It reminds me of the tomato. Some people don't know the tomato is a fruit. Some people know the tomato is a fruit and treat them like a fruit. Some people know the tomato is a fruit and still treat them like a vegetable because it is what makes sense. In practice, I have never saw someone mistreating a badly written number in a way that had any impact. If they don't know themselves the concept of significant digit and believe the number is 1.234567890 precisely and not 1.234567889, what would be the wrong decision they will take that they would not have taken if the number was 1.234567889? It starts to matter when you have 10% or ~100% uncertainty, in which case, writing 1.2 or 1 is still not enough to convey the meaning to someone who does not get significant numbers, because for them, 1 = 1.0000 anyway. In this case, you need to explicitly explain the limitation on decision due to the uncertainty. In practice, splitting hairs on the significant digit convention is just missing the point: if you apply the convention to people who are not informed on the precision, they will make bad decision anyway even if technically the number has the correct significant number.
- nicklecompte 3y agoThe problem is that it's not just "presentation": being sloppy about significant digits (or precision generally) early in the computation leads to bad statistical reasoning much later in the problem. If your variable is x +/- 0.05, then 1/(x +/- 0.05) != 1/x +/- 0.05. If you're not careful about this when doing computations, you'll end up with answers that aren't actually meaningful. The computational implementation of these equations is only concerned with machine epsilon, but each one of those 20 database columns has a real-world +/- delta which isn't being correctly considered.
- cauch 3y agoBut the error propagation is not transmitted by the significant number. x, y being written with the correct number of significant number will not lead to f(x, y) being written to the correct number of significant number. Usually, the best approach is to propagate the uncertainty, for example by saving the uncertainty as another variable in the database and using it directly when the number is used. If you do that, there is no practical needs to lose time to format the numbers. Using significant numbers seems a "cheap trick" that risk to mislead you more often than help.
- nicklecompte 3y ago> But the error propagation is not transmitted by the significant number. x, y being written with the correct number of significant number will not lead to f(x, y) being written to the correct number of significant number. Significant figures are not a convention for making your deliverable pretty. They have semantic meaning. Don't think about dumb rules from high school chemistry, think about the actual problem. There are two entwined sources of uncertainty I am referring to: 1) measurement uncertainty, due to a lack of precision in the instrument (or the quantity itself, e.g. many financial computations are not meaningful if they involve fractional cents) 2) computational uncertainty, which is exclusively due to algebraic propagation of measurement uncertainty Far too many data scientists don't care about the first category of uncertainty because they don't care about where the data came from. And they don't even realize the second category is a problem. Let's look at a specific example. Somebody tells us that they measured the side of a square as 1.0m. Their tape measure only went down to centimeters, so the uncertainty is +/- 0.01m. What is the area of this square? Let's look at it two ways: 1) The smallest possible side length is 1.00m - 0.01m = 0.99m, so the smallest possible area is 0.98m^2. The largest possible side length is 1.01m, so the largest possible area is 1.02m^2. Thus the area is 1.00m^2 +/- 0.02m^2. 2) The side length is (1.00m +/- 0.01m). So the area is (1.00m +/- 0.01m)(1.00m +/- 0.01m) = 1.00m^2 +/- 0.02m^2 +/- 0.0001m^2 ~ 1.00m^2 +/- 0.02m^2 So the uncertainty is not +/- 0.01, it is +/- 0.02. This can add up quite dramatically. In general if you have x +/- delta, then f(x +/- delta) is not going to be f(x) +/- delta or f(x) +/- f(delta). It needs to be handled carefully.