4 ms·
The code in the statistics library is actually quite a bit more complicated than I had expected. Is there a reason why a function like `mean` isnt just: `sum(da
by eachro 5y ago
The code in the statistics library is actually quite a bit more complicated than I had expected. Is there a reason why a function like `mean` isnt just: `sum(data) / len(data)`?
- mattkrause 5y agoThis is the complete code (from here: https://github.com/python/cpython/blob/3.9/Lib/statistics.py https://github.com/python/cpython/blob/3.9/Lib/statistics.py) if iter(data) is data: data = list(data) n = len(data) if n < 1: raise StatisticsError('mean requires at least one data point') T, total, count = _sum(data) assert count == n return _convert(total / n, T) The _sum function is a little more involved, but not appreciably so. In general though, a literal translation of a formula is not always great in terms of numerical stability/error. For example, you might have learned to calculate the variance as mean(X^2) - mean(X)^2, but this can lead to a huge loss of precision that more complicated approaches avoid.
- abecedarius 5y agoA couple things here I don't understand: > if iter(data) is data: Wouldn't it be cheaper like `if type(data) is iter:`? And why convert `data` to a list at all, to check for length 0, given that `_sum(data)` will return the count?
- turndown 5y agoPerhaps it's cheaper to just get the conversion and length check out of the way rather than have to execute _sum and find out from checking T.
- abecedarius 5y agoBut the path where the cost of _sum is extra is the error path; plus that'll end up being 0 iterations anyway.
- slaymaker1907 5y agoOne trick that often improves numerical stability with computing the sum is to first sort the data.