13 ms·
PEP 450: Adding A Statistics Module To The Standard Library
- aristus 13y agoAbout damned time. Writing your own stats library is like writing your own crypto.
- aidos 13y agoYou wouldn't write your own - numpy / scipy have everything you'll need.
- dalke 13y agoSome of the things I need are: - fewer dependencies for my package I've written the average() and standard_deviation() functions at least a couple of dozen times, because it doesn't make sense to require numpy in order to summarize, say, benchmark timing results. - reduced import time NumPy and SciPy were designed with math-heavy users in mind, who start Python once and either work in the REPL for hours or run non-trivial programs. It was not designed for light-weight use in command-line scripts. "import scipy.stats" takes 0.25 second on my laptop. In part because it brings in 439 new modules to sys.modules. That's crazy-mad for someone who just wants to compute, say, a Student's t-test, when the implementation of that test is only a few dozen lines long. (Partially because it depends on a stddev() as well.) Sure, 0.25 seconds isn't all that long, but that's also on a fast local disk. In one networked filesystem I worked with (Lustre), the stat calls were so slow that just starting python took over a second. We fixed that by switching to zip import of the Python standard library and deferring imports unless they were needed, but there's no simple solution like that for SciPy. - less confusing docstring/help Suppose you read in the documentation that scipy.stats.t implements the Student's t-test as scipy.stats.t. >>> import scipy.stats >>> scipy.stats.t <scipy.stats.distributions.t_gen object at 0x108f87390> It's a bit confusing to see scipy.stats.distributions.t_gen appear, but okay, it's some implementation thing. Then you do help(scipy.stats.t) and see Help on t_gen in module scipy.stats.distributions object: class t_gen(rv_continuous) | A Student's T continuous random variable. | | %(before_notes)s | ... | | %(example)s Huh?! What's %(before nodes)s and %(example)s? The answer is, scipy.stats auto-generates various of the distribution functions, including things like docstrings. Only, help() gets confused about that because help() uses the class docstring while SciPy modifies the generator instance's docstring. Instead, to see the correct docstring you have to do it directly: >>> print scipy.stats.t.__doc__ A Student's T continuous random variable. Continuous random variables are defined from a standard form and may require some shape parameters to complete its specification. Any optional keyword parameters can be passed to the methods of the RV object as given below:
- cdavid 13y agoscipy.stats distribution objects are a bit particular, that's a bit unfair to pin point them. Generally, numpy and scipy have much better docstrings than python stdlib itself.
- dalke 13y agoWell, help(scipy.optimize.nonlin.Anderson) has the same problem, but you're right in that that failure mode is rare, and that numpy/scipy has good documentation. However, in the context of a stats library, I think it's okay to point out that scipy.stats has some annoying parts. ;) In all honesty, I seldom use NumPy and rarely use SciPy, so I can't judge that deeply. I know that when I read their respective code bases I get a bit bewildered by the many "import *" and other oddities. It doesn't feel right to me. I know the reason for most of the choices - to reduce API hierarchy and simplify usability for their expected end-users - but their expectations don't match mine. So I looked at more of the documentation. I started with scipy/integrate/quadpack.py. The docstring for quad() says, in essence, "this docstring isn't long enough, so call quad_explain() to get more documentation." I've never seen that technique used before. The Python documentation says "see this URL" for those cases. Again, this is a difference in expectations. I argue that NumPy and Python have different end-users in mind. Which is entirely reasonable - they do! But it means that it's very difficult to simply say "add numpy to part of the standard library." There's also a level of normalization that I would want should numpy be part of the standard library. For example, do out of range input raise ValueError or RuntimeError? scipy/ndimage/filters.py does both, and I don't understand the distinction between one or the other. Now, in the larger sense, I know the history. RuntimeError was more common in Python, and used as a catch-all exception type. Its existence in numpy reflects its long heritage. It's hard to change that exception type because programs might depend on it. But it means that integrating all of numpy into the standard library is not going to work: either it breaks existing numpy-based programs, or the merge inherits a large number of oddities that most Python programmers will not be comfortable with.
- cdavid 13y agoActually, I don't think the import * in numpy is anything else than historical artefact. Numpy just happens to be one of the oldest, still widely used python library (considering numpy started as numeric), as you point out. As for import speed, have you considered using lazy import in your script ? I don't see numpy being integrated in python anytime soon. I don't think it would bring much, and one would have to drop performance enhancement that rely on blas/lapack. I think installing has improved a lot, and once pip + wheel matures, it should be easy to pip install numpy on windows.
- coldtea 13y agoWell, one of the things I need is for it to be built in the standard library, so not, numby/scipy doesn't have everything I need.
- tvst 13y agoNo mention of Pandas? http://pandas.pydata.org/ http://pandas.pydata.org/
- yati 13y agoI think Pandas is a great candidate for inclusion in the stdlib if this ever happens - and hopefully, numpy/parts of scipy will also be thrown in :)
- lifeisstillgood 13y agoI think the idea is to include a small independent stats package not a full featured still developing third party. Any number of people need std dev easily and reliably available. If you need numpy on top of that, you know you do and can afford the effort. For 99% of my work numpy and the associated compilation overhead is unneeded - fits my brain, fits my needs
- westurner 13y agoSo let's amortize the cost of compiling and/or installing fast binaries by only relying on plain Python. It would be great if there was a natural progression (and/or compat shims) for porting from this new stdlib library to NumPy[Py] (and/or from LibreOffice). (e.g. "Is it called 'cummean'")?
- lifeisstillgood 13y agoI guess that's the point of the stats-battery - pure python stats with no / minimal cost to migrate to numpy e.g. From stats import mean ... from numpy import mean
- goronbjorn 13y agoThey probably left out pandas because it depends on numpy and also this point: > For many people, installing numpy may be difficult or impossible. that's as true, and arguably more, for pandas.
- bachback 13y agoNice proposal. I think the problem is numpy itself. If you could just do pip install numeric_package then nobody can complain. I don't quite understand why a package has to depend on LINPACK. I will probably switch to julia-lang, because numpy is (at least for me) not that great to work with.
- stiff 13y agoNumPy is a full MAT-LAB for Python, not a simple drop-in statistics library. It has to depend on LINPACK because writing a full linear algebra library that performs well is damn hard and takes several researcher-years. Most serious scientific computing libraries and utilities depend on it, including Julia I think. There is certainly room for simpler libraries for people not seriously into numeric computations, as the document linked well indicates.
- mvanveen 13y agonumpy has all sorts of awful C bindings which make it less than versatile in environments where you want pure Python. It's great from a performance point of view, but horrible for compatibility. Google App Engine used to suffer because of this (more specifically, it still only restricts your runtime to pure Python, but now you can import numpy at least). I believe the PyPy folks have also had their own set of struggles with numpy compatibility, although I'm not sure what the state of that is at present. In any case, I think these compatibility concerns alone make a strong argument for including simple Statistics tooling into the standard library.
- asgard1024 13y agoHow does that work for SQLite 3, which _is_ part of Python library? I would actually prefer to have numpy included before those statistics functions.
- mvanveen 13y agoOK, that's a pretty good counterexample. Touché. :-p GAE just avoids it alltogether (except locally, where you have CPython and use it to stub out core services hosted on the cloud runtime). You simply can't import sqlite3 on GAE when running on cloud runtime, nor can you really use it as an external dependency. I'm not really up on the details, but the PyPy website claims they've gotten around this by implementing a pure Python equivalent of the CPython stdlib library (http://pypy.org/compat.html http://pypy.org/compat.html). I would put forward that SQLite3 is probably a pretty easy include in most C projects compared to whatever numpy would likely require. That said, I'm not qualified to assess this, being neither a numpy, Python core, or sqlite3 dev. All of this aside, it's worth mentioning that the entire standard lib includes and depends on some other C-only libraries. So it's not unprecedented. In principle, you'd want the standard lib to have as much pure Python as possible (PyPy kind of takes this to the ultimate extreme from what I can gather), but this isn't always practical (great example of "practicality beats purity" if you ask me). Speaking of which, if it's cool to have `sqlite3` in the standard lib as part of the included batteries, why not mean and variance and the like? :D
- matiasb 13y agoNice idea
- rev 13y agoKudos to PHP for apparently being ahead of the curve among dynamic languages with regard to statistics. Another interesting, yet unmentioned option is Clojure/Incanter.
- draegtun 13y agoAnd another option is PDL - http://pdl.perl.org http://pdl.perl.org NB. And I believe Perl6 (spec) includes PDL - http://perlcabal.org/syn/S09.html#PDL_support http://perlcabal.org/syn/S09.html#PDL_support
- scribu 13y agoI upvoted this comment, then I realized that those PHP stats functions aren't in the standard library. They're in a PECL extension (equivalent to Python C extensions).
- zokier 13y agoReminds me of the story that made rounds here couple of years ago: The Python Standard Library - Where Modules Go To Die https://news.ycombinator.com/item?id=3913182 https://news.ycombinator.com/item?id=3913182
- fiatmoney 13y agoIt's not a terrible idea to support the absolute basics like mean & variance, but anything beyond that (particularly things like models or tests) is not a good idea for a standard library. Once you hit even something simple like a linear regression you have issues of how to represent missing or discrete variables, handling colinearity, or whether to do online or batch modes which can give different results. Tests in particular are fraught because if you're going to make them available for general consumption they need a good explanation of when they're appropriate, which is basically a semester course in statistics and well out of scope for standard library docs. Basically, the idea of "batteries included" should also mean that if something looks like you can put a D-cell in there, you're unlikely to blow your arm off.
- daniel-levin 13y agoAgreed. I'm studying statistics at the moment and I'm continually reminded of how easy it is to choose the wrong model / distribution and be incorrect because of some non-obvious and technical reason. For example, just the other day, I wanted to use the binomial distribution to solve a problem. To use this distribution, the trials must be independent of one another. In that particular problem, there was a subtle condition that made the trials non-independent. I arrived at correct-appearing answers (0 <= P <= 1) that were actually all wrong. Statistics is way too easy to break to be used naively.
- lutusp 13y ago> Statistics is way too easy to break to be used naively. Fair enough, but the same argument could be made about using an unskewed standard distribution on non-symmetrical datasets, a common error even among people who should know better. I think binomial functions should be included, on the ground that they're very useful and their probability of misuse is only equal to the continuous statistical forms, not more so.
- dandellion 13y agoHell, sometimes they use a dictionary when they should be using a list. Almost everything can be used wrongly by a begginer, which doesn't mean it shouldn't be there. I think having a basic stats module always handy would be very convenient.
- cabalamat 13y ago> For many people, installing numpy may be difficult or impossible. For example, people in corporate environments may have to go through a difficult, time-consuming process before being permitted to install third-party software. I do not regard this as a good justification for putting something in the standard library! If you don't have root access, use vitualenv (which you might want to do anyway) and install the package somewhere under your home directory.
- nknighthb 13y agoHaving the technical capability to do something is not the same as having permission to do it. The PEP refers to corporate policies. Installing third-party software in your home directory would be just as much a violation as doing it in a location that requires root.
- siddboots 13y agoNumPy is not nearly portable enough to do what you are describing as a user on a windows machine. You cannot simply do a `pip install numpy` into a virtualenv. Instead you must either install the package system-wide, or compile it yourself, which means getting a working MinGW environment or similar. Edit: Although, I do agree that NumPy being difficult to install is not, on its own, a good justification for the PEP.
- cdavid 13y agonumpy is quite portable, I am not sure what you mean by not nearly portable enough. The reason why you can't do pip install numpy is pip's fault, there is nothing that numpy can do to make that work. Note that easy_install numpy does work on windows (without the need for a C compiler).
- bachback 13y agothe problems I have had with numpy are endless. Usually I'll just prefer to write my own, because it's quicker. if you have tried to get numpy running on a cloud machine you'll know what I'm talking about. basically you will have to know how to compile from source, know some gcc, etc. the last time I tried to get it running I promised myself never to use numpy again.
- lutusp 13y agoGreat idea, but while assembling this library, don't leave out permutations, combinations, and the binomial Probability Mass Function (PMF) and Cumulative Distribution Function (CDF). Small overhead, easy to implement, very useful. More here: http://arachnoid.com/binomial_probability http://arachnoid.com/binomial_probability
- enalicho 13y agoPermutations and combinations already exist within the itertools module.
- lutusp 13y ago> Permutations and combinations already exist within the itertools module. Not exactly. Given argument lists, Itertools provides result lists (actually, iterators for that purpose) with the original elements permuted and combined, but doesn't provide numerical results for numerical arguments, as shown here: http://arachnoid.com/binomial_probability http://arachnoid.com/binomial_probability I was referring to permutation and combination mathematical functions, not generator functions.
- megrimlock 13y agoPut differently, you want the functions that count the number of permutations and combinations possible, not functions that yield/generate the actual permutations and combinations.
- lutusp 13y agoYes, exactly, for a number of reasons including the problem of large arguments and results.
- blt 13y agoHmm. Python has default Bignum promotion so the naive N choose K implementation will not suffer from overflow. I guess it could be slow for large values, but how many casual users need to calculate N choose K for large N?
- clutchski 13y agoBatteries included is a fine philosophy when starting a language to encourage early adoption, but at this point, I don't think it's worth adding new libraries to the stdlib. Here's why: - It's very easy to find and install third party modules - Once a library is added to stdlib, the API is essentially frozen. This means we can end up stuck with less than ideal APIs (shutil/os, urllib2/urrlib, etc) or Guido & co are stuck in a time consuming PEP/deprecate/delete loop for even minor API improvements. - libraries outside of the stdlib are free to evolve. users of those libraries who don't want to stay on the bleeding edge are free to stay on old versions.
- dalke 13y agoThose are good reasons for rejecting any addition to the standard library. However, new libraries are sometimes added to the standard library, which means the reasons you listed can be overcome by even better reasons for inclusion. What are those reasons for why a new library can be included, and why aren't those reasons appropriate justification for including this proposed statistics package?
- Peaker 13y ago> However, new libraries are sometimes added to the standard library, which means the reasons you listed can be overcome by even better reasons for inclusion. Or overcome by bad judgement.
- __e 13y agoThe PEP acknowledges the existence of high-end statistics libraries. It also notes that the alternative to such libraries are DIY implementations - which are often incorrect in their implementation. The PEP proposes adding simple, but correct support for statistics. Apart from high-end libraries being an overkill and DIY implementations being incorrect, the PEP also cites resistance to third party software in corporate environments. This problem is more social than technical though, and I'm not sure what weight must be attached to it
- pmr_ 13y agoThere is a different perspective on the frozen APIs. Using an API from the stdlib gives you the certainty that your program is not going to break with a minor python version bump. This might not matter for all software but is crucial for others.
- ot 13y agoJust out of curiosity, I submitted this yesterday: https://news.ycombinator.com/item?id=6190603 https://news.ycombinator.com/item?id=6190603 The URL was http://www.python.org/dev/peps/pep-0450/ While this is http://www.python.org/dev/peps/pep-0450 That is, exactly the same except for a trailing slash. Doesn't the deduplication algorithm handle this case?
- daGrevis 13y agoTechnically speaking, they are separate URLs that may lead to separate resources. For example, Google engine treats them as separate URLs. That's the reason why opening http://www.python.org/dev/peps/pep-0450 http://www.python.org/dev/peps/pep-0450 redirects to http://www.python.org/dev/peps/pep-0450/ http://www.python.org/dev/peps/pep-0450/ . HN engine should follow redirect to avoid situations like this.
- ot 13y agoTechnically speaking, there are no equivalent URLs in general, different strings may lead to different resources. Still, there are a number of common sense heuristics to normalize URLs, that HN applies to do de-duplication. I was wondering what is the rationale for not having trailing slash removal among them. I mean, is there any legitimate website that serves a different resource if you remove the trailing slash?
- keeperofdakeys 13y agoWithout actually checking redirects, not breaking a few edge cases is much better than a few submissions being duplicated.
- zeckalpha 13y agoOr it could check for a 3xx HTTP status.
- gsnedders 13y agoPer RFC3986/7, http://example.com/%60 http://example.com/%60 and http://example.com/a http://example.com/a are equivilant. (Indeed, all major browsers will request the latter regardless of what is input.) Equally, punycode encoded IRIs and the original IRI are equivilance. There is a whole section on equivilance in both of the RFCs (3967 includes 3986 by reference, so is a superset).
- bayesianhorse 13y agoI'm against this. Either you have to create a new statistics module or you would have to include numpy/pandas/statsmodels into the standard library. In both cases it would essentially freeze the modules for further development outside the python release cycle...
- andrewflnr 13y agoI'm in favor. I was surprised and annoyed to find there wasn't a standard library for doing excel-level statistics. If you throw basic least-squares linear regression in there too, I can eliminate Excel from my physics classes.
- cindaydavilla 13y agomy classmate's ex-wife makes $63 an hour on the laptop. She has been without work for seven months but last month her check was $18401 just working on the laptop for a few hours. Read more here... max38.cℴm
- bthomas 13y agoOne side effect is that this would accelerate the adoption of Python 3 in the scientific community
- Demiurge 13y agoI would like having these simple functions, but I think they can just go into 'math' library.