4 ms·
Passing the torch of NumPy and moving on to Blaze
- carterschonwald 14y agoCongrats to Travis and the rest of the Continuum analytics team on the Darpa XDATA funding! As someone working to build tools in the same space as Continuum (and perhaps as a competitor), having your competitors (Continuum) be intelligent, nice, interesting folks who really understand the problem domain is pretty darn great. Point being: the numerical computing / data analysis landscape is going to be seing a lot of great tools emerge and/or mature over the next year, and I have no doubt that 30-50% of them will be coming from Continuum Analytics. [edit: to the substantial enrichment of high level tools for extending numerical within python / and likely generally!] I can only hope that I execute my tool building work at WellPosed well enough that I can call them a competitor for years to come!
- xaa 14y agoAs someone who has dabbled in using Python for numerical computing in several small projects, I wonder: what would be the motivation for further investment in Python as a numerical platform, considering all Python's problems with concurrency. Real threads will never come to Python. MPI is a real pain unless you are running very large computations. This will only become more true as time progresses. Am I missing something?
- dagw 14y agoIs concurrency really that important for numeric work? Surely parallelism is what you care about. Many numpy primitives are already parallel since it basically just hands off to your BLAS library. Beyond that there is numexpr which is really good at doing parallel evaluation of large array expressions. If your problem isn't solved by any of these, there are other powerful solutions like IPython, Parallel Python and even multiprocessing from the standard library If you need even more performance, cython has some support for semi-automated parallelization, and if all else fails drop down to C and use OpenMP or whatever else you like. So while concurrency is a problem in python, numeric parallelism is an area where many good solutions exist.
- hippyloopy 14y agoWhy should I have to revert to a C library in order to do anything in Python? What's the point of using Python if every time I want to do something in parallel I'm going to have to write a C library? Python people have their heads in the sand! If we have a hundred core processors, running Python on a single core is not going to be a tractable solution to any problem. Your BLAS may be parallel, but any time you go back into the Python driver code suddenly it's a massive bottleneck. Distributed Python is a messy hack that wastes all the amazingly tuned shared memory support in the processor. Writing C extensions goes against the whole point of using Python. "if all else fails" The problem with Python is that the moment you want to do something in parallel, which in the next decade will be everyone, "all else fails" is your starting point!
- sqrt17 14y agoThe model that has worked amazingly well for Python (and Matlab, and R, and probably a number of others) is to encapsulate the hard stuff - say BLAS for linear algebra, or GraphLab for loopy belief propagation - together with all the amazingly tuned shared memory support, concurrency, parallelism, data locality, whatnot - in C-level modules written by expert people and expose a powerful API that doesn't expose you to the nontrivialities of concurrent or parallel programming. If you spend lots of time in the driver code, you most certainly won't be happy with Python, R, or Matlab, but then Cython (and possibly Numba at some point) help push this "lots of time" further and further down. "all else fails" is the starting point of pretty much everyone doing real work. What's your alternative here? Most of the time, specialized libraries will be both more convenient and more efficient than rolling your own with fine-grained concurrency/parallelism in Java or PyPy.
- StefanKarpinski 14y agoThe paper "Evaluating the Design of the R Language" [1] is a great read on this subject. A key figure they found (p. 17 end of top paragraph) is that in a realistic corpus of work, only 22% of compute time was spent in C/Fortran "kernels" as opposed to R code. So the effectiveness of the "two language" design is somewhat limited, even for scientific workloads where kernels like BLAS, FFTs, etc. apply (and there are many areas where they don't really). [1] http://r.cs.purdue.edu/pub/ecoop12.pdf http://r.cs.purdue.edu/pub/ecoop12.pdf
- deleted 14y ago[deleted]