4 ms·
The problem with the Python ecosystem is that there exist a lot of libraries for scientific computation mostly incompatible with each other (blaze, numpy, numba
by pathsjs 11y ago
The problem with the Python ecosystem is that there exist a lot of libraries for scientific computation mostly incompatible with each other (blaze, numpy, numba, numexpr, dask...). Some of them try to reintroduce types in order to compile to native code, thereby losing the advantages of dynamic typing. In short, it is a mess, and everyone is trying to add on this because the existing ecosystem is big and the cost of rewriting everything is deemed too high, so people try to fix python limitations at the library level, sometimes making python more cumbersome to write tha a conventional statically typed language (I am looking at theano).
I don't think that Lua is that much of an improvement, and I do not have big hopes for the JVM. The growing number of projects trying to improve the situation with value types, off-heap allocations and C compatibility are a sign that the VM is not suitable for scientific computing and will not be much better for a while.
I instead hope that Nim will take a leading role for scientific computation. It has an easy syntax and is flexible enough to write libraries such as numpy. At the same time it is fast and C compatibility is trivial. Of course, there is a library ecosystem to build, but I think that this is more doable than trying to work against the language itself. I am trying to start this effort with https://github.com/unicredit/linear-algebra https://github.com/unicredit/linear-algebra but the road is very long and I hope it takes off.
I cannot really comment on Julia as I do not know it welll enough
- osullivj 11y agoSo is this official Unicredit open source? I'm impressed if it is...
- pathsjs 11y agoYes, there are a few open-source projects that you can check on our github :-)
- shoyer 11y ago> The problem with the Python ecosystem is that there exist a lot of libraries for scientific computation mostly incompatible with each other (blaze, numpy, numba, numexpr, dask...). I actually disagree completely with your specific examples. These projects have mostly orthogonal goals and are actually quite compatible with each other -- they all speak NumPy arrays. In fact, three of them (Blaze, Numba and Dask) are sponsored by the same company (Continuum Analytics). Yes, the SciPy ecosystem is a complex beast. There are lots of complementary projects and its not always clear what the best tool for the job is. There are certainly improvements we could make for easier cross-compatibility between libraries (and there are likely projects that could be consolidated), but the number of options you have for scientific computing in Python is an indication of a very robust ecosystem.
- pathsjs 11y agoI am not quite sure. Let me be more specific. Dask requires you to encode computations as a graph explicitly, essentially forcing you to write an AST manually. This means that you are sidestepping the normal language mechanisms to encode program flow, and doing so is essentially incompatible with anything else. Moreover, it does not use numpy arrays - it uses its own internal format that you can convert to a numpy array when needed as explained the overview: http://dask.pydata.org/en/latest/array-overview.html http://dask.pydata.org/en/latest/array-overview.html Numba is another cool project, but - again - it deviates from the mainstream Python. Not every Python function is compilable by Numba, and arithmetic is fixed-size. This means that existing Python code may or may not be compilable by Numba, and even if it works one has to carefully check that arbitrary precision arithmetic is not used. So, it is nice when it works, but is little more than a way to write C-like code with a Pythonish syntax (in which case I greatly prefer Nim, that has actual dispatching on types, generics and so on). Blaze is in a strange position which I do not fully understand. Apparently it uses Numpy arrays, but at the same time the Numpy author states it should be a Numpy replacement: http://technicaldiscovery.blogspot.it/2012/12/passing-torch-of-numpy-and-moving-on-to.html http://technicaldiscovery.blogspot.it/2012/12/passing-torch-... Numexpr requires you to write your code as strings, so it is essentially a separate interpreter. It is nice that it is fast, but it does not play with normal Python code. Not a single project of these works on PyPy, which is the only fast interpreter for non-numeric Python code, nor on Jython, which is the only multithreaded Python interpreter. Do not misunderstand me: I enjoy Python for many things, and internally we use it a lot. But I am frustrated that a lot of effort seems to be dedicated to fix things at the language level starting from a powerful library ecosystem, where I would like to see something solid at the language level that evolves libraries instead
- shoyer 11y agoI appreciate where you are coming from, but on many of these projects you are just misinformed. The entire point of dask is to enable parallel and bigger than memory computations that are not be possible with NumPy arrays. It uses an internal graph representation for deferred computation because deferred computation is basically the only sensible way to do these computations. But in fact, a major point of dask is that users should not need to write these task graphs explicitly. The dask.array API is actually designed to be almost a perfect match for NumPy. The docs do a nice job of explaining the internal abstraction, but understanding it is by no means necessary to use it. (Dask is not my project, but I have made quite a few contributions to it.) Numba certainly does deviate from mainstream Python, but in my experience it works very much in line with the needs and expectations of scientific python users. As for Blaze, it's changed a long way since Travis wrote that blog post in 2012. This website provides a better overview of what Blaze is today: http://blaze.github.io/ http://blaze.github.io/. Confusingly, the name is used for both an ecosystem and one of its major components more specifically. Dask is a component of the larger Blaze project. > Jython, which is the only multithreaded Python interpreter. This is not quite true. Python supports multithreading, and CPython's GIL is less of a obstacle to numerical computing than you might think: http://matthewrocklin.com/blog/work/2015/03/10/PyData-GIL/ http://matthewrocklin.com/blog/work/2015/03/10/PyData-GIL/
- vegabook 11y agoI am intrigued by Nim. Does it have a REPL though? I think an official REPL (not a third party hack) is mandatory for scientific/numerical computing. Ocaml is the only compiled language I can think of which has a credible REPL, officially supported.
- pathsjs 11y agoIt currently does not have a REPL, and that is its main limitation right now. I think it is still suitable for more complex workflows where you would not do it in the REPL anyway - think training a large neural network. I hope something will come after 1.0 is out...
- sea6ear 11y agoSome other (somewhat similar) compiled languages with official REPLs include: - F# - Scala - Haskell (repl is a little different compared to regular code)