4 ms·
I am not quite sure. Let me be more specific. Dask requires you to encode computations as a graph explicitly, essentially forcing you to write an AST manually.
by pathsjs 11y ago
I am not quite sure. Let me be more specific.
Dask requires you to encode computations as a graph explicitly, essentially forcing you to write an AST manually. This means that you are sidestepping the normal language mechanisms to encode program flow, and doing so is essentially incompatible with anything else. Moreover, it does not use numpy arrays - it uses its own internal format that you can convert to a numpy array when needed as explained the overview: http://dask.pydata.org/en/latest/array-overview.html http://dask.pydata.org/en/latest/array-overview.html
Numba is another cool project, but - again - it deviates from the mainstream Python. Not every Python function is compilable by Numba, and arithmetic is fixed-size. This means that existing Python code may or may not be compilable by Numba, and even if it works one has to carefully check that arbitrary precision arithmetic is not used. So, it is nice when it works, but is little more than a way to write C-like code with a Pythonish syntax (in which case I greatly prefer Nim, that has actual dispatching on types, generics and so on).
Blaze is in a strange position which I do not fully understand. Apparently it uses Numpy arrays, but at the same time the Numpy author states it should be a Numpy replacement: http://technicaldiscovery.blogspot.it/2012/12/passing-torch-of-numpy-and-moving-on-to.html http://technicaldiscovery.blogspot.it/2012/12/passing-torch-...
Numexpr requires you to write your code as strings, so it is essentially a separate interpreter. It is nice that it is fast, but it does not play with normal Python code.
Not a single project of these works on PyPy, which is the only fast interpreter for non-numeric Python code, nor on Jython, which is the only multithreaded Python interpreter.
Do not misunderstand me: I enjoy Python for many things, and internally we use it a lot. But I am frustrated that a lot of effort seems to be dedicated to fix things at the language level starting from a powerful library ecosystem, where I would like to see something solid at the language level that evolves libraries instead
- shoyer 11y agoI appreciate where you are coming from, but on many of these projects you are just misinformed. The entire point of dask is to enable parallel and bigger than memory computations that are not be possible with NumPy arrays. It uses an internal graph representation for deferred computation because deferred computation is basically the only sensible way to do these computations. But in fact, a major point of dask is that users should not need to write these task graphs explicitly. The dask.array API is actually designed to be almost a perfect match for NumPy. The docs do a nice job of explaining the internal abstraction, but understanding it is by no means necessary to use it. (Dask is not my project, but I have made quite a few contributions to it.) Numba certainly does deviate from mainstream Python, but in my experience it works very much in line with the needs and expectations of scientific python users. As for Blaze, it's changed a long way since Travis wrote that blog post in 2012. This website provides a better overview of what Blaze is today: http://blaze.github.io/ http://blaze.github.io/. Confusingly, the name is used for both an ecosystem and one of its major components more specifically. Dask is a component of the larger Blaze project. > Jython, which is the only multithreaded Python interpreter. This is not quite true. Python supports multithreading, and CPython's GIL is less of a obstacle to numerical computing than you might think: http://matthewrocklin.com/blog/work/2015/03/10/PyData-GIL/ http://matthewrocklin.com/blog/work/2015/03/10/PyData-GIL/
- tadlan 11y agoNumba has loop fusion, runtime, and soon multithreading parallel vectorized functions. How is that c like?