6 ms·
Granted special Numba class syntax, and nowhere near as elegant or integrated as Julia. Dask distributed interfaces directly with HDFS...no need for serializa
by tanlermin 10y ago
Granted special Numba class syntax, and nowhere near as elegant or integrated as Julia.
Dask distributed interfaces directly with HDFS...no need for serialization.
Dynd is specifically designed to deal with arrays of custom types...criticism sounds like it would be more aptly directly towards numpy.
once it gains Numba binding it will lose the Python overhead.
Re Dask...is 1ms per task too much? It has a threading scheduler that can work with numpy arrays to release the GIL....so not bound by that with numerical code.
Regarding regular list and text processing... Julia's poor performance with heterogenous data actually makes it slower than python for cleaning the corresponding dirty and text datasets.
- tavert 10y agoInterfaces to HDFS how? PyObject has to get translated somewhere. Arrays of custom element types is a good first start, but boy is dynd a seriously overengineered way to accomplish that. I'm referring to array structure, sparsity, symmetry, linear algebraic properties that should be reflected in the type system. Python's type system is lousy for this, and C++ isn't extensible. Julia can fix performance on heterogeneous data with gradual compiler improvements, and the string representation is due for a major rework. Python can't fix the fact that the language and libraries were not designed to be efficiently JITted, and extension interfaces are closely coupled to the CPython interpreter's API.
- tanlermin 10y agoI don't see why all that can't be reflected in dynd type system and array metadata. It's just early and foundations are still being laid. Revelation julia union etc performance...are these improvents a given, hypothetocal or hope? Doesnt fast code on these types fly kn the face of julia static optimization ethos? Also serious question: dask's use of fast python datastructures like dictionaries gives it a 1ms per task overhead. Is that slow? How does it compare to other dag frameworks like julia etc
- tavert 10y ago> Doesnt fast code on these types fly [in] the face of julia static optimization ethos? What does that even mean? Julia's union types aren't intentionally slow, they just aren't implemented very efficiently yet. Major revisions of how they're implemented are definitely on the roadmap, not far away. For fine-grained parallelism of the type you'd use MPI for, 1ms overhead (assuming that's pure overhead above and beyond the actual cost of data movement) could be significant, sure. If you have calculations that need to go for thousands of individually cheap iterations, it adds up.
- tanlermin 10y agoIt means that I thought Julia wanted to avoid inching too much towards tracing JIT optimization. Well, luckily dask isn't meant for MPI style parallelism. I was wondering how it compared so something like computeframework. Good to hear that union type performance will be optimized. Are there any issues I can follow?