4 ms·
Interfaces to HDFS how? PyObject has to get translated somewhere. Arrays of custom element types is a good first start, but boy is dynd a seriously overenginee
by tavert 10y ago
Interfaces to HDFS how? PyObject has to get translated somewhere.
Arrays of custom element types is a good first start, but boy is dynd a seriously overengineered way to accomplish that. I'm referring to array structure, sparsity, symmetry, linear algebraic properties that should be reflected in the type system. Python's type system is lousy for this, and C++ isn't extensible.
Julia can fix performance on heterogeneous data with gradual compiler improvements, and the string representation is due for a major rework. Python can't fix the fact that the language and libraries were not designed to be efficiently JITted, and extension interfaces are closely coupled to the CPython interpreter's API.
- tanlermin 10y agoI don't see why all that can't be reflected in dynd type system and array metadata. It's just early and foundations are still being laid. Revelation julia union etc performance...are these improvents a given, hypothetocal or hope? Doesnt fast code on these types fly kn the face of julia static optimization ethos? Also serious question: dask's use of fast python datastructures like dictionaries gives it a 1ms per task overhead. Is that slow? How does it compare to other dag frameworks like julia etc
- tavert 10y ago> Doesnt fast code on these types fly [in] the face of julia static optimization ethos? What does that even mean? Julia's union types aren't intentionally slow, they just aren't implemented very efficiently yet. Major revisions of how they're implemented are definitely on the roadmap, not far away. For fine-grained parallelism of the type you'd use MPI for, 1ms overhead (assuming that's pure overhead above and beyond the actual cost of data movement) could be significant, sure. If you have calculations that need to go for thousands of individually cheap iterations, it adds up.
- tanlermin 10y agoIt means that I thought Julia wanted to avoid inching too much towards tracing JIT optimization. Well, luckily dask isn't meant for MPI style parallelism. I was wondering how it compared so something like computeframework. Good to hear that union type performance will be optimized. Are there any issues I can follow?