4 ms·
There is work underway to bridge this gap. The fundamental bet by my company is actually basically the same one that Julia makes: higher level languages combin
by pwang 12y ago
There is work underway to bridge this gap. The fundamental bet by my company is actually basically the same one that Julia makes: higher level languages combined with dynamic compilation are the only route to performance in our modern world of heterogeneous hardware.
My company sells a Python-to-GPU and Python-to-x86 compiler. It matches or beats hand-rolled C and C++ for some cases, and is loads easier to deal with when it comes to CUDA work. It's still in its infancy, but the approach (in my mind) is definitely validated.
If you think about it, why should C be the speed king? Most programmers cannot reason about cache coherency to save their lives and that's the dominant performance cost for real performance. Many of the big numerical codes follow some high level patterns of data movement; if a compiler has greater visibility into the structure of both the data and the algorithm, it has an easier time parallelizing and optimizing, than if it has to rely on the programmer to "lower" the algorithm to a layer that's just a hair above RTL. "C is just portable assembly" and all that.
Additionally, FWIW, if you really want to pick nits, a lot of NumPy is actually not written in C, but rather a custom macro template system which then creates C code for each of the core types (int8, uint8, int16, etc. etc.). So even for the low-level guts, code generation (albeit a very simple mechanism) is the current approach.