6 ms·
How does this compare to CuPy (https://cupy.chainer.org/ https://cupy.chainer.org/) ? It is now independent from Chainer, is highly compatible with numpy and su
by lhenault 9y ago
How does this compare to CuPy (https://cupy.chainer.org/ https://cupy.chainer.org/) ?
It is now independent from Chainer, is highly compatible with numpy and supports both CUDA and CuDNN.
- Loic 9y agoFor the CUDA part I cannot tell, but Numba is also compiling on the fly your Python code into machine code using LLVM. This where it shines. For example, instead of pushing your code into Cython or a Fortran library, you can keep writing in simple Python and get your code to run in some cases nearly as fast as Fortran. This is my use case. I haven't used the CUDA features yet.
- fnl 9y agoBut LLVM doesn't support vectorizing, like AVX or SSE4, right? So I don't think that would be nearly as fast as fully (Intel-) CPU optimized code... EDIT: Let me hedge that a bit, to advanced AVX instructions, as LLVM can do simple loops and such.
- lliiffee 9y agoI believe that LLVM 6 has finally introduced this, e.g. see http://llvm.org/docs/Vectorizers.html#vectorization-of-function-calls http://llvm.org/docs/Vectorizers.html#vectorization-of-funct...
- fnl 9y agoOh, cool, I see LLVM now even sports (basically all of) SSE4.2 and AVX-512. As always, that project amazes... :-)
- HelloNurse 9y agoDo you mean LLVM 5?
- fnl 9y agoYes, indeed even LLVM 4 already improved its AVX-512 support. http://releases.llvm.org/4.0.0/docs/ReleaseNotes.html http://releases.llvm.org/4.0.0/docs/ReleaseNotes.html Really impressive how many new things came to LLVM this year!
- Joky 9y agoUh, what you're pointing at was introduced in 2012 in LLVM.
- fnl 9y agoOnly in parts, not all instructions, and some functionality it did have was buggy. 4 and 5 are much more advanced/competitive on SIMD issues, it seems. Edit: Oh, sorry you meant that other guy's link to LLVM's vectorization tutorial. Ignore my reply ...
- radarsat1 9y agoYour comment surprised me as clang is regarded to be pretty competitive these days (compared to gcc). I don't know the current state of things but a quick search revealed at least one sentence, http://llvm.org/docs/Vectorizers.html http://llvm.org/docs/Vectorizers.html, "the loop below will be vectorized on Intel x86 if the SSE4.1 roundps instruction is available." So it seems SSE4 is supported by LLVM? I was going to say maybe it's a new thing, but the following post also talks about SSE4 and is from 2011: http://blog.llvm.org/2011/12/llvm-31-vector-changes.html http://blog.llvm.org/2011/12/llvm-31-vector-changes.html Maybe it only supports a subset of SSE4? Do you know the details, compared to other compilers?
- fnl 9y agoThe source of my confusion is that the last time I looked into LLVM's SIMD support was in the context of looking at Rust, a bit more than a year ago or so, and back then my conclusion was that neither (Rust or LLVM, then in version 3) are very good tools for that. It seems I was very wrong, at least on the LLVM side. EDIT: Sorry, to reply to your question, my concern is not GCC vs. clang; If you want max out your vector ops, I would suggest you should compare to ICC as the "standard", at least on Intel CPUs.
- radarsat1 9y agoYep, I understand about icc. I actually wonder what is so difficult about optimising the way icc does, what does it actually do so much better than gcc? Anyone know of a good analysis?
- fnl 9y agoI'd mostly attribute that to the MKL and their ability to just have to deal with their own instructions. But that's just an "educated guess".
- mattnewton 9y agoProbably also having full time people working in the same company as the hardware guys who's job is to make the hardware look good.
- Coding_Cat 9y agoI'm not sure what parts are clang and what parts are LLVM, but I recently did some tests comparing g++ to clang++ for auto-vectorization and they were very on par. I'd even say clang was a little better than g++ with full optimizations turned on on average. This was on an AVX2 machine, testing the (auto)-vectorization performance of expression templates. Anything with a compile-time unknown stride or a random gather failed horribly with both. Using (semi-)explicit vectorization turned out to be much faster still.
- fnl 9y agoImpressive; I didn't take a look since the 3.x series, so I am totally stunned by the amount of "love" that LLVM has received lately (4, which was released just a few months ago, and up to the coming version 6, that is under development still).
- marmaduke 9y agoIt does and this can be done sometimes automatically by Numba/LLVM https://github.com/numba/llvmlite/issues/270 https://github.com/numba/llvmlite/issues/270
- keldaris 9y agoWhile your statement can certainly be correct for a sufficiently stringent definition of "advanced" (and replacing "instructions" with "code patterns", etc.), in my experience in using both clang for C++, Julia (a language largely reliant on the LLVM optimizer) and LDC (a D language compiler on top of the LLVM) vectorization support is competitive (and sometimes superior) to GCC. Comparisons with ICC are somewhat more complicated, ICC is unparalleled for specific patterns and often fails horribly in other things, the variance there is quite huge. Disclaimer: by "vectorization" I'm referring to SSE4, AVX and AVX2, I haven't had a chance to try out AVX512 yet.
- dagw 9y agoOne advantage of Numba is that it doesn't require CUDA. You can easily write your code so that if the machine it's running on has CUDA then it will use that and if it doesn't it will just JIT it for the CPU.