4 ms·
This has been my general experience writing numerical linear algebra methods in Julia. Optimizing a specific algorithm, eliminating unnecessary allocations is t
by spacedome 6y ago
This has been my general experience writing numerical linear algebra methods in Julia. Optimizing a specific algorithm, eliminating unnecessary allocations is the first thing I do, and can give large performance gains, especially for iterative methods.
It can be a bit annoying, and can result in much less readable code (eg having to explicitly write things like mul!(C, A, B)). It ends up looking a lot more like the FORTRAN it is meant to replace, and if you aren't careful you lose the ability to use generic types, though multiple dispatch is a great solution. The worst case I have found is iteratively calling linear algebra solvers (Eig, SVD, LU) which do not have any option for preallocating and reusing work arrays. The only way to do this now is to call the BLAS/LAPACK routines directly with ccall, which is a huge pain. There was an attempt to fix this with PowerLAPACK.jl, but it seems abandoned, so for the moment writing optimized methods in FORTRAN/C and calling from Julia is sometimes still preferable.
Julia does quite well though, so this only seems necessary for things like core NLA libraries, for example I would not try rewriting ARPACK in Julia for a performance gain. The gain in flexibility for writing these in Julia is absolutely huge though, and I definitely recommend it. The Grassmann.jl library has many examples of what a language like Julia makes possible, or DifferentialEquations.jl.
- Xcelerate 6y agoI love Julia and have used it for years, but this is really my only complaint with my language. I used to do HPC work and ended up spending a lot of time tracking down where tiny bits of memory were being allocated in for loops. A lot of my routines ended up having “blocks” of memory to pass in as arguments, which as you mention, moves the language a bit closer to C or Fortran IMO.
- snicker7 6y agoMinimizing allocations (and deallocations) and optimizing memory layout are just things you have to do for HPC in any language. Though I think that it is easier in Julia than in C/Fortran.
- etik 6y agoThere is a recent effort [1] to provide low-level support for faster operations by transforming user code to take advantage of a compiler's instruction set, memory packing, etc. This is being expanded upon to essentially provide a Julia-native BLAS. Some of the benchmarks are even competitive with or beat Intel MKL (calibrate that statement appropriately to your level of trust in benchmarks). I wouldn't count out a Julia ARPACK implementation just yet. [1] LoopVectorization: https://github.com/chriselrod/LoopVectorization.jl https://github.com/chriselrod/LoopVectorization.jl Announcement post and discussion: https://discourse.julialang.org/t/ann-loopvectorization/32843 https://discourse.julialang.org/t/ann-loopvectorization/3284...
- stabbles 6y agoYou're probably confusing ARPACK with BLAS / LAPACK. Pure julia ARPACK already exists, e.g. https://github.com/haampie/ArnoldiMethod.jl/ https://github.com/haampie/ArnoldiMethod.jl/. A competive BLAS-gemm is implemented here https://github.com/YingboMa/MaBLAS.jl/blob/master/src/gemm.jl https://github.com/YingboMa/MaBLAS.jl/blob/master/src/gemm.j... (single-threaded). A LAPACK-like library could be https://github.com/JuliaLinearAlgebra/GenericLinearAlgebra.jl https://github.com/JuliaLinearAlgebra/GenericLinearAlgebra.j...
- etik 6y agoI was mostly referring to the parent comment's suggestion that low level numerical libraries wouldn't benefit from a pure Julia implementation, specifically to the statement that it was still better to write optimized C/FORTRAN and call from Julia. Indeed, MaBLAS, which you linked, is built on top of LoopVectorization.jl. I don't know how well ArnoldiMethod.jl compares with ARPACK, but if there is a gap my suggestion is simply that these recent developments might help bridge it :)
- stabbles 6y agoFWIW, MaBLAS currently does not depend on LoopVectorization.jl, the code to generate kernels is all handwritten.
- 6y ago