3 ms·
Regarding your performance work, you've nerd sniped me into looking for analytical tricks to speed it up ;) We'll see... Regarding the readability, I suppose i
by ianhorn 5y ago
Regarding your performance work, you've nerd sniped me into looking for analytical tricks to speed it up ;) We'll see...
Regarding the readability, I suppose it's a matter of taste at this point, but if I swapped `import numpy as np` for `from numpy import newaxis, exp, sqrt`, then IMO, the numpy is more readable:
A = a[newaxis]
M = exp(1j * k * sqrt(A**2 + A.T**2))
But then my tastes also run towards preferring the handwritten definitions to be in terms of vectors and transposes instead of elements and indices, at least until you get to many indexed tensors.
Side note: any idea why the python and fortran are exp(i k... while the julia is exp((100 +i) i...)? Is it something I overlooked?
- eigenspace 5y ago> Regarding your performance work, you've nerd sniped me into looking for analytical tricks to speed it up ;) We'll see... Looking forward to it! These microbenchmarks are very fun to explore. > Side note: any idea why the python and fortran are exp(i k... while the julia is exp((100 +i) i...)? Is it something I overlooked? Oh, that is a partially applied edit I guess. When the author posted on the julia forum, people pointed out that since the exponent was pure imaginary, it could be speeded up even more with cis(...) instead of exp(im * ...) but then the author claimed that exp((100 + im) * ...) was more representative of his actual workflow and I guess changed the julia version in his blogpost but not the Python or Fortran versions.
- ianhorn 5y agoI got bored with trying to find an analytical boost, but I benchmarked a couple IMO super readable python versions (basically what's in my original comment after making the (100+i)i change): https://colab.research.google.com/drive/1ABrZJlm8pwB6_Sd6ayOPj_zf9rxUWk8Z?usp=sharing https://colab.research.google.com/drive/1ABrZJlm8pwB6_Sd6ayO... On my macbook, using XLA's jit in python gave about a 12-15x speedup on CPU over OP's solution, which was pretty cool, but I'm too lazy to figure out how to install and benchmark Julia on my machine. Applying a 12-15x speedup would at least beat the Julia MT solution in OP, and you've got to admit `exp(CONST * sqrt(A**2 + A.T**2))` is a pretty clean way to do it. Then I ran on whatever GPU colab decided to give me (a P100), and for just adding a decorator, it's a 1000x-1900x speedup (better as n goes up). Hence my current honeymoon period with jax. I love the speed vs readability tradeoff.
- celrod 5y ago> I'm too lazy to figure out how to install and benchmark Julia on my machine. Assuming you're eager to try once you've found out: You should be able to simply download and unpack a binary from: https://julialang.org/downloads/ https://julialang.org/downloads/ I'd strongly recommend going with the current stable release (1.6.1). To start the Julia REPL: bin/julia To install packages: using Pkg Pkg.add("BenchmarkTools") Pkg.add("LoopVectorization") then you can copy/paste code from eigenspaces link to discourse to run Julia benchmarks. E.g., you can copy/paste from https://discourse.julialang.org/t/i-just-decided-to-migrate-from-python-fortran-to-julia-as-julia-was-faster-in-my-test/61188/19?u=elrod https://discourse.julialang.org/t/i-just-decided-to-migrate-... if you define using LoopVectorization const var"@tvectorize" = var"@tturbo" (because `@tvectorize` has since been renamed on the latest releases)
- celrod 5y agoAlso, note that Julia starts with only a single thread by default. You'd either need to add `-t4` when starting julia, e.g. `julia -t4`, or set the environmental variable `JULIA_NUM_THREADS=4` for 4 threads, for example.
- zwaps 5y agoLooks good, but I actually love that Julia gets at least the same performance right out of the box, with full interoperability and the great multiple dispatch logic. Would be interesting to run your optimized code against the Julia optimized code that was linked to in this thread. Or run a Julia GPU benchmark as well. As for loops versus vectorization - I personally am used to vectorized code as it has always been a requirement. Sometimes it makes sense to use (e.g. Matrix Algebra stuff). On the other hand, I have also been in many situations where I could not vectorize my code. More code than I'm happy to admit includes list or dict comprehensions or even loops. Sure, not all performance critical. Nevertheless, the prospect of performance no matter what is pretty exciting. Finally, threading and multiprocessing in python is a headache compared to Julia imo. Sometimes I just miss the good ol' "parfor" from Matlab. All in all, Julia is a really appealing value proposition to people doing numeric computing.