4 ms·
No, the parent comment is correct. Tulllio.jl has been getting up to openBLAS levels of performance Julia also is on the cusp of support for tensor cores, mix
by dklend122 6y ago
No, the parent comment is correct. Tulllio.jl has been getting up to openBLAS levels of performance
Julia also is on the cusp of support for tensor cores, mixed precision multiplication and bfloat16 codegen. New PR just merged last week
- xiphias2 6y agohttps://discourse.julialang.org/t/blas-vs-cublas-benchmark/46205/3 https://discourse.julialang.org/t/blas-vs-cublas-benchmark/4... Here's an OpenBLAS vs CuBLAS performance comparision. Basically CPU is at this point outdated technology for matrix multiplication. Your second comment is more interesting, I'd be happy if Julia gets to the point when it can beat PyTorch or even JAX on training Lambda networks or Performers.
- dklend122 6y agoI understand CPU is slower, but it's not going anywhere, and matmul is just a benchmark. It's extremely impressive that a high level (using index notation for multiple backends) Julia library can compete with hand tuned kernels. Also bodes well for Julia compiler tech generally that can easily be extended Julia is going to get to the point of beating pytorch, just taking a but more time due to smaller team and approaching it from a more general position..so when it does it will be more flexible and ergonomic and easily extensible to new techniques, in pure Julia. 1.6 will be a big step as much of the compiler hacks underlying the current ad/gpu codegen (which just fell out accidentally of lispy design) will be replaced with proper tooling for composable compiler passes on typed IR. This will be a phase change imo. There's already a new faster AD that's almost ready for debut based on 1.6 tech
- xiphias2 6y agoThat sounds amazing, thanks for the update! I stopped using Julia because it was very frustrating that I payed thousands of dollars for a laptop with GeForce RTX 2070 card and the I couldn't make use of it, it felt like I wasted a lot of money...that was my main decision for moving to PyTorch. I know that Julia as a language should be able to beat it easily, but it felt to me that the community was prioritizing CPU over GPU for a long time. I'm happy that it's changing now.
- eigenspace 6y agoThe GPU stuff is basically just a teaser into the future, it's still much more immature than that CPU stuff for Tullio. > Basically CPU is at this point outdated technology for matrix multiplication. This is an incredibly naive statement. There are many many circumstances where you need fast matmul on the CPU even if you have a GPU available because it takes too long to send the memory to the GPU and fetch it once the kernel runs. If you have many chained matmuls, then yes the GPU is your friend. That said, regardless of how useful the CPU is in practice for matrix multiplication, BLAS kernels are some of the most overengineered, most optimized pieces of code in wide use. Tullio is incredibly flexibly and general. The fact that matrix multiplication falls out as a special case in Tullio and can outperform OpenBLAS and match MKL is absolutely stunning.
- xiphias2 6y agoCan you give concrete examples when data processing on the CPU is clearly better than on the GPU? Before I was working with time series, and there the CPU really shined, but right now I'm classifying sick patients using DNA methylation data (about 100-300 GB / dataset). I'm using a convolutional network that already gives me state of the art results, but I want to experiment with self attention based neural networks. While I'm doing this, most of the biologists are still using logistic regression that can be computed on the CPU. I see more and more domains (even simulations) where neural networks can outperform classical methods if you know how to apply them, and CPUs can never match the GPU performance. Also as I wrote, the slowest part of all these networks is the matrix multiplication (self attention is especially depending on it).
- eigenspace 6y ago> Can you give concrete examples when data processing on the CPU is clearly better than on the GPU? Pretty much anything where the matmul is important to performance, but doesn't dominate it and requires serial steps. A classic example off the top of my head would be solving a matrix differential equation[1]. Here's an example in Julia using DifferentialEquations.jl and CUDA.jl for the GPU part: This is solving the differential equation du/dt = A * u for the cases where u and A are arrays on the GPU versus when they are arrays on the CPU. I'm doing the solving in-place to try and minimize the amount of expensive allocations. using DifferentialEquations, CUDA function mysolve(u0, A, tspan) f!(du, u, A, t) = mul!(du, A, u) prob = ODEProblem(f!, u0, tspan, A) solve(prob) end mysolve (generic function with 1 method) julia> let n = 50, u0 = rand(Float32, n), A = randn(Float32, n, n) # Make GPU versions of u0 and A cuu0 = cu(u0) cuA = cu(A) # Do a run of the code so there's no compiler latency biasing the results mysolve( u0, A, (0.0, 1.0)) mysolve(cuu0, cuA, (0.0, 1.0)) # Now run the compiler code and time it's execution: @time mysolve( u0, A, (0.0, 1.0)) @time mysolve(cuu0, cuA, (0.0, 1.0)) end; 0.000314 seconds (621 allocations: 98.281 KiB) 0.007218 seconds (10.33 k allocations: 382.266 KiB) In this case, the CPU code was ~20x faster than the GPU code. True, I could get a better GPU (I'm using a RTX 2060), but I could also get a better CPU (Ryzen 5 2600). There's not much point in using a GPU for this example, and I think that just comes down to the fact that there's a non-trivial amount of serial work that needs to be done between the matmuls. [1] https://en.wikipedia.org/wiki/Matrix_differential_equation https://en.wikipedia.org/wiki/Matrix_differential_equation