4 ms·
It answers your question because it's a specific case of KA's compile structured loops with annotations to low level CUDA code. If you could express the probl
by dklend122 6y ago
It answers your question because it's a specific case of KA's compile structured loops with annotations to low level CUDA code.
If you could express the problem in generalized index notation, tullio.jl is an even higher level abstraction that uses KA
I don't know about the numpy function you mentioned, but here's a KA matmul. Definitely doesn't look like it has any CUDA specific things. Just a regular loop with some restrictions
@kernel function matmul_kernel!(a, b, c)
i, j = @index(Global, NTuple)
# creating a temporary sum variable for matrix multiplication
tmp_sum = zero(eltype(c))
for k = 1:size(a)[2]
tmp_sum += a[i,k] * b[k, j]
end
c[i,j] = tmp_sum
end
Wish I had time to write the code and compare to pytorch for you though
- dklend122 6y agoHere's a Tulio example. Tulip also provides source to source derivatives. Both KA and Tulio work in the CPU with the same code using Tullio, OffsetArrays # A convolution with cyclic indices mat = zeros(10,10,1); mat[2,2] = 101; mat[10,10] = 1; @tullio kern[i,j] := 1/(1+i^2+j^2) (i in -3:3, j in -3:3) @tullio out[x,y,c] := begin xi = mod(x+i, axes(mat,1)) # xi = ... means that it won't be summed, yj = mod(y+j, axes(mat,2)) @inbounds trunc(Int, mat[xi, yj, c] * kern[i,j]) # and disables automatic @inbounds, end (x in 1:10, y in 1:10) # and prevents range of x from being inferred. It's mostly just math!