Y
HN Search
Hacker News Search
new
|
comments
|
top
|
jobs
maleadt
searching PlanetScale…
1.
▲
2.
▲
3.
▲
4.
▲
5.
▲
6.
▲
6 ms
·
1.
▲
Technical Preview: Programming Apple M1 GPUs in Julia with Metal.jl
(juliagpu.org)
39 points
by
maleadt
4y ago
|
3 comments
2.
▲
by
maleadt
5y ago
Here's a screenshot: https://julialang.org/assets/blog/nvvp.png . Or a recent PR when you can see NVTX ranges from Julia: https://github.com/JuliaGPU/CUDA.jl/pull/760
3.
▲
by
maleadt
5y ago
Yeah, see this section of the documentation: https://juliagpu.gitlab.io/CUDA.jl/development/profiling/ . CUDA.jl also supports NVTX, wraps CUPTI, etc. The full extent of the APIs and tools is available. Source
4.
▲
by
maleadt
6y ago
Impressive! The PTX to SPIR-V compiler must have been quite a bit of work; what's the coverage of the ISA like? With oneAPI I had hoped to get the inverse, a oneAPI implementation for NVIDIA hardware, but I don't think the CUDA dr
5.
▲
by
maleadt
6y ago
You can emit `trap` or `exit` in the PTX code (although that has exposed many bugs in the PTX assembler because it does not expect that kind of often divergent control flow). But even if you'd just have the kernel return and otherwise
6.
▲
by
maleadt
6y ago
Yes, that's fair. I focused on CUDA.jl because it is the most mature, easiest to install, etc. but as I mentioned we're actively working on generalizing that support as much as possible, and as a result support for AMD (AMDGPU.jl)
7.
▲
by
maleadt
6y ago
Since we use fat array objects, and not raw pointers, we know the size of the array and can perform bounds checks at run time. We then have a mechanism to throw an exception and signal it to the CPU to display it there. That's obviousl
8.
▲
by
maleadt
6y ago
The Julia array abstractions make it so that most code is vendor-neutral already, and you execute on whatever GPU back-end you want by using an appropriate array type. For vendor-neutral kernel programming there's GPUArrays.jl and Kern
9.
▲
by
maleadt
6y ago
I'm looking at targeting it from Julia, and the lower-level (Level Zero) API seems rather nice, resembling the CUDA driver API but building on SPIR-V. It's also nice how the API is decoupled from the implementation, so let's
10.
▲
by
maleadt
6y ago
Great talk! Any thoughts on Intel's oneAPI?
11.
▲
by
maleadt
6y ago
That link should probably have been https://juliacomputing.com/industries/gpus.html , but that's rather old content. A better overview is https://juliagpu.org/ , and you can find a demo of Julia
12.
▲
by
maleadt
7y ago
Our view is that to get performance out of a system (here CUDA), it's better not to start abstracting it right away. So we have CUDAnative.jl and CUDAdrv.jl for fairly low-level CUDA programming, albeit in a high-level language. Howeve
13.
▲
by
maleadt
7y ago
Author here, happy to answer any questions! We've been developing and maintaining this toolchain for a while now, so the relevant packages (CUDAnative.jl for kernel programming, CuArrays.jl for a GPU array abstraction) are much more ma
14.
▲
by
maleadt
9y ago
As the author of the underlying framework: the dependency on CUDA is unfortunate indeed, but it was the only viable option at the time. OpenCL tooling was (is) way more fragmented, there's no unified compilation target or back-end, dif
15.
▲
by
maleadt
11y ago
But it relies on Pocket successfully extracting the content. Plenty of sites yielded a broken ereader version on my Kobo (mainly missing or unreadable figures).
16.
▲
by
maleadt
12y ago
A colleague of mine mentioned how this has been the case for quite a bit longer than Haswell, see "Compiler mitigations for time attacks on modern x86 processors" (from 2012). Or Agner Fog's timings pages, where many DIV inst
17.
▲
by
maleadt
12y ago
It is definitely possible to do that in Julia. The reason a didn't yet is purely a manner of priorities, I first focused on wrapping the basic primitives (calling a kernel, marshalling arguments, etc) in a user-friendly way.
18.
▲
by
maleadt
12y ago
OpenCL.jl is purely the runtime part, ie. it still requires you to write manual OpenCL code, after which you can use the julia wrapper to manage that code. My project also provides compiler support for lowering Julia to CUDA assembly, so yo
19.
▲
by
maleadt
12y ago
Thanks for the kind words! This was exactly what I was aiming for: the code (or insights) to be reusable without too much hassle. More so because part of it was developed in the scope of my PhD; I wouldn't want to know how many failed
20.
▲
by
maleadt
13y ago
Wouldn't it be interesting to make the page editable, for example in typical wiki-style like on kernelnewbies.org? Since much of your reader base consists of LLVM developers I'm sure this could further improve the content, as well
21.
▲
by
maleadt
15y ago
Carmack has said somewhere (I can't retrieve the source right now) that an alternative, slower method will be used in the GPL'd codebase, effectively working around Creative's patent claims.
22.
▲
by
maleadt
15y ago
But you're right about the subreddits, and this is one of the more important features which manages to keep acquainted users from leaving the site, even if they are tired of the now 4chan-esque frontpage.