4 ms·
Sweet! Did you consider using AutoGEMM.py from clBLAS instead of a static GEMM kernel ? I was considering using a polyhedral compiler macro (like PPCG), for wr
by akssri 10y ago
Sweet! Did you consider using AutoGEMM.py from clBLAS instead of a static GEMM kernel ?
I was considering using a polyhedral compiler macro (like PPCG), for writing OpenCL kernels in matlisp, but it's not clear how optimal this would be.
- dragandj 10y agoclBLAS is, in my opinion, hard to build AND hard to integrate. On top of it, this approach gives better performance in most cases even on AMD, and especially on Nvidia. Now, I have AMD hardware, but it is better to create an overall more encompassing library, thus I avoided clBLAS :) When I need to write my own OpenCL kernels, I use ClojureCL - it gives me easy management while still retaining full control of the kernels and their performance.
- akssri 10y agoI found the latest version of clBLAS on Fiji achieves a fantastic ~4 Tflops (on 2^n matrices). NVblas has probably had more resources allocated to it that clBLAS. I'd be positively surprised if the kernels in Neantherdal beat those. Do you plan on adding benchmarks for the GPU calls ? I can help running the clBLAS benchmarks, if you like, since I have a tuned setup. If you want to take a look at it, the AutoGemm generator seems to be a simple python script written in order to overcome the limitations of the C preprocessor. I was considering using its tiling structure, since I already have a Lisp->OpenCL compiler in place (and have had no luck beating it). See, for instance, https://github.com/matlisp/matlisp-opencl/blob/master/tests/blocked-gemm.lisp#L25 https://github.com/matlisp/matlisp-opencl/blob/master/tests/... https://github.com/matlisp/matlisp-opencl/blob/master/src/tensor/copy.lisp#L15 https://github.com/matlisp/matlisp-opencl/blob/master/src/te...
- dragandj 10y agoWhat is the theoretical MAX flops on that Fiji card? I achieve 3.75 TFLOPS on Hawaii, which has much less power than Fiji...
- akssri 10y agoI think it's about 5.6 Tflops. Wow, 3.75 Tflops on Hawaii is very good indeed; I agree this is not something that clBLAS would beat by a wide-margin if at all.
- dragandj 10y agoJudging by this page, https://en.wikipedia.org/wiki/List_of_AMD_graphics_processin... https://en.wikipedia.org/wiki/List_of_AMD_graphics_processin..., Fiji has more than 8 TFLOPS.
- akssri 10y agoAh, you're right; 5.6 Gflops is for Hawaii. That's doubly impressive. I'll make sure to try your kernel then. Thank you!
- akssri 10y agoAh, is this https://github.com/CNugteren/CLBlast https://github.com/CNugteren/CLBlast ? I was looking at, https://github.com/uncomplicate/neanderthal/blob/master/src/opencl/uncomplicate/neanderthal/opencl/kernels/amd_gcn/blas.cl https://github.com/uncomplicate/neanderthal/blob/master/src/...
- dragandj 10y agoYep, CLBlast. My old kernels are from the pre-CLBlast era. Now they are deprecated.
- dragandj 10y agoActually, the man who writes those amazing kernels is Cedric Nugteren. I call HIS kernels.
- deleted 10y ago[deleted]