5 ms·
I have found writing CUDA code is much simpler than writing correct multi-threaded AVX2/AVX-512 code.
by shubuZ 5y ago
I have found writing CUDA code is much simpler than writing correct multi-threaded AVX2/AVX-512 code.
- dragontamer 5y agoIf you need CPU-side SIMD, then try ispc: https://ispc.github.io/ https://ispc.github.io/ Its pretty much the OpenCL-model, except it compiles into AVX2 code / AVX512 code. Very similar to CUDA / OpenCL style programming. Its not single-source like CUDA, but it largely accomplishes the programming model IMO.
- gnufx 5y agoWhy not a standard? OpenMP is more than pretty much C(++) and Fortran, and has offload inspired by the needs of the Sierra supercomputer.
- dragontamer 5y agoOpenMP 4.5+ looks very promising to me (particularly the "simd" keyword associated with for-loops). But open-source implementations of OpenMP are somewhat lackluster... at least last time I checked it out. Maybe its time I revisit it. I've always thought the OpenMP spec was being written by highly competent programmers. They seem to "get" what is needed. But the question is if I can get my hands on any OpenMP implementation to actually play with. Not all of us can afford IBM's compiler suite!
- jpf0 5y agoLLVM has an openMP implementation
- dragontamer 5y agoThe task-based parallelism in LLVM leaves much to be desired however. Ideally, you'd want a more efficient implementation. But yeah, good enough to play with. But maybe not good enough to achieve high levels of performance. The SIMD stuff is probably simple enough to implement... maybe I should checkout how well LLVM works with OMP SIMD keywords.
- jpf0 5y agoCan you comment on experience (or contact me) regarding implementation efficiency? We have recently implemented task-based parallelism in the J language with openMP[0]. Improvements or critiques are appreciated. SIMD instructions there have been coded directly rather than via pragmas. [0] https://www.monument.ai/m/parallel https://www.monument.ai/m/parallel
- dragontamer 5y agoI can't say that my critiques are based off of personal experience. But mostly about microbenchmarks I've read that other people have talked about. I am probably a bit out of date, since its been a while since I last played with OpenMP. I'm looking at the benchmarks I used to look at, and they're all from 2014 or earlier. So maybe I really should double-check modern implementations. We all know GCC 4.x and LLVM 3.x are an eternity ago, so I probably should revisit their performance. For example: https://www.phoronix.com/scan.php?page=article&item=llvm_clang_openmp&num=1 https://www.phoronix.com/scan.php?page=article&item=llvm_cla... And back then, it was pretty well known that OpenMP implementations were slower than commercial (such as Intel ICC or IBM's OpenMP implementation).
- gnufx 5y agoIn LLVM or in libomp? I don't know what omp simd is likely to get you over autovectorization. I know of cases where it was thought necessary (-fopenmp-simd, without -fopenmp) but wasn't with recent GCC.
- dragontamer 5y agoAutovectorization has issues with function calls. "#pragma omp declare simd" applies over a function call, which then allows that function to be used inside of a "#pragma omp for simd" loop. A few keywords here and there really help the autovectorizer achieve closer to CUDA-like environments (like... actually having your SIMD code extend "through" a function call, so you can start splitting up the work a bit better). EDIT: Here's an example from Intel's ICC: https://software.intel.com/content/www/us/en/develop/documentation/cpp-compiler-developer-guide-and-reference/top/optimization-and-programming-guide/vectorization/explicit-vector-programming/simd-enabled-functions.html https://software.intel.com/content/www/us/en/develop/documen...
- gnufx 5y agoI should really find time to do comparisons with the compilers to hand, including XL, on the NAS benchmarks, as I've never seen that, though it must have been done. I think we have a "community" version of XL, i.e. no support, like basically everything else. I wasn't aware there was anything much wrong with GCC and libgomp or libomp, but then I haven't measured.
- einpoklum 5y agoIt's much more limited in expressivity than OpenCL/CUDA, IIANM.
- 37ef_ced3 5y agoUse a domain-specific compiler to generate custom, stand-alone, massively multi-threaded AVX-512 inference C code: https://NN-512.com https://NN-512.com The generated code is easily twice as fast as TensorFlow's AVX-512 kernels (Intel's oneAPI).
- etaioinshrdlu 5y agoYour project is well engineered but no matter how many times you post it here, you won’t get real traction without a different development model, such as building on github. This is especially true in a high churn, high cost field like ML. I also think you are being anonymous unnecessarily.
- johndough 5y agoSome user feedback: I tried to get the ResNet50 example to work, but I gave up after 2 hours. There really should be an end-to-end example like "./resnet50_example.py monkey.jpg". A few points where I struggled: - What are the names in the ResNet50Params struct? I looked at a few popular ResNet50 implementations, but I could not find any correspondence and since the names have been sorted, the order of the members might as well be random. I thought about matching the parameters by shape instead, but the chance of getting that correct is almost zero. Why no natural sort? Or even better, a naming scheme that sorts in the same order as execution order, for example with zero-prefixed numbers like "layer05" - 6 MB C files are somewhat ridiculous. There should at least be a download link somewhere. Copy & pasting from the website would take forever because of scrolling. Also it is impossible to verify that that code monstrosity is not doing anything evil. - Go seems like an odd choice when almost every machine learning project these days is developed in Python. - GCC throws a few hundred warnings: ResNet50.c: In function ‘ResNet50ThreeArrangeDats7Callee1’: ResNet50.c:57009:15: warning: unused variable ‘rel27’ [-Wunused-variable] 57009 | ptrdiff_t rel27 = j62-0; | ^~~~~ and ResNet50.c: In function ‘ResNet50NetCreate’: ResNet50.c:77852:20: warning: taking address of packed member of ‘struct ResNet50Params’ may result in an unaligned pointer value [-Waddress-of-packed-member] 77852 | params1->bn1Means, | ~~~~~~~^~~~~~~~~~ - You would get more feedback if there was a platform to discuss such issues (e.g. GitHub) instead of hijacking random threads on HackerNews. - EDIT: I just realized I could probably just write a parser for that graph file format. This would be much easier if it was something standard like JSON instead. - EDIT2: I think some shapes in the graph file definition are incorrect (possibly reversed per block?). For example, tensor "one1" should have "ToChannels=64" instead of "ToChannels=256".