3 ms·
* the SB-SIMD contrib now supports ARM64. (Thanks to Sylvia Harrington) * AVX512 instructions are now supported on X86-64. (Thanks to Robert Smith and Arthu
by wk_end 2mo ago
* the SB-SIMD contrib now supports ARM64. (Thanks to Sylvia Harrington)
* AVX512 instructions are now supported on X86-64. (Thanks to Robert Smith and Arthur Miller)
* additional support for SIMD instructions on ARM64 and X86-64. (Thanks to Arthur Miller)
These seem like pretty awesome additions. Does anyone know how SIMD works in SBCL? Is this at the codegen layer? i.e. can it auto-vectorize or anything like that? Or are these intrinsics you have to explicitly ask for?
- wild_egg 2mo agoNo auto-vectorization yet, but you may find this illuminating: https://www.stylewarning.com/posts/nbody https://www.stylewarning.com/posts/nbody
- BoingBoomTschak 2mo agoAnd the very recent https://old.reddit.com/r/lisp/comments/1v6kyyn/so_they_say_lisp_is_slow/ https://old.reddit.com/r/lisp/comments/1v6kyyn/so_they_say_l...
- e12e 2mo agoLooks like you use it explicitly? https://github.com/sbcl/sbcl/tree/master/contrib/sb-simd https://github.com/sbcl/sbcl/tree/master/contrib/sb-simd
- peri-cl 2mo ago> "AVX512 instructions are now supported on X86-64" This is the first news I've seen on HN in weeks that I am genuinely excited about! I have several AVX-2 hobby projects in Common Lisp, and an AVX-512 machine. It's an unexpected surprise to read this morning that this very useful ISA is suddenly unlocked. I'll be trying it out right away. (edit: Looks like they mean *compiler* support for AVX-512, but not SB-SIMD definitions (yet). So I believe the only way for end-users to call AVX-512 instructions right now is to write custom VOP's). > "Or are these intrinsics you have to explicitly ask for?" There are language SIMD types and you explicitly use SIMD functions that operate on them. I believe it's essentially the same idea as C intrinsics. You can for example write (interactively) (u32.8+ (make-u32.8 0 1 2 3 4 5 6 7) (u32.8 10)) ;; => #<SB-EXT:SIMD-PACK-256 10 11 12 13 14 15 16 17> And that's VPADDD under the hood. Or reading an array (loop with sum = (f32.8 0.0f0) for index below (* 8 (floor length 8)) by 8 do (setq sum (f32.8+ (f32.8-aref array index) sum)) finally (return sum)) which compiles down to a small loop around the vector insts VMOVUPS YMM1, [RDX+RCX*2+1] VADDPS YMM0, YMM1, YMM0 I much prefer it to writing intrinsics in C (and the results are just as good). It's an interactive, exploratory, coding: I write small modular functions, SBCL compiles them on the fly, I glue them together with high-level language constructs.
- Archit3ch 2mo ago> I much prefer it to writing intrinsics in C (and the results are just as good). It's an interactive, exploratory, coding: I write small modular functions, SBCL compiles them on the fly, I glue them together with high-level language constructs. Same here, but in Julia instead. Sometimes I drop down to LLVM intrinsics (e.g. to force lop3.lut on GPUs).
- amno 2mo ago> This is the first news I've seen on HN in weeks that I am genuinely excited about! I have several AVX-2 hobby projects in Common Lisp, and an AVX-512 machine. It's an unexpected surprise to read this morning that this very useful ISA is suddenly unlocked. I'll be trying it out right away. Nice to hear! :) It is the basic compiler support, and lots of instructions added. However, you can't use knor, knoq and similar since they require scheduling of k-masks. Not done yet. But you can certainly use some of avx512 instructions and add yourself if you need some that is not available already. Check https://github.com/sbcl/sbcl/blob/master/src/compiler/x86-64/avx512-insts.lisp https://github.com/sbcl/sbcl/blob/master/src/compiler/x86-64....
- amno 2mo ago> Is this at the codegen layer? On implementation level, it is a codegen layer. It uses a system of macros to generate instructions from a database of instructions. The database is specified manually: https://github.com/sbcl/sbcl/tree/master/contrib/sb-simd/code/instruction-sets https://github.com/sbcl/sbcl/tree/master/contrib/sb-simd/cod... At compile-time, they are converted into "VOPs", i.e. intrinsic functions, which are used by the compiler to emit the actual machine instructions. > can it auto-vectorize or anything like that? Unfortunately, it can't. > are these intrinsics you have to explicitly ask for? Yes, more like higher-level intrinsics. You get quite some automation, but you are requesting manually what you need. More like a DSL, than pure intrinsics. This is how you can use it (as an example): (defun count-lines-and-words-ascii (sap size ws-init-state) (declare (type fixnum size) (type (unsigned-byte 8) ws-init-state) (type sb-sys:system-area-pointer sap) (optimize (speed 3) (safety 0))) (loop with loop-end of-type fixnum = (logandc2 size 127) for i of-type fixnum from 0 below loop-end by 128 with 0x0 of-type u8.32 = (u8.32 #x00) with 0x20 of-type u8.32 = (u8.32 #x20) with 0x0A of-type u8.32 = (u8.32 #x0A) with wa of-type u64.4 = (u64.4 0) with la of-type u64.4 = (u64.4 0) with ws-prev of-type u8.32 = (u8.32 ws-init-state) for c1 = (u8.32-sap-ref sap (+ i 0)) for c2 = (u8.32-sap-ref sap (+ i 32)) for c3 = (u8.32-sap-ref sap (+ i 64)) for c4 = (u8.32-sap-ref sap (+ i 96)) do (flet ((process-chunk (curr prev) (let* ((ctrl (u8.32-sat- (u8.32- curr 9) 4)) (ws (u8.32-or (u8.32= ctrl 0x0) (u8.32= curr 0x20))) (ws-shift (u8.32-alignr ws (u8.32-permute128 prev ws #x21) 15)) (wmask (u8.32-andc1 ws ws-shift)) (lmask (u8.32= curr 0x0A))) (values wmask lmask ws)))) (multiple-value-bind (wm lm prev) (process-chunk c1 ws-prev) (psetf wa (u64.4+ wa (u8.32-sad wm 0x0)) la (u64.4+ la (u8.32-sad lm 0x0)) ws-prev prev)) (multiple-value-bind (wm lm prev) (process-chunk c2 ws-prev) (psetf wa (u64.4+ wa (u8.32-sad wm 0x0)) la (u64.4+ la (u8.32-sad lm 0x0)) ws-prev prev)) (multiple-value-bind (wm lm prev) (process-chunk c3 ws-prev) (psetf wa (u64.4+ wa (u8.32-sad wm 0x0)) la (u64.4+ la (u8.32-sad lm 0x0)) ws-prev prev)) (multiple-value-bind (wm lm prev) (process-chunk c4 ws-prev) (psetf wa (u64.4+ wa (u8.32-sad wm 0x0)) la (u64.4+ la (u8.32-sad lm 0x0)) ws-prev prev))) finally (return (loop for j from loop-end below size with words of-type fixnum = (sum-lanes wa) with lines of-type fixnum = (sum-lanes la) with prev-ws of-type boolean = (logbitp 31 (u8.32-movemask ws-prev)) with tlines of-type fixnum = 0 with twords of-type fixnum = 0 for byte of-type fixnum = (sb-sys:sap-ref-8 sap j) for curr-ws of-type boolean = (or (= byte 32) (<= 9 byte 13)) do (when (= byte 10) (incf tlines)) (when (and (not curr-ws) prev-ws) (incf twords)) (setf prev-ws curr-ws) finally (return (values (the fixnum (+ lines tlines)) (the fixnum (+ words twords)) nil))))))