4 ms·
No, not really. VLIW is considered: Multiple Instruction Multiple Data, in each line of assembly you can send out something like 4 (or 8) instruction each with
by wespiser_2018 4y ago
No, not really.
VLIW is considered: Multiple Instruction Multiple Data, in each line of assembly you can send out something like 4 (or 8) instruction each with a different target, and it will work as long as there aren't dependency issues.
GPUs are still Single Instruction Multiple Data (SIMD), for every vector operation you are doing operation: adding vectors, taking a dot production you are only executing a single op at a time.
SIMDs are really close to the RISC/CISC paradigm, and there's various extensions for other types of SIMD processing in different ISAs used today. VLIW is a much different set of assumptions, requiring the compiler to program in the same instruction level parallelism that a superscalar chip will parallelize via it's architectural features (pipelines/branch prediction/et cetera).
- trelane 4y ago> GPUs are still Single Instruction Multiple Data (SIMD), for every vector operation you are doing operation: adding vectors, taking a dot production you are only executing a single op at a time. Sort of. It's both, really. On nvidia at least, the threads in the warp are simd, but between warps it's mimd. And that's before we get into SMs.
- fulafel 4y agoQualcomm basebands apparently use VLIW in baseband in an architecture they called Hexagon: https://www.llvm.org/devmtg/2017-02-04/Halide-for-Hexagon-DSP-with-Hexagon-Vector-eXtensions-HVX-using-LLVM.pdf https://www.llvm.org/devmtg/2017-02-04/Halide-for-Hexagon-DS... VLIW seems quite suited to DSP / SDR applications.
- fulafel 4y agoBetter PDF with architecture picture about where it's used: https://developer.qualcomm.com/download/hexagon/hexagon-dsp-architecture.pdf https://developer.qualcomm.com/download/hexagon/hexagon-dsp-... More links: https://en.wikichip.org/wiki/qualcomm/microarchitectures/hexagon https://en.wikichip.org/wiki/qualcomm/microarchitectures/hex... https://pages.cs.wisc.edu/~danav/pubs/qcom/hexagon_micro2014_v6.pdf https://pages.cs.wisc.edu/~danav/pubs/qcom/hexagon_micro2014... https://blog.tensorflow.org/2019/12/accelerating-tensorflow-lite-on-qualcomm.html https://blog.tensorflow.org/2019/12/accelerating-tensorflow-...
- hajile 4y agoAMD's latest RDNA3 is explicitly VLIW with the VLIW instructions working on SIMD units. Ironically, they are also running head first into the compiler issues with almost nothing around taking advantage of their dual issue potential.