5 ms·
SIMD and VLIW are the future of microprocessors, and unfortunately it doesn't seem like this ISA will be able to support them.
by dbdjxjcnd 11y ago
SIMD and VLIW are the future of microprocessors, and unfortunately it doesn't seem like this ISA will be able to support them.
- creshal 11y agoRISC-V has a vector mode that can be used for SIMD applications. VLIW has been the future since the 80s, and we're still waiting for the magic wonder compilers that can actually spit out efficient VLIW code. Even GPUs have abandoned VLIW (AMD TeraScale) in favour of RISC (Nvidia, AMD GCN).
- _yosefk 11y agoVLIW DSPs are in every phone. VLIW CPUs are almost certainly a bad idea and as to GPUs, AFAIK VLIW needs to coexist with barrel threading there which might create problems. But VLIW certainly has its place.
- creshal 11y agoIn DSPs, sure, but not in general-purpose CPUs or GPUs.
- adapteva 11y agoMost DSP archs came out of thr 90's and many of the cores today are a reflection of that trend (ceva etc). I worked on the TigerSharc DSP for 8 years and can tell you that from an implementation standpoint they can be a nightmare! Not sure they do have a viable place long term from an economical perspective.
- wsxcde 11y agoI suspect the DSP makers just use VLIW because of interia. They probably don't have the money or incentive to revisit their old decisions. Also wouldn't you say that most of the stuff that used to be implemented on DSPs is now moving into ASICs? I wouldn't be so sure that VLIWs are going to be around forever.
- sklogic 11y agoVLIW is many times cheaper (in terms of power and area) than OoO. It is not going anywhere from the low budget range.
- wsxcde 11y agoBut VLIW and OoO aren't the only two design points. The renewed interest in traditional inorder vector processors. In any case, I do think the point of DSPs was to be area/power efficient for certain specialized algorithms, and a lot of these are just becoming ASICs/specialized accelerators today.
- Marat_Dukhan 11y agonVidia Kepler and Maxwell (i.e. the two latest archs ATM) use VLIW instruction encoding
- creshal 11y agoDo they? I can't find anything about it. VLIW for GPUs is always only mentioned in the context of AMD TeraScale – which has been obsoleted in favour of a RISC architecture five years ago.
- Marat_Dukhan 11y agoI guess that's because unlike AMD, nVidia doesn't officially document the GPU instruction set. But if you disassemble .cubin with nvdisasm, you'd see that code of Kepler/Maxwell is organized in bundles of 4/8 words, where the first word doesn't encode any instruction. Here is what Scott Gray of Nervana Systems, who developed a native assembler for Maxwell, write about it[1]: "Starting with the Kepler architecture Nvidia has been moving some control logic off of the chip and into kernel instructions which are determined by the assembler. This makes sense since it cuts down on die space and power usage, plus the assembler has access to the whole program and can make more globally optimal decisions about things like scheduling and other control aspects. The op codes are already pretty densely packed so Nvidia added a new type of op which is a pure control code. On Kepler there is 1 control instruction for every 7 operational instructions. Maxwell added additional control capabilities and so has 1 control for every 3 instructions." [1] https://github.com/NervanaSystems/maxas/wiki/Control-Codes https://github.com/NervanaSystems/maxas/wiki/Control-Codes
- cmrx64 11y agoFalse. http://hwacha.org/ http://hwacha.org/
- analognoise 11y agoInteresting project; no updates for a year. I wonder how the project is going?
- sagark 11y agoWe've recently released a couple of tech reports on Hwacha: Hwacha Vector-Fetch Architecture Manual: https://www.eecs.berkeley.edu/Pubs/TechRpts/2015/EECS-2015-262.pdf https://www.eecs.berkeley.edu/Pubs/TechRpts/2015/EECS-2015-2... Hwacha Microarchitecture Manual: https://www.eecs.berkeley.edu/Pubs/TechRpts/2015/EECS-2015-263.pdf https://www.eecs.berkeley.edu/Pubs/TechRpts/2015/EECS-2015-2... Preliminary Evaluation Results: https://www.eecs.berkeley.edu/Pubs/TechRpts/2015/EECS-2015-264.pdf https://www.eecs.berkeley.edu/Pubs/TechRpts/2015/EECS-2015-2... M.S. Thesis on Mixed Precision in Hwacha: https://www.eecs.berkeley.edu/Pubs/TechRpts/2015/EECS-2015-265.pdf https://www.eecs.berkeley.edu/Pubs/TechRpts/2015/EECS-2015-2...
- analognoise 11y agoInteresting project; no updates for a year. I wonder how the project is going?
- _chris_ 11y agoThere's a chapter in the current RISC-V manual that explains how you could make a RISC-V-like VLIW ISA. But VLIW is most certainly NOT the "future". VLIW demands an even "mix" of instruction types, and that's largely incompatible with general-purpose application code. And the concept of baking into your ISA what the designer believes is the "perfect functional unit mix" is an anti-pattern. What's the perfect mix depends on the benchmark, and it changes from basic block to basic block. A history of failed VLIW projects can attest to this. A dynamic superscalar is far superior, even in power-efficiency.
- protomyth 11y ago"VLIW demands an even "mix" of instruction types,and it's hugely incompatible with general-purpose application core." Has anyone ever done a study / experiment of a VLIW with multiple hardware threads and how that would impact the need for an even mix?
- _chris_ 11y agoThe studies I've seen have shown that the FU mix is heavily skewed and changing on every basic block, particularly when you offload the DLP to a more efficient vector/SIMD unit. I'm not sure I see how MT would solve the mix problem, if each thread gets an issue cycle (and each thread itself has a bad mix).
- protomyth 11y agoI was just thinking that the the bad mixes might on average fill in the gaps to what the processor actually had for resources.
- creshal 11y agoEven GPUs, which have arguably far more predictable instruction mixes, struggled massively to get a useful utilization on VLIW architectures – AMD tried two different ones from 2006 to 2011 before they finally gave up on the concept and started using RISC architectures like Nvidia had been using all the time.
- legulere 11y agoBy what I've been told, VLIW makes only really sense in some use cases such as DSP processors. With general purpose computing it happens way too often that you can't find enough instructions that are independent of each other. For the case where it is possible to execute instructions in parallel, you can make your CPU superscalar. The simple nature of the RISC-V ISA probably should make supercalarability easy and performant.
- zhemao 11y agoWe already have a RISC-V superscalar out-of-order core, the Berkeley Out-of-Order Machine (BOOM). https://github.com/ucb-bar/riscv-boom https://github.com/ucb-bar/riscv-boom