3 ms·
Compiler does the reordering. Not the GPU itself. Even in OoO CPU’s the ordering by HW doesn’t affect FP results. As a rule of thumb shaders in graphics are com
by sharpneli 6y ago
Compiler does the reordering. Not the GPU itself. Even in OoO CPU’s the ordering by HW doesn’t affect FP results. As a rule of thumb shaders in graphics are compiled as if one had passed -ffast-math to the compiler. Works perfectly in most cases, but not everywhere.
Nowadays MS has actually published the spec for DX that was closed previously. See https://microsoft.github.io/DirectX-Specs/d3d/archive/D3D11_3_FunctionalSpec.htm#FloatingPointRules https://microsoft.github.io/DirectX-Specs/d3d/archive/D3D11_... for differences of strict IEEE behaviour. In Cuda and OpenCL one can get way closer. As an example for performance reasons one might want to flush denorms to zero. But DX mandates that. So no denormals for you. In CL and Cuda they’re usable by default.
As for the command queues. I’ve often used them in Cuda. Just to get overlap between kernel executions. In DX12 one can do that by omitting barriers. DX11 allows no such feat.
- Const-me 6y ago> Compiler does the reordering. It does, but you can always disassemble the DXBC and see what happened to your HLSL code. > DX11 allows no such feat. ID3D11DeviceContext::Dispatch is asynchronous just like CopyResource. Dispatch multiple shaders, and unless they have data dependencies (i.e. same buffer written by one as UAV and read by the next one as SRV) they'll happily run in parallel. No need for manual shenanigans with command queues.
- sharpneli 6y ago> It does, but you can always disassemble the DXBC and see what happened to your HLSL code. And you can always disasm X86 code and see what -ffast-math did. Doesn't mean that everyone would be fine with just mandating it everywhere with no option to disable it. Even then the DX functional spec gives some leeway. As an example if you write x*y+z it will compile it into mad instruction. And that's just specified as that the precision must not be worse as the worst possible ordering of separate instructions. So which it is? Depends on the vendor. This is completely fine for graphics, but not fine for all workloads. > No need for manual shenanigans with command queues Unless you do access same buffer from multiple places in a way that's still spec conformant, just in a way that the DX11 implementation cannot detect.
- Const-me 6y ago> Doesn't mean that everyone would be fine with just mandating it everywhere Practically speaking, often I enable it everywhere even on CPU (or similar options in visual C++). When more precision is needed, FP64 is the way to go. Apart from rare edge cases, you won't be getting many useful mantissa bits in these denormals, or with better rounding order. > So which it is? Depends on the vendor. Yeah, but on the same nVidia GPU, I'm pretty sure mad in DXBC does precisely the same thing as fma in CUDA PTX. > just in a way that the DX11 implementation cannot detect It doesn't detect much. If you want to allow shaders to arbitrarily read and write the same buffer, bind that buffer as UAV and you'll be able to run many of these shaders in parallel, despite a single queue. P.S. AFAIK the main use case for these queues is high-end graphics, to send relatively cheap GPU tasks at huge rate (like 1MHz of them), from many CPU cores in parallel. In GPU compute, at least in my experience, the tasks tend to be much larger.