4 ms·
For the last few months, I’ve been working on a vendor-agnostic inference library. Have 3 backends so far: legacy D3D11 for compatibility, D3D12, and Vulkan 1.3
by Const-me 1mo ago
For the last few months, I’ve been working on a vendor-agnostic inference library. Have 3 backends so far: legacy D3D11 for compatibility, D3D12, and Vulkan 1.3. I have reasons to believe nVidia deliberately crippling Vulkan API for their consumer GPUs. Couple examples to be specific.
nVidia driver sets quite low number for VkPhysicalDeviceLimits::maxTexelBufferElements. I don’t think that’s a hardware limit because I have D3D12 backend doing the same thing on the same hardware.
Vulkan performance is not great on nVidia. On all AMD cards I am testing, Vulkan 1.3 is the fastest backend. On nVidia however, D3D12 is faster despite tensor cores (WMMA / wave matrix multiply accumulate / cooperative matrices) are only available through Vulkan, D3D12 is using shader cores exclusively.