3 ms·
Driver developers would like for it to be their concern and their concern only, but this doesn't scale when an application developer needs their app to run on h
by oddity 5y ago
Driver developers would like for it to be their concern and their concern only, but this doesn't scale when an application developer needs their app to run on hardware from vendors with different levels of driver quality. Vulkan and DX12 were built with the assumption that application developers would take on more of the burden that the driver developers were doing for them (or not doing, in the case of some vendors). The problem is that hardware vendors will always build a lower level as long as it's economical to do so.
AMD GPUs starting with GCN 1.2 include a microcoded hardware scheduler for virtualizing queues. NVIDIA also has a custom (slightly different functionality) control processor (https://riscv.org/wp-content/uploads/2017/05/Tue1345pm-NVIDIA-Sijstermans.pdf https://riscv.org/wp-content/uploads/2017/05/Tue1345pm-NVIDI...). NVIDIA is certainly the worst of the two for this, but in both cases they're convenient places to shove proprietary secrets while still keeping the driver open source. But even if you see an open source driver on one platform, you have very little guarantee that similar optimizations are running when you use a closed source driver on an entirely different platform.
But, honestly, open source driver/firmware is a low bar. Some of the nastier secrets are in _hardware_. Things like instruction X halving the throughput of a seemingly unrelated shader only on hardware Y might not be obvious unless someone tells you, and they won't tell you in public because it might give away IP. These are the kinds of things you'd hope to find in microbenchmarks, but the search space at the API level is so large that I doubt anyone could come up with a reasonable model on their own. Smaller (or even larger) devs just can't microbenchmark at the scale needed to find these issues.
I don't think hardware vendors are in the wrong for this. Hardware vendors have very understandable motives to guard their IP for competitive and forward-compatibility reasons. Honestly, I think most of the development in graphics APIs has been for the better overall, but that doesn't change that using these APIs is undeniably miserable. With CPUs, we dodge the problem by recommending that people minimize context switches (an IP boundary between different layers of the computing stack) so that most people can assume that most of the code that ran is theirs when they measure things. For the same reasons, I recommend that most people minimize graphics API exposure.
- shmerl 5y agoYeah, from what I've seen some weird edge cases often are fixed after bug reports, not becasue there is some sweeping benchmark that catches them. I suppose GPUs are in general more complicated to address than CPUs in this sense. And it's not always even some intentional secrecy, possibly GPU makers themselves don't anticipate every problem in advance. Driver developers just find some of these issues after some known case exposes them and try to work around that.