4 ms·
"FP32 is less common in modern ML workloads and often less optimized on recent hardware compared to FP16 or BF16, which may partly explain why it’s easier to ac
by ekelsen 1y ago
"FP32 is less common in modern ML workloads and often less optimized on recent hardware compared to FP16 or BF16, which may partly explain why it’s easier to achieve performance gains over PyTorch with FP32 kernels."
People haven't spent time optimizing the fp32 versions of these kernels in years. This will be much more interesting if they can improve the kernels where developer effort has gone and that are actually used.
- suddenlybananas 1y agoI wonder if it's using known improvements from the fp16/bf16 kernels that are transferable to fp32?
- moralestapia 1y ago>People haven't spent time optimizing the fp32 versions of these kernels in years. Wow, so, you're basically saying the AI created new algos in a domain with no pre-existing solutions? Awesome!
- Aurornis 1y agoNo one said the AI created new algorithms nor that there weren’t pre-existing solutions. The implication was that the FP32 versions of these kernels have lagged behind the more popular versions. There was opportunity to translate the advancements from other kernels into these. Someone would need to look closely to see exactly what was done, but it’s premature to suggest anything like “new algos” or “no pre-existing solutions” This is a great use case for LLMs, though. I often do something similar where I make improvements to something I use most frequently and ask an LLM to translate that pattern to other similar parts of the code.
- moralestapia 1y ago>The implication was that the FP32 versions of these kernels have lagged behind the more popular versions. Help me understand this 'cause I'm a bit slow these days ... Does that mean optimized FP32 versions of these kernels were already there or not?
- almostgotcaught 1y ago> Help me understand this 'cause I'm a bit slow these days ... If I do `sed 's/f32/f16/g' kernel.cu` does this count as AI? Help me understand because I'm a little slow when it comes to all the dumb shit people attribute to LLMs these days...
- moralestapia 1y agoIndeed, you're slow on these news. >sed 's/f32/f16/g' kernel.cu This is not what's happening here, it's a completely different thing, read TFA.
- imtringued 1y agoYou are a blatant troll and you know that.
- Dylan16807 1y ago> Does that mean optimized FP32 versions of these kernels were already there or not? If you're trying to support your original point with that argument, then you're using some pretty awful definitions of the terms "new algos" and "no pre-existing solutions".
- uoaei 1y agoThe hype cycle in action, folks. Pay heed.
- deleted 1y ago[deleted]
- vlovich123 1y agoThe solution not existing in PyTorch does not mean the solution doesn’t exist elsewhere on the internet. Remember - PyTorch is largely maintained by employees of companies that have their own priorities for the SW and those priorities may not include hyper optimizing fp32 kernels. That being said, it is cool if AI is enabling lower cost adoption of better more optimized kernels with less effort.
- imtringued 1y agoRead the article before spouting lies. Actually never mind that. Read the damn comment you're responding to. There have been human written kernels for both fp16 and fp32 for a long time. Here is the corrected version of your comment: "Wow, so, you're basically saying the AI created the same but faster algos in a well known domain with established pre-existing solutions, whose overall impact on the runtime of practical workloads is insignificant? Awesome!"
- deleted 1y ago[deleted]
- adrian_b 1y agoI believe that these good results are explained at least in part by the fact that NVIDIA does not provide detailed enough documentation for their GPUs. For a processor with well-documented microarchitecture, for which a programmer or a compiler can deterministically write an optimal program, it is much less likely that applying ML/AI can be successful, except as a substitute for searching already known solutions. On the other hand, for less documented microarchitectures, like of the NVIDIA GPUs, finding an optimal program may be impossible other than by doing a random search guided by examples of previous optimized programs, and possibly doing some reverse-engineering work to determine the real behavior of the GPU in some circumstances. Improving over something like this is likely to be feasible for ML/AI, where training over known good programs may be able to extract some of the undocumented behavior that may be non-obvious for humans reading those examples.
- almostgotcaught 1y ago[flagged]
- speerer 1y agoThe point was about being able to write an optimal program with certainty, not about just getting the thing to operate.
- almostgotcaught 1y ago[flagged]
- throwaway81523 1y agoThe running time of a CUDA kernel is apparently impossible to determine except by experiment and measurement, and might be nondeterministic. By contrast for a more typical CPU, there's a compiler whose assembly output you can examine, and there's a processor manual that gives the cycle timing of each instruction. So you can compute the running time at least of inner loops that stay in cache, and that sort of thing.