11 ms·
FP8 is ~100 tflops faster when the kernel name has "cutlass" in it
- giingyui 1y agoAnd what’s the downside of using that kernel name? It can’t just be that it’s faster and nothing else. Unless they included lots of sleep(x) calls.
- KomoD 1y agoactual link: https://github.com/triton-lang/triton/pull/7298 https://github.com/triton-lang/triton/pull/7298
- bede 1y agoThank you, perhaps the parent can be edited to use this URL instead
- orlp 1y agoGenuineIntel moment.
- hofrogs 1y agoI'm interested in that story, what are you referring to with "GenuineIntel"?
- orlp 1y agoIntel's C++ compiler is known to add branches in its generated code checking if the CPU is "GenuineIntel" and if not use a worse routine: https://en.wikipedia.org/wiki/Intel_C%2B%2B_Compiler#Support_for_non-Intel_processors https://en.wikipedia.org/wiki/Intel_C%2B%2B_Compiler#Support....
- pieterbreed 1y agoIs this for the runtime of the compiled code or for the compiling machine? Do they generate slow code if the compiler is running on non-intel?
- SSLy 1y agothe runtime. patching cpuid makes the code go faster
- kstrauser 1y agoFor the compiled code. Its output deliberately runs slower on non-Intel CPUs.
- Uvix 1y agoRuntime of the compiled code. The ostensible intent is so that new processors can use new features like SIMD, while offering a fallback for older ones. In practice, they’re detecting an Intel processor, not just the specific feature.
- microtonal 1y agoAlso MKL: https://danieldk.eu/Intel-MKL-on-AMD-Zen https://danieldk.eu/Intel-MKL-on-AMD-Zen
- bayindirh 1y agoEven in the middle of that turmoil, we managed to compile some code with Intel's ICC and make it go faster on AMD Opterons, breaking Intel's own numbers. When my colleague said that they managed to go faster than intel with icc with some hand tuned parameters, I remember answering "youdidwat?". Good times.
- reitzensteinm 1y agoOr maybe Quack III: Arena. https://m.slashdot.org/story/21054 https://m.slashdot.org/story/21054
- iforgotpassword 1y agoI think that was the first case (to go public), but I remember reading about this in game magazines a couple times after this, for both ATI and nvidia.
- BoredPositron 1y agoNow I want a Quake shooter but with ducks.
- carlos22 1y agoNot ducks, but chickens, was very popular in Germany back in the day: https://en.wikipedia.org/wiki/Crazy_Chicken https://en.wikipedia.org/wiki/Crazy_Chicken
- avhception 1y agoOh wow, that was a blast from the past. The Moorhuhn craze! Many people, including me, didn't have an internet connection back in the day. The Sneakernet went into overdrive so get everyone a copy!
- supportengineer 1y agoA Duck Hunt, if you will…
- deleted 1y ago[deleted]
- dahauns 1y agoAah, that brings back memories... Interestingly, most benchmark controversies back in the day are now expected behaviour, i.e. game-specific optimizations with no (well, in this age of upscalers and other lossy optimization techniques, probably even somewhat) visible image degradation. A gaming-specific driver with no game-specific improvements in its changelog would be considered strange, and it very much works with executable detection. Back in the day, there was still the argument that drivers should not optimize for benchmarks even when visually identical, because it wouldn't show the hardware's real world potential. Kinda cute from today's perspective. :) But of course there were the obvious cases... The Quack3 lowering filtering quality as shown above, of course (at least that one was put into the driver as a togglable setting later on). But the most cheeky one has to be nVidia's 3dmark03 "optimizations", where they blatantly put static clip planes into the scenes so that everything outside the predefined camera path from the benchmark sequence would simply be cut from the scene early (which e.g. fully broke the freelook patched into 3dmark and would generally break any interactive application)
- _zoltan_ 1y ago[flagged]
- koakuma-chan 1y agois 100 tflops a lot?
- brightmood 1y agoyea
- saagarjha 1y agoIt's like 5-10% here
- irrelative 1y agoCorrect, this is the actual headline too. 100 tflops sure seems like it'd be more than that, but here we are. If the headline was "FB8 is ~7% faster when kernel name has 'cutlass' in it...", it wouldn't seem sensational.
- saagarjha 1y agoI think the interesting part is that it improves performance measurably at all, not the actual number. These people are trying to hit 90+% MFU (though most don't reach it) so this does actually translate to many millions of dollars for them.
- progx 1y ago5060 ti +~15%
- HideousKojima 1y agoAccording to Terminator 3 Skynet used a mere 60 TFLOPS
- IAmBroom 1y agoHow much is that in jiggawatts per parsec?
- nolok 1y agoIntel's quest to move from "trusted by default / the reference" to "check for scam" is getting worse every release. And it's 100% self inflicted. How weird.
- aleph_minus_one 1y agoIn my understanding of the PR, it rather seems that it is NVidia is the company that is cheating. :-)
- pkhuong 1y agoNVIDIA-inflicted in this case.
- hvenev 1y agoIn `libnvidia-nvvm.so` the string `cutlass` appears right after `Memory Dependence Analysis` and `memdep`. Perhaps it acts as an optimization attribute of some sort, where the compiler is allowed to make assumptions about the kernel's behavior that are not valid in general?
- high_na_euv 1y agoThats very likely imo
- jdright 1y agoyes, that is a very usual way (known practices) of vendors applying specific optimizations for known things. It is also part of the benchmarks game they play against each other.
- MichaelZuo 1y agoIt’s really strange for established companies to waste their credibility on games like that…
- MangoToupe 1y agoI was pretty young at the time, but I recall the market for graphics being a lot wider open at the time Quake was released. Remember 3dfx? They produced the Voodoo series of graphics cards. They're barely a distant memory now. Quake was also the standard for a game that was willing to fully exploit the hardware of the time.
- IAmBroom 1y agoNever underestimate how much human ego will control actions.
- MBCook 1y agoThe link is long dead and the Wayback machine doesn’t have a copy. But in 2001 ATI was caught applying optimizations to Quake 3 when someone realized if you renamed the executable from “quake” to “quack” the score dropped a ton. It was a big scandal. I know that’s common now but that wasn’t a thing that was done at the time.
- PLenz 1y agoThe Volkswagon emissions testing model
- fambalamboni 1y ago[dead]
- rowanG077 1y agoLet's hope for Nvidia this is an innocent optimization only valid for internal kernels that cannot be applied in general.
- jagrsw 1y agoIn which case checking for a string inside arbitrary name is sloppy (a bug).
- high_na_euv 1y agoI have small experience with compilers and llvm but youd be shocked how many things rely on names and parsing names If you have hundreds of passes that are complex and rely on various "contracts" like type names or some shit, then really crazy things like this can happen unintentionally and not maliciously
- diggan 1y agoWeb-developers are well aware of this too. Sincerely, Mozilla/5.0 (X11; Linux x86_64; rv:139.0) Gecko/20100101 Firefox/139.0
- bravesoul2 1y agoFunny we send a browser wars tombstone in every request!
- antonvs 1y agoLet's have a moment of silence for Gecko/20100101
- halJordan 1y agoWhy would i be shocked that a name is informative. Like... are you surprised that wrought iron is wrought? Or cast iron is made from a cast?
- IAmBroom 1y agoDog piles are often neither composed of dogs, nor actual piles. Names can be both informative, and misdirecting, at the same time.
- the8472 1y agoSome names are standardized items, like memcpy. Matching those is ok, nothing sneaky going on there. Matching something vendor-specific in a general-purpose API is different story.
- Arch-TK 1y agoI wish people either learned how to use git or just wholesale stopped using it.
- tempaway43563 1y agoSo, what is Cutlass, can someone explain whether checking for kernel names makes sense here or is a form of cheating? https://docs.nvidia.com/cutlass/index.html https://docs.nvidia.com/cutlass/index.html
- rurban 1y agoThat's strange because the cutlass docs explicitly does NOT mention fp8 support. So it looks like it can be used nevertheless with fp8 by using the name hack.
- mlazos 1y agoIt supports e5m2 and e4m3 right in the doc linked.
- gpm 1y agoGithub version: https://github.com/NVIDIA/cutlass https://github.com/NVIDIA/cutlass I wonder if we search the comments if we can find something referencing this.
- zahlman 1y agoThis tweet appears to be taking the original material out of context to misrepresent it: > Rewrite the attention kernel to be persistent. This gives better performance at low-contexts. However, fp16 at large context has suffered a bit due to a ptxas instruction scheduling issue in the softmax partition. fp8 is ~100 tflops faster when the kernel name has "cutlass" in it. The charitable reading is that, on certain kernels, using fp8 rather than fp16 values gives better performance. (Although I can't even see how the numbers relate to a "~100 tflops faster" claim in any respect, nor does it even list any kernel names or suggest a control kernel!) But this is being presented as if someone has uncovered evidence of cheating on benchmarks.
- saagarjha 1y agoI think you're the one doing that to the tweet, actually.
- zahlman 1y agoWhat are you talking about? When I view the tweet, the only text I see is: > > fp8 is 100 tflops faster when the kernel name has "cutlass" in it > kms
- saagarjha 1y agoAnd it includes a link to show that this is the context it came from.
- zahlman 1y agoAnd when I look at the link, the part I quoted is the relevant text I see. In order to get to the part that you're trying to hold me accountable for, I would furthermore have to click onto the commits tab and search through a 93-commit PR. I thought today I was using a site where trying to think the best of people and propose that someone had taken something out of context, based on the immediately available context having a simpler explanation, would not get me treated like a corporate shill (for a company I don't even care about). Apparently I was wrong.
- arzookanak 1y ago[dead]
- spoaceman7777 1y agoSeems this is likely due to ongoing work on FP8 support on nvidia/cutlass. From my reading, the alternative code path was likely added recently for testing by external contributors to the cutlass project, and other involved parties. (Rather than attempting to distribute custom packaged internal builds of cuda.) This ticket is a good starting place to see the chain of issues around the ongoing work: https://github.com/NVIDIA/cutlass/pull/2037 https://github.com/NVIDIA/cutlass/pull/2037
- Xss3 1y agoThe real answer