4 ms·
Here is the old archived blog post where x264 team answered in the comments why they do it that way. https://web.archive.org/web/20091223024333/http://x264dev.m
by Glacia 3y ago
Here is the old archived blog post where x264 team answered in the comments why they do it that way.
https://web.archive.org/web/20091223024333/http://x264dev.multimedia.cx/?p=191#comments https://web.archive.org/web/20091223024333/http://x264dev.mu...
- exDM69 3y agoThis is dated to 2009. Probably sound advise back then. Compilers are much much better with SIMD code than they were then. Today you'll have to work real hard to beat LLVM in optimizing basic SIMD code (edit: when given SIMD code as input, see comment below). I happen to know because this "hacker news bro" has been dealing with SIMD code for longer than that.
- kierank 3y agoFFmpeg code by definition is not "basic SIMD code". And it supports numerous other compilers other than LLVM.
- jbk 3y ago> Today you'll have to work real hard to beat LLVM in optimizing basic SIMD code. On dav1d, we see just a 800% increase… I know it’s negligible, but…
- exDM69 3y agoCompared to what? Scalar loopy C code sure. The auto vectorization is not great. But give LLVM some SIMD code as input, and it will be able to optimize it, and it does a great job with register allocation, spill code, instruction scheduling etc. Instruction selection isn't as great and you still need to use intrinsics for specialized instructions. And you get all of this for all CPU architectures and will deal with future microarchitecture changes for free. E.g. more execution ports added by Intel will get used with no code changes on your side. With infinite time you can still do better by hand, but it gets expensive fast, especially if you have several CPU architectures to deal with.
- jbk 3y ago> Scalar loopy C code sure. The auto vectorization is not great. Stop considering people as idiots. People do that because it’s a LOT faster, not just a bit. If you are so able, please show us your results. Dav1d is full open source, fully documented, and with quite simple C code. Show your results.
- Const-me 3y ago> show us your results Not GP but here’s an example where intrinsics outperformed assembly by an order of magnitude: https://news.ycombinator.com/item?id=36624240 https://news.ycombinator.com/item?id=36624240 They were AVX2 SIMD intrinsics versus scalar assembly, but I doubt AVX2 assembly gonna substantially improve performance of my C++. The compiler did a decent job allocating these vector registers and the assembly code is not too bad, not much to improve. It’s interesting how close your 800% to my 1000%. For this reason, I have a suspicion you tested the opposite, naïve C or C++ versus SIMD assembly. Or maybe you have tested automatically vectorized C or C++ code, automatic vectorizers often fail to deliver anything good.
- Glacia 3y agoSo you took asm code that had no SIMD instructions in it, made your own version in c++ with intrinsics and figured out that, yes, SIMD is faster? Realy? I think you're completely missing what are we talking about here.
- Const-me 3y ago> Realy? No, not really. My point is, in modern compilers SSE and AVX intrinsics are usually pretty good, and assembly is not needed anymore even for very performance-sensitive use cases like video codecs or numerical HPC algorithms. I think in the modern world it’s sufficient for developers to be able to read assembly, to understand what compilers are doing to their codes. However, writing assembly is not the best idea anymore. Assembly is unreliable due to OS-specific shenanigans, result in bugs like that one: https://issues.chromium.org/issues/40185629 https://issues.chromium.org/issues/40185629 Assembly complicates builds because inline assembly is not available in all compilers, and for non-inline assembly every project uses a different version: YASM, NASM, MASM, etc.