7 ms·
Hello. I'm a programmer. I noticed you couldn't optimize my code to use SIMD so I went ahead and used inline assembly. It will probably take another 30 years b
by qompiler 14y ago
Hello. I'm a programmer.
I noticed you couldn't optimize my code to use SIMD so I went ahead and used inline assembly. It will probably take another 30 years before you can actually think like a human and perform optimizations like this.
- Scaevolus 14y ago30 years? ICC already does automatic vectorization pretty well. LLVM and GCC have implementations that need more tuning. I bet they'll be solid within 5 years. You still beat them with inline assembly, but you probably won't accelerate the code vectorwidth times anymore.
- raverbashing 14y agoExactly ICC is the best one, but GCC from 4.0 could do it "automatically" (being very loose about the term) And no one does inline assembly for that, they use intrinsics
- cwzwarich 14y ago> And no one does inline assembly for that, they use intrinsics Depends on your architecture. With ARM, compilers will often not set the alignment bits in vector loads and stores, and this can be a big performance hit depending on the microarchitecture. In general, compilers sometimes deal poorly with instructions that have particularly strange register constraints, or loads/stores with address writeback.
- raverbashing 14y agoGood point
- pjmlp 14y agoCommercial compiler vendors tend to invest more in code generation quality as open source developers. After all they need to provide reasons why you would buy them.
- potkor 14y agoThis paper from 2011 compares autovectorization in GCC against two proprietary compilers, from Intel and IBM: http://polaris.cs.uiuc.edu/~garzaran/doc/pact11.pdf http://polaris.cs.uiuc.edu/~garzaran/doc/pact11.pdf GCC got a 23% speedup vs 15% for XLC and 28% for ICC. GCC did well considering the proprietary compilers' narrow focus and corporate resources. BTW, Fortran compilers for the 1980's supers did autovectorization pretty well. Funny that we've made it so hard today.
- exDM69 14y ago> 30 years? ICC already does automatic vectorization pretty well. LLVM and GCC have implementations that need more tuning. I bet they'll be solid within 5 years. Unfortunately, the automatic vectorization compilers do is not very smart and is targeted at simple scalar loops in existing code bases. In many cases, a decent programmer can easily outperform the compiler with simple SIMD tricks, especially in cases where you can use clever vector shuffling tricks (e.g. sum of 4 numbers with 2 additions + 2 vector shuffles). To get the best of both worlds, use compiler/machine specific intrinsic functions or vector extensions with a smart optimizing compiler. Clang's vector extensions work really well.
- raverbashing 14y agoHow would you sum 4 numbers with 2 adds + 2 shuffles (sorry, I'm rusty in SIMD) But remember, similar (old, but simpler) tricks were adopted by compilers, like xor to load 0, lea to do math (not sure about this one), shifts instead of multiply/divide etc
- exDM69 14y agoSum of elements in a 4-vector with SIMD in OpenGL/OpenCL style vector shuffle syntax: vec4 s1 = v + v.yxwz; // s1 = (v.x+v.y, v.y+v.x, v.z+v.w, v.w+v.z) vec4 sum = s1 + s1.zwxy; // sum = sum of v More practical examples, related to 3d graphics: https://github.com/rikusalminen/threedee-simd https://github.com/rikusalminen/threedee-simd
- nwmcsween 14y agoGcc does not vectorize to simd ins without more work than just using intrins. Gcc will vectorize to word size though
- lucian1900 14y agoActually, most modern compilers can vectorize for you. They are limited by imperative idioms of course, but much of the time they do a good job.
- anonymouz 14y agoIgnoring for a moment that compilers are getting better at automatic vectorization, using a compiler allows you to write highly optimized inline assembly for the (usually) tiny parts of the code where it really matters, and gives you the convenience of a high level language everywhere else.
- deleted 14y ago[deleted]
- friendly_chap 14y agoHello, I am your project manager, I thought you were working on the project XYZ, please come into my office.
- MrScruff 14y agoYou obviously don't work in games...
- friendly_chap 14y agoObviously not everyone works in games.
- bad_user 14y agoIt takes a really good developer with vast knowledge to do optimizations such as using "inline assembly" for stuff like SIMD. And even though there are enough good developers able to do this, the mother of all problems when developing software is managing complexity. Yes, you can take a subroutine and apply local optimizations on it. Building complex software in assembly that on the whole is better optimized than what a compiler can do is next to impossible. Speaking of SIMD and stuff like it, there are already optimizations that LLVM is doing, but such optimizations are hard to apply ahead of time because (1) if you want to distribute those binaries easily, then you need to compile for the common denominator (which is less of an issue with LLVM) and (2) your programming language sucks. It's not the compiler's fault, but rather your own fault that you're using a programming language so confusing that inferring intent from your code is next to impossible. How can the compiler know that you're sorting freaking numbers if you're specifying exactly how bits move around in memory while doing so? If you're speaking about virtual machines though, there are projects out there for .NET or Scala for instance that can recompile/retarget code at runtime to run with SIMD instructions or on your GPUs if you have any. All you need is a virtual machine that runs bytecode and a programming language (slash compiler) that lets you access at runtime the syntax trees of the routines that you want to optimize and that lets you generate new bytecode. So you can easily shove this kind of optimizations in libraries for special-purpose and descriptive DSLs (e.g. LINQ). Of course, it gets tricky and doing stuff like this at runtime has overhead, but it's better than what 99.99% of developers can do, not to mention that good developers first and foremost ship.
- eru 14y ago> It's not the compiler's fault, but rather your own fault that you're using a programming language so confusing that inferring intent from your code is next to impossible. And that's the long-running argument for why high-level languages have the possibility to be compiled to faster code eventually. We are mostly still waiting for the languages and compilers. (Even though ghc gives us hope.)
- pjmlp 14y agoHello I am a processor. I have picked up your clever optimized SIMD and decided it was better to change the execution order, because another unit was bored without nothing to do. Unfortunately the L2 cache seems to have some issues getting all required data for your instructions due to the way all your threads are manipulating the required addresses from multiple cores.
- qompiler 14y agoI have to admit this is the painful truth. The speed of the vector operations didn't actually improve using SSE3. But that doesn't dismiss the issue with compilers not being able to actually vectorize properly.
- potkor 14y ago> It will probably take another 30 years before you can actually think like a human and perform optimizations like this. ITYM "It will probably take another 30 years before I use a language that tells you what the required semantics of the code are, meanwhile I'll use my requirements interpretation privileges to relax the semantics in a few places"
- exDM69 14y agoHello, this is the compiler calling back at you. I am very sorry that I cannot write efficient SIMD code but neither can you. If we can work together, we will outperform our individual selves as a team. So you go ahead and write clever SIMD code but please use the SIMD intrinsics I can understand, not inline assembler which I cannot do anything with. You are very good in expressing algorithms in a SIMD friendly way and do a decent job in instruction selection. You, however, are not very good in doing instruction scheduling and register allocation, so let me handle that and we can achieve a result that can keep the CPU pipelines busy. When you make the tiniest change to the program, I can re-do instruction scheduling and register allocation in an instant, when that would take you hours to rewrite the whole algorithm to use different registers. my point: SIMD intrinsics + a smart C compiler produces a lot better code than a programmer writing assembly. Clang + vector extensions in particular is very good at it.
- loup-vaillant 14y agoReminds me that the best current chess player is actually a team mixing humans and computers. Even in that domain where computers are undoubtedly better than humans, the best solution still involve humans.
- EvilLook 14y agoWho else is going to drive the computer to the event and set it up until Google get their why and we all have self-driving cars?
- Jabbles 14y agoThat sounds interesting. Do you have a paper on that?
- loup-vaillant 14y agoHere is an entry point: https://en.wikipedia.org/wiki/Advanced_Chess https://en.wikipedia.org/wiki/Advanced_Chess I don't recall any specific paper, but I think tournament were played.
- dbecker 14y agoHello. I'm another programmer. I noticed that it takes forever to understand inline assembly. I don't foresee myself ever thinking like a computer and reading assembly as well as a high-level language.