5 ms·
Man, when I see this stuff I sure hope there is maturation of auto-vectorization at the compiler level in clang etc. Even more useful would be compiler-level f
by hackcrafter 10y ago
Man, when I see this stuff I sure hope there is maturation of auto-vectorization at the compiler level in clang etc.
Even more useful would be compiler-level feedback of how to stay within the constraints needed to auto-vectorize your C/C++ for loop. (I need to make this data access const etc)
As far as I know, the Intel compiler is ahead of MSVC/clang on this front without reverting to OpenMP or other annotations on your code.
- revelation 10y agoThe changes necessary to make efficient use of these instructions go well beyond just inner loops (e.g. struct of arrays vs. array of structs) that the compiler can't transparently perform for you.
- wyldfire 10y ago> Even more useful would be compiler-level feedback of how to stay within the constraints needed to auto-vectorize your C/C++ for loop. (I need to make this data access const etc) Specifically for clang you can use the new opt-viewer [1] to annotate the source to find hints on what could/couldn't be optimized/vectorized. Requires compiler flags only found in 4.x/trunk, unfortunately. [1] https://github.com/llvm-mirror/llvm/tree/master/utils/opt-viewer https://github.com/llvm-mirror/llvm/tree/master/utils/opt-vi...
- hackcrafter 10y agoThat's really neat, is there any documentation/use-case examples for this outside the source code?
- wyldfire 10y agoIt was presented at the dev conf but I don't think the slides or video are up yet. :( I've used it with CPython by adding `CFLAGS="-fsave-optimization-record -fdiagnostics-show-hotness"` to the env for `./configure`. Then run opt-viewer like so: ./opt-viewer.py $(find build/ -name '*.yaml') -source-dir src/ output/ If that's not clear I can whip up a quick demonstration project.
- mpweiher 10y agoYes, semi-automatic is definitely a better idea than fully automatic. See "The Death of Optimizing Compilers" [1][2][3]. [1] https://cr.yp.to/talks/2015.04.16/slides-djb-20150416-a4.pdf https://cr.yp.to/talks/2015.04.16/slides-djb-20150416-a4.pdf [2] https://news.ycombinator.com/item?id=9396950 https://news.ycombinator.com/item?id=9396950 [3] https://news.ycombinator.com/item?id=9202858 https://news.ycombinator.com/item?id=9202858
- DannyBee 10y agoHe's quite wrong. In fact, he quotes me. I'm the one who said plenty of people have programs with flat profiles. He simply asserts "this view is obsolete", with no data. I pointed out to him that my data was based on the tens/hundreds of thousands of different programs running in google datacenters for many years. The truth is, for larger customers who actually care about optimizing, profiles are pretty flat. Because if they weren't, you'd hire people, switch languages, switch to gpus, make asics, whatever, until they were. And if you didn't, you clearly don't care that much because it's not worth it, monetarily (IE not "important"). Which is fine, but it in fact, goes very against his point, because those people want optimizing compilers that do a great job without having to spend the money/time to mess with it by hand. He believes you hand these things to algorithm designers, but he misses that this just makes them flat again! As far as i can tell, he's just handwaving. I've talked with numerous colleagues elsewhere, and they see the same. I have yet to hear of a lot of stories of incredibly important code that can't possibly be made faster, where people don't want compilers to do it. Because it would just be dumb to have that code like that. There may be code where, temporarily, it has to be like that while another solution becomes more cost effective, that happens for sure, but it's usually 2-3 year time periods. Because otherwise, if my business could make more money/save time/etc by shoving that code on an FPGA, by convincing intel to add instructions, by funding compiler work, or whatever, not doing it is just dumb. So they do it. And then everything is flat again. The only time people are doing those things by hand is for very short time periods. In the end, they want and need the optimizing compiler to catch up and be able to optimize it well for the architectures. Even on GPU's, for example, people now depend heavily on optimizing compilers producing good code, even though, a few years ago, literally everything was done by hand. Will there come a time where things can't be made faster, and we really do have hotspots again? Maybe. We may change computing models so that doesn't happen :)
- jcranmer 10y agoIntel is abandoning their compiler infrastructure and moving everything to Clang/LLVM. This does mean pushing their autovectorization work into LLVM, although judging from the quality of conversation in the vectorization BoF at the latest developer's meeting, it's not clear how much work they wish to put in to actually making acceptable upstreamable patches.
- yvdriess 10y agoDo you have any citations to back this up? I am a user and not seeing any of this taking place.
- robinhoodexe 10y agoIt sure would be nice if MKL could be integrated in clang. It's still the fastest LAPACK implementation. I use it on a relatively powerfull cluster (~11k cores) at university for doing quantum chemical calculations in Dalton, and having more research software as open source would benefit everyone in the end I believe.
- hackcrafter 10y agoHere here! I run into this often, it is amazing the speedup MKL provides for linear algebra heavy scientific computations and it is under this weirdly commercial, but permissively redistrubutable framework which means a lot of folks are using it unkowningly in a gray licensing area. See the pre-build python Numpy/Scipy packages that use it and are often use by data science types: http://www.lfd.uci.edu/~gohlke/pythonlibs/#numpy http://www.lfd.uci.edu/~gohlke/pythonlibs/#numpy
- infinite8s 10y agoNot sure about Chris's licensing of MKL, but Continuum has a redistributable license of MKL packaged into their numpy builds for Anaconda - https://www.continuum.io/blog/developer-blog/anaconda-25-release-now-mkl-optimizations https://www.continuum.io/blog/developer-blog/anaconda-25-rel...
- 10y ago
- StephanTLavavej 10y agoMSVC provides that feedback. See https://msdn.microsoft.com/en-us/library/jj658585.aspx https://msdn.microsoft.com/en-us/library/jj658585.aspx - the compiler option you want is /Qvec-report:2 which "Outputs an informational message for loops that are vectorized and for loops that are not vectorized, together with a reason code."
- jxy 10y agoFor effective use of such vector instructions, languages like C/C++ really fail at giving enough hints to their compiler. An advanced inspection tool in clang/gcc would certainly help humans to write compiler friendly code, but the real advance can only be taken with an improved programming language that designed specifically for such use. Perhaps the HN crowd is more knowledgable than me, but I fail to recognize any potentially useful language on the market to date. Perhaps Haskell or OCaml with compilers helped by advanced AI?