5 ms·
Maybe this will need C based highly optimized backend as well, however, the code looks very clean and beautiful to me. I hope someday differentiable programming
by chunsj 7y ago
Maybe this will need C based highly optimized backend as well, however, the code looks very clean and beautiful to me. I hope someday differentiable programming becomes possible in common lisp.
- darkpuma 7y agoBLAS/LAPACK is fortran, not C.
- stephencanon 7y agoBLAS is a set of subroutine interfaces that originated in Fortran, but can be implemented in any language; most mainstream implementations use one or more of Fortran, C, C++, assembly, and DSLs.
- darkpuma 7y agoIt's my impression that common implementations of BLAS have fortran at the core because fortran (moreso than C) lends itself to automatic vectorization.
- gh02t 7y agoThis is partially true (or at least once was, the performance gap is pretty slim nowadays), but most of the high-performance BLAS implementations like OpenBLAS or MKL have their hot path code written in hand-tuned assembler anyway.
- stephencanon 7y agoBLAS inner loops are usually explicitly hand-vectorized, so Fortran’s autovectorization advantages don’t help. I’ve written many BLAS kernels in assembly over the years.
- gnufx 7y agoNot arguing with that, but I think the jury is out on whether they need to be hand vectorized. With recent GCC, generic C for BLIS' DGEMM gives about 2/3 the performance of the hand-coded version on Haswell, and it may be somewhat pessimized by hand-unrolling rather than letting the compiler do it. The remaining difference is thought to be mainly from prefetching, but I haven't verified that. (Details are scattered in the BLIS issue tracker earlier this year.) For information of anyone who doesn't know about performance of level 3 BLAS: It doesn't come just from vectorization, but also cache use with levels of blocking (and prefetch). See the material under https://github.com/flame/blis https://github.com/flame/blis. Other levels -- not matrix-matrix -- are less amenable to fancy optimization, with lower arithmetic density to pit against memory bandwidth, and BLIS mostly hasn't bothered with them, though OpenBLAS has.
- stephencanon 7y agoHaswell is now a 6 year-old core, so compiler cost models and vectorization support for AVX2 + FMA have mostly caught up. The state of autovectorization targeting AVX-512 or even NEON on new cores is quite a bit less satisfactory for now. It's long been the case that compilers generate adequate vector code for five-year old cores. A large part of the wins for hand-vectorization (and assembly) come from targeting cores that haven't even been released yet, or were just made available. Also, 2/3 is ... significantly worse than I would expect, actually. My experience writing GEMM (I did it professionally for a decade) is that getting to 80% of peak is mechanical, and the remaining 10-15% is where all the real work is. It's also a mistake to ignore L1 and L2 BLAS; for data that fits in the cache hierarchy, these become important, and you can absolutely squeeze out significant gains from hand optimization. For the HPC community, such small problems are uninteresting, but for everyday consumer computing, you can reap huge benefits here.
- gnufx 7y agoI wasn't in a position to make realistic comparisons on SKX at the time, but it seems that Zen is similar to Haswell (and Haswell/Broadwell is still pretty relevant in HPC). I'll eventually get back to it and try on SKX. I expect you can do better with the generic code, as suggested. The point is that received wisdom is that you can't just rely on compiler vectorization but need careful hand-(un?!)rolled code. The story around BLIS was that GCC wouldn't vectorize loops that at least GCC6 actually does (better than the SSE3 code in the how-to-optimize-dgemm tutorial). Then the suggestion was to force it with the OpenMP simd pragma, which is worse on Haswell because it doesn't use fma, but I don't know if that's relevant for, say, POWER using the generic kernels. I'm not actually ignoring level 1 and 2, although my interest is HPC. BLIS' reduction-type loops in generic C at least get vectorized now via -funsafe-math-optimizations. Small dgemm is important in some HPC applications too, for which BLIS doesn't do too well, but there's libxsmm.
- gnufx 7y agoYou'll typically need "restrict"ed arguments for vectorization in C, but then GCC usually does vectorize the loops OK in BLAS-like code. There may be advantages from Fortran arrays; I forget the details.
- stephencanon 7y agoFortran gets you both restrict and re-association allowed by default, plus the nice syntax for vector ops. That's about it, though.
- gnufx 7y agoRight, though a LAPACK implementation is probably going to be mostly Fortran from the reference version, typically with some optimized routines, e.g. what OpenBLAS imports. Something to note is that unfortunately, the interface is the obsolete Fortran77-style, not modern Fortran, though that does make it more likely you can interchange implementations at run time as appropriate.
- jlarocco 7y agoSBCL is actually really good at optimizing this type of code, and I think it can even do some vectorization. If it's not not good enough, there are a few CL bindings and wrappers around BLAS and LAPACK. A good one is Matlisp, which is available in QuickLisp. http://quickdocs.org/matlisp/ http://quickdocs.org/matlisp/
- fiddlerwoaroof 7y agoSbcl has all the pieces in place for vectorization, but I don’t think anyone has actually hooked them up. However the great thing about lisps is that the compiler and optimizer are written in the language they compile, so an end user can do this sort of thing themselves: https://www.pvk.ca/Blog/2013/06/05/fresh-in-sbcl-1-dot-1-8-sse-intrinsics/ https://www.pvk.ca/Blog/2013/06/05/fresh-in-sbcl-1-dot-1-8-s...
- pjmlp 7y agoOptimizing Lisp compilers are quite good. Current generations tends to be unaware that once upon a time C compilers only generated turtle code, it was the investment into optimizing compilers and abuse of UB that made current compilers able to generate fast code.
- jstimpfle 7y agoAnother C hate post? It seems way overblown. When you compare trivial compilation of C code with optimized compilation, you would see maybe a factor of 5 in difference in code size. On today's machines, I think you can expect to get more than 50% of performance from trivial compilation, compared to optimized code. Depending on workload you it can be 80-90%. (I have a small amount of experience, based on my own very shitty compiler, to back up that claim). Even on machines from "once upon a time", where performance was more directly correlated with code size, it couldn't have been that bad, assuming the factor of 5. And also seeing the fact that C was used for systems development from the start. And that's still ignoring (or at least contradicting) the common lore that "once upon a time", LISP was comparatively more inefficient than C than it is today! And I don't believe that abuse of UB is important to generate fast code. After all, fast code doesn't necessarily contain UB.
- pjmlp 7y agoI have written plenty of them since C vs C++ on comp.lang.c were a thing. Once upon a time Common Lisp compilers were already good enough versus Fortran. https://people.eecs.berkeley.edu/~fateman/papers/lispfloat.pdf https://people.eecs.berkeley.edu/~fateman/papers/lispfloat.p... As for C's performance on the old days, "Oh, it was quite a while ago. I kind of stopped when C came out. That was a big blow. We were making so much good progress on optimizations and transformations. We were getting rid of just one nice problem after another. When C came out, at one of the SIGPLAN compiler conferences, there was a debate between Steve Johnson from Bell Labs, who was supporting C, and one of our people, Bill Harrison, who was working on a project that I had at that time supporting automatic optimization...The nubbin of the debate was Steve's defense of not having to build optimizers anymore because the programmer would take care of it. That it was really a programmer's issue.... Seibel: Do you think C is a reasonable language if they had restricted its use to operating-system kernels? Allen: Oh, yeah. That would have been fine. And, in fact, you need to have something like that, something where experts can really fine-tune without big bottlenecks because those are key problems to solve. By 1960, we had a long list of amazing languages: Lisp, APL, Fortran, COBOL, Algol 60. These are higher-level than C. We have seriously regressed, since C developed. C has destroyed our ability to advance the state of the art in automatic optimization, automatic parallelization, automatic mapping of a high-level language to the machine. This is one of the reasons compilers are ... basically not taught much anymore in the colleges and universities." -- Fran Allen interview, Excerpted from: Peter Seibel. Coders at Work: Reflections on the Craft of Programming
- yarrel 7y agoIt's unlikely to need an optimized C backend :-) - http://www.iaeng.org/IJCS/issues_v32/issue_4/IJCS_32_4_19.pdf http://www.iaeng.org/IJCS/issues_v32/issue_4/IJCS_32_4_19.pd...