3 ms·
I often write small bits in C because it's simply fun. And I tend to get bogged down over-abstracting when I'm working in something more modern. The simplicity
by Coding_Cat 11y ago
I often write small bits in C because it's simply fun. And I tend to get bogged down over-abstracting when I'm working in something more modern. The simplicity of C is a boon, allowing me to focus on one small problem and easilly study the performance. (I have spend the last 2 days making a thread-pool and BLAS system in C for fun, initially results are quite good. My 4x4 matrix-matrix multiplication algorithm beats Intel's own example :D).
I wouldn't be quick to use it for a large scale project which does not require 110% performance but for anything were performance is your nr. 1 goal I've found it quite refereshing.
- exDM69 11y ago> My 4x4 matrix-matrix multiplication algorithm beats Intel's own example You know you can't make such statements without putting your money where your mouth is. We want to see the code. Here's three different 4x4 matrix multiplication routines I've written (the mmmul functions). Depends on your cpu which is fastest. https://github.com/rikusalminen/threedee-simd/blob/master/include/threedee/matrix.h https://github.com/rikusalminen/threedee-simd/blob/master/in...
- Coding_Cat 11y agoI'm in a hurry right now, but I'll link to it later. It's on a i7-4750HQ, it was about 10% faster as measured with rdtsc and looping it a couple of million times. Granted, Intel's implementation was for 8x8 (floats), perhaps that makes a difference in the instruction pipelining. I'll see if it does later.
- Coding_Cat 11y agoI'm waiting on a plane, so it's a bit messy, and I haven't tested the float version yet but: https://github.com/RDeckers/c_doodles/blob/master/Linalg/src/linalg.c https://github.com/RDeckers/c_doodles/blob/master/Linalg/src... Compile with gcc -O3. MxM_4x4 is my function, say we have AB=C then it loads B transposed into registers, then for each row it uses 4 multiplications to calculate to row(A)column(B) products, 2 HADD for summing half of the generated products, then 2 permutations to line up the vectors, and one final sum to compute a row of C. MxM_4x4_2 is a direct port to doubles instead of floats of intel's example on the provided link. When compiled with -O3 my compiler produces the same code in terms of assembly instructions as Intel's example explictlly writes. (Note the name of the repo, I know it's not pretty ;) )
- girvo 11y ago> And I tend to get bogged down over-abstracting when I'm working in something more modern. The simplicity of C is a boon, allowing me to focus on one small problem and easilly study the performance Interestingly, this is why I enjoy working with Go.