5 ms·
The 387 instructions for elementary functions do not have the reputation of having a good accuracy/time ratio compared to, say, state-of-the-art implementations
by pascal_cuoq 12y ago
The 387 instructions for elementary functions do not have the reputation of having a good accuracy/time ratio compared to, say, state-of-the-art implementations using polynomial approximations.
Do you have actual measurements that reveal some actual irony, or is it only superficial irony in that expl() is computed using the (microcoded, slow) hardware instruction and exp() isn't?
Note: For most functions, you can actually make a pretty accurate double-precision implementation by calling the corresponding 387 instruction and then rounding the result to double-precision. Not a correctly rounded implementation (that would require a wider intermediate result than 80-bit) but pretty good nonetheless.
- derf_ 12y agoWe didn't measure average case performance, but for worst-case performance the "orders of magnitude" comment is absolutely backed by measurement. See also my comment at https://news.ycombinator.com/item?id=8830667 https://news.ycombinator.com/item?id=8830667 In our case, we only needed about 12 bits of accuracy anyway, so we wound up writing our own implementation using polynomial approximations (which have the benefit of predictable running times regardless of the quality of your C library). The fact that you need to do this in any context where worst-case performance is a concern might be surprising to some people. It was to me before I read glibc's code.
- derf_ 12y agoI take it back, we did measure average case performance of expf() (though not expl()), and it was slower than exp(): 16:12:48 < gmaxwell> ha. So, completely eliminating all the doubles made celt slower. 16:13:23 < gmaxwell> I'm guessing that this is because I used the C99 float versions of the math functions and they suck. 16:13:57 < gmaxwell> looks like 3% slower or so. 16:14:05 < gmaxwell> <3 GLIBC. 16:26:02 < gmaxwell> with float approx the no-doubles is .9% faster. Whoohoo. ;) 17:24:06 < derf> Yes. Don't use the float versions of libc functions, at least not on x86-64. 17:24:24 < Dark_Shikari> what the heck did they do ? 17:24:53 < derf> They reset the rounding modes on entrance and exit of expf(), for example, regardless of whether or not they're already correctly set up. 17:25:07 < derf> Apparently this is really slow. 17:25:32 < gmaxwell> derf: I was hoping that they'd be fixed by now. "Float approx" refers to our custom polynomial approximation of exp(). Given that, I think it's safe to infer that expl() would also have been slower than exp() in the average case. That was in 2010, though. This appears to have been fixed for expf() in 2012 (which now uses SSE/AVX instead of x87 asm on x86-64), but not for expl().
- siddhesh 12y ago> 17:24:53 < derf> They reset the rounding modes on entrance and exit of expf(), for example, regardless of whether or not they're already correctly set up. I fixed this around the same time I worked on the mp improvements: commit 2506109403de69bd454de27835d42e6eb6ec3abc Author: Siddhesh Poyarekar <siddhesh@redhat.com> Date: Wed Jun 12 10:36:48 2013 +0530 Set/restore rounding mode only when needed The most common use case of math functions is with default rounding mode, i.e. rounding to nearest. Setting and restoring rounding mode is an unnecessary overhead for this, so I've added support for a context, which does the set/restore only if the FP status needs a change. The code is written such that only x86 uses these. Other architectures should be unaffected by it, but would definitely benefit if the set/restore has as much overhead relative to the rest of the code, as the x86 bits do.
- derf_ 12y agoThat patch doesn't affect the x87 asm in sysdeps/x86_64/fpu/e_expl.S, though, does it? Good to hear it's fixed for some cases, though.
- pascal_cuoq 12y agoWhat I meant is that the assembly instruction is faster only because it does not provide the same level of accuracy as the generic implementation. Everyone's needs are different, but it does seem that it is a little bit early, and perhaps that it will always be too early, for a general-purpose libm to aim for correctly rounded transcendental function. Even the best implementations only work in one rounding mode, leading to the pleasantries discussed in a sibling of this comment, and the worst-case execution time can be a problem for some. It is possible to aim for 0.51 ULP with implementations that still work(-ish) if the rounding mode is not round-to-nearest, and perhaps this is the sweet spot that libms that want to be useful should aim for.