2 ms·
> http://goto.ucsd.edu/~andrysco/errol/ http://goto.ucsd.edu/~andrysco/errol/ seems to be the artifact in question. The artifact reveals that authors of the pa
by mraleph 11y ago
> http://goto.ucsd.edu/~andrysco/errol/ http://goto.ucsd.edu/~andrysco/errol/ seems to be the artifact in question.
The artifact reveals that authors of the paper benchmarked Grisu3 in the debug mode (low compiler optimization level, ASSERTions enabled).
Changing their scripts to actually build Grisu3 in release mode turns tables:
==== Absolute Results ====
Errol3 1196 cycles
Grisu3 802 cycles
Dragon4 6176 cycles
Grisu3 w/fallback 833 cycles
==== Relative Speedup of Errol ====
Grisu3 0.67x
Dragon4 5.16x
Grisu3 w/fallback 0.70x
(these are results on my machine with both Grisu3 and Errol3 compiled with -O3, original scripts build Errol with -O2 which makes it yet more slower than Grisu3).
So ultimately it seems that Grisu3 is still faster.
I also wanted to try benchmarking it on ARM - but it actually fails to build due to its dependency on __uint128_t which does not seem to be supported by my cross-compilation tool chain (at least for 32bit ARMs).