4 ms·
In all fairness it needs to be said that the libc's implementation has to consider portability to more "exotic" architectures. For example, not every CPU allows
by nynyny7 5y ago
In all fairness it needs to be said that the libc's implementation has to consider portability to more "exotic" architectures. For example, not every CPU allows to make unaligned 32-bit or 64-bit writes, or it takes a huge penalty for such writes.
- gpderetta 5y agoglibc has hand crafted assembler implementations of memcpy (often specialized for specific size ranges) for many architectures.
- asdfasgasdgasdg 5y agoDoes glibc not have feature detection and conditional compilation for cases like this? That is surprising to me.
- rwmj 5y agoIt does. Each subdirectory of sysdeps/ can contain specific implementations per platform, arch, etc. eg: the aarch64 assembler memset is: https://sourceware.org/git/?p=glibc.git;a=blob;f=sysdeps/aarch64/memset.S https://sourceware.org/git/?p=glibc.git;a=blob;f=sysdeps/aar... There's also the "ifunc" mechanism which can be used to make the choice at runtime, eg: https://sourceware.org/git/?p=glibc.git;a=blob;f=sysdeps/aarch64/multiarch/memset.c https://sourceware.org/git/?p=glibc.git;a=blob;f=sysdeps/aar...
- nynyny7 5y agoWhat do I mean by "not every CPU allows to make unaligned 32-bit or 64-bit writes"? Let's test the code (as of commit eac67b6) on a Raspberry Pi 400: pi@rasppi400:~/memset_benchmark $ uname -a Linux rasppi400 5.10.63-v8+ #1459 SMP PREEMPT Wed Oct 6 16:42:49 BST 2021 aarch64 GNU/Linux pi@rasppi400:~/memset_benchmark $ ./bench_memset size, alignment, offset, libc, local 0, 16, 0, 1237452, 834116, 1.483549, 1, 16, 0, 1612697, 945325, 1.705971, 2, 16, 0, 1779538, 945320, 1.882472, 3, 16, 0, 1557081, 945324, 1.647140, 4, 16, 0, 1779527, 889736, 2.000062, 5, 16, 0, 1557103, 1000940, 1.555641, 6, 16, 0, 1779551, 1000944, 1.777873, 7, 16, 0, 1557111, 1000945, 1.555641, 8, 16, 0, 1334654, 889723, 1.500078, Bus error pi@rasppi400:~/memset_benchmark $ gdb ./bench_memset [...] (gdb) run Starting program: /home/pi/memset_benchmark/bench_memset size, alignment, offset, libc, local 0, 16, 0, 1557105, 722928, 2.153887, 1, 16, 0, 1557103, 889797, 1.749953, 2, 16, 0, 1557107, 889849, 1.749855, 3, 16, 0, 1557108, 889759, 1.750033, 4, 16, 0, 1557117, 889789, 1.749985, 5, 16, 0, 1557110, 889745, 1.750063, 6, 16, 0, 1557116, 889754, 1.750052, 7, 16, 0, 1557110, 889758, 1.750038, 8, 16, 0, 1557109, 889803, 1.749948, Program received signal SIGBUS, Bus error. small_memset (n=<optimized out>, c=<optimized out>, s=0x29690) at /home/pi/memset_benchmark/src/lib.c:33 33 *((uint64_t *)last) = val8;
- rwmj 5y agoIs that 32 bit ARM code on 64 bit kernel? I thought ARM (since v6) allows unaligned access, although it might have to be emulated through the kernel which is going to be super-slow. On SPARC you have no choice, align or die!
- nynyny7 5y agoYes it is 32 bit code on a 64 bit kernel. I didn't debug what the instruction is that ultimately causes the bus error. pi@rasppi400:~/memset_benchmark $ uname -a Linux rasppi400 5.10.63-v8+ #1459 SMP PREEMPT Wed Oct 6 16:42:49 BST 2021 aarch64 GNU/Linux pi@rasppi400:~/memset_benchmark $ pi@rasppi400:~/memset_benchmark $ file ./bench_memset ./bench_memset: ELF 32-bit LSB executable, ARM, EABI5 version 1 (SYSV), dynamically linked, interpreter /lib/ld-linux-armhf.so.3, for GNU/Linux 3.2.0, BuildID[sha1]=ebeb69b6cb9664d78c1256a2c862f3d28f11e15e, with debug_info, not stripped
- nynyny7 5y agoPS: It's STRD, which as far as I understand the Arm Architecture Reference Manual always requires word alignment. Program received signal SIGBUS, Bus error. small_memset (n=<optimized out>, c=<optimized out>, s=0x29690) at /home/pi/memset_benchmark/src/lib.c:33 33 *((uint64_t *)last) = val8; 1: x/i $pc => 0x11c8c <local_memset+1560>: strd r0, [r7, #-8] (gdb) info registers r0 0x0 0 r1 0x0 0 r2 0x0 0 r3 0x1475 5237 r4 0x5f5e100 100000000 r5 0x11674 71284 r6 0x29690 169616 r7 0x29699 169625
- dividuum 5y agoCan't look right now, but you might not have benchmarked against libc, but against an optimized version included in Raspbian (https://github.com/simonjhall/copies-and-fills https://github.com/simonjhall/copies-and-fills). I'm not sure if that's still active in the latest Raspberry Pi OS releases.
- dolmen 5y agoIn all fairness we need the fastest memset on every architecture. Whatever the cost of maintenance.