3 ms·
With glibc, you can use -ftls-model=initial-exec (or the corresponding variable attribute) to get offset-based TLS. The offset is variable per program (unlike l
by fweimer 2y ago
With glibc, you can use -ftls-model=initial-exec (or the corresponding variable attribute) to get offset-based TLS. The offset is variable per program (unlike local-exec), but the same for all threads, so it's more efficient. Using too much initial-exec TLS (potential across multiple shared objects) eventually causes dlopen to fail because the TCB cannot be resized. This is not a problem if the shared objects are loaded through dependencies at process start.
If initial-exec TLS does not work due to the dlopen issue, on x86-64 and recent-enough distributions, you can use -mtls-dialect=gnu2 to get a faster variant of __tls_get_addr that requires less register spilling. Unfortunately glibc and GCC originally did not agree on the ABI: https://gcc.gnu.org/bugzilla/show_bug.cgi?id=113874 https://gcc.gnu.org/bugzilla/show_bug.cgi?id=113874 https://sourceware.org/bugzilla/show_bug.cgi?id=31372 https://sourceware.org/bugzilla/show_bug.cgi?id=31372 This has been fixed for RHEL 10 (which switched to -mtls-dialect=gnu2 for x86-64 for the whole distribution, thereby exposing the ABI bug during development). As the ABI was fixed on the glibc side in dynamically-linked code, the change is backportable, but it's a bit involved because the first XSAVE-using change upstream was buggy, if I recall correctly. But the backport is definitely something you could request from your distribution.
Note that there was a previous bug in __tls_get_addr (on all architectures that use it), where the fast path was not always used after dlopen: https://sourceware.org/bugzilla/show_bug.cgi?id=19924 https://sourceware.org/bugzilla/show_bug.cgi?id=19924 This bug introduced way more overhead that just saving registers. I expect that quite a few distributions have backported the fix. This breaks certain interposed mallocs due to a malloc/TLS cyclic dependency, but there is a workaround for that: https://sourceware.org/git/?p=glibc.git;a=commitdiff;h=018f0fc3b818d4d1460a4e2384c24802504b1d20 https://sourceware.org/git/?p=glibc.git;a=commitdiff;h=018f0...
The other issue is just that the C++ TLS-with-constructors design isn't that great. You can work around this in the application by using a plain pointer for TLS access, which starts out as NULL and is initialized after a null check. To free the pointer on thread exit, you can use a separate TLS variable or POSIX thread-specific data (pthread_key_create) to register a destructor, and that will only be accessed on initialized and thread exit.
This sort of question is probably more suited to libc-help: https://sourceware.org/mailman/listinfo/libc-help/ https://sourceware.org/mailman/listinfo/libc-help/
- dzaima 2y ago> Using too much initial-exec TLS (potential across multiple shared objects) eventually causes dlopen to fail because the TCB cannot be resized. This feels like the perfect situation to preallocate a gigabyte or something of virtual memory for extending the TLS, similar to how the stack is. But, testing on my system, looks like the allowed initial-exec TLS size is just ~1700 bytes.
- o11c 2y agoRemember that you pay this cost for every thread, not just the main thread. Frankly, if your name isn't `libGL.so` you shouldn't even try to mix initial-exec with dlopen. Just link your libraries normally dammit!
- dzaima 2y agoBut with virtual memory it wouldn't be much of a cost at all (..on 64-bit systems, that is; things are more sad on 32-bit if one cares about those). Just some kernel-internal data structure configuration to ensure that future memory page allocations don't overlap this one. dlopen is a requirement for importing native libraries in non-compiled languages; and, regardless, I as a library author don't get to choose whether users will avoid using dlopen and so have to assume worst-case.
- loeg 2y agoVirtual memory still costs, you know, something like 0.2% of virtual memory space in page table entries. 1 GB of VMA per thread is 2MB of real RAM cost per thread. And there's absolutely no need for that kind of space use -- the thread-local variable can just be a pointer to a heap-allocated large object.
- Dylan16807 2y agoIn addition to the ways that page table entries can be avoided, the system can use large pages for all the areas you aren't using yet, cutting the overhead to 4KB.