5 ms·
A few things worth reading on the topic: Why the C++ committee didn't go with 'Stackful coroutines AKA Fibers' http://www.open-std.org/JTC1/SC22/WG21/docs/pape
by leeter 6y ago
A few things worth reading on the topic:
Why the C++ committee didn't go with 'Stackful coroutines AKA Fibers'
http://www.open-std.org/JTC1/SC22/WG21/docs/papers/2018/p1364r0.pdf http://www.open-std.org/JTC1/SC22/WG21/docs/papers/2018/p136...
Raymond Chen's thoughts on where they are still useful (hint they really aren't):
https://devblogs.microsoft.com/oldnewthing/20191011-00/?p=102989 https://devblogs.microsoft.com/oldnewthing/20191011-00/?p=10...
- jlokier 6y agoInteresting first link, thanks. Though I'm not impressed with the thread-local storage errors argument, as it's just an argument against low quality implementations. In a good implemention of fibers or green threads, TLS APIs should be replaced by fiber-local storage using the same APIs and fast access method, and FLS doesn't have those errors. Perhaps there's another argument against FLS, but the TLS errors argument should have considered FLS rather than leaving it implied that there's no solution.
- gpderetta 6y agoI think the actual issue with TLS is that compilers will happily hoist TLS address computations across function calls. It they didn't do that, TLS would be much safer (at least as safe as with stackless coroutines). TLS Address computation is pretty much free in non-PIE code on x86-64 linux so normally GCC doesn't do it, but on shared libraries or on other architectures, you can randomly get corruptions. I do not think that FLS (at least by default) is the right solution as it would consume a lot of memory, increase the switch overhead and add complexity for little benefit.
- HexDecOctBin 6y agoWait, are you saying that compilers can cache the pointer to TLS across function calls, even when inlining is disabled? Does this mean that the OP article is wrong when it says to use a separate getter/setter function for each TLS variable (even if we manually disable inlining for those accessor functions)? Also, regarding shared libraries, how can compiler cache the TLS address, since the compiler versions that compiled the shared library may be different that the version which is compiling the executable that links to it, and thus may have different caching behaviours?
- gpderetta 6y agothread_local x; void foo() { x = 10; // x in thread 1 foo(); // moves this coroutine to thread 2 x = 20; // might erroneously write to x of thread 1 The issue with dynamic libraries is that TLS computation is slightly more expensive (as the TLS offset is not statically known) so the compiler is more likely to cache the TLS address computation. So the above function might be more likely to "miscompiled" if is part of a .so translation unit. edit: also yes, forcing inlining off is not guaranteed to disable cross procedure optimizations. At the very least you have to disable function cloning. In fact I see that GCC now has attribute noipa which is probably a better solution.
- HexDecOctBin 6y agoYes, but what you describe only happens if you try to access the TLS variables directly. If I put all the TLS variable in a .so file, and only access them through getter/setter functions (also part of the .so), can compiler optimisations still screw things up? It seems to me that unless there is some runtime optimisation going on, this method should be safe.
- gpderetta 6y agoPutting the code to a separate .so is very unlikely to break, unless C++ compilers start doing code generation at dynamic link time. LTO is a thing though, putting it in a separate translation unit doesn't help when static linking.
- jlokier 6y ago> I think the actual issue with TLS is that compilers will happily hoist TLS address computations across function calls. TLS has its own issues, but in every situation where this optimisation is fine with TLS, it's fine with FLS too.¹ If you call a function in a fibers / stackful coroutines environment, and the function has a context switch inside (say it does some blocking I/O), by the time the function returns the original context is valid again. > I do not think that FLS (at least by default) is the right solution as it would consume a lot of memory, Depends on what you store in FLS. As with TLS, there are situations when it's beneficial for memory, compared with storing replicated copies of data all over the place in arguments and data structures, and situations when it isn't beneficial. In some it will consume less memory than stackless coroutines, which store copies of the same data in async structures anyway. > increase the switch overhead If we're still discussing the article's case that stackless coroutines are better, FLS switch overhead is usually lower than stackless/async overhead. In the stackless/async version you have compiler-generated loads and stores to context-specific structs that need to be executed every time something context switches for an await. FLS overhead is updating a single pointer. Even that may be free, because you're going to update some kind of context structure pointer anyway. Even in the base case, stackless will update an equivalent pointer at absolute minimum, but often more. This is really the same argument for TLS being sometimes useful in threaded programs specifically for performance, as opposed to explicitly passing around copies of pointers and storing them in data structures all over the place. (Of course there are non-performance reasons why people choose TLS or avoid it too.) > add complexity for little benefit The benefit is faster programs, and a simpler programming model. You may reasonably disagree :-) But consider, the Linux kernel is basically all fibers with FLS, and that's a top performer :-) With regard to programming model: In Linux, it was found to be very much simpler to use fibers (called "tasks" in Linux) than to build an async I/O state machine for filesystems. This is why Linux didn't get useful async I/O in filesystems for a long time, because the state machine version of filesystem operations, though people attempted, was just far too complex in practice to implement for all corner cases, so its async properties were unreliable. -- ¹ At least no problems that are specific to FLS. If you're pointing out there are issues with TLS (without fibers) due to shared library loading/unloading or lazy memory instantiation, that's an issue but it does not support the article's implication that TLS is ok while FLS is not.