4 ms·
Can’t speak for OP, but for certain low-latency applications like trading, one wants to isolate a core, pin a thread to it, and have that thread spinning hot on
by caffeine 4y ago
Can’t speak for OP, but for certain low-latency applications like trading, one wants to isolate a core, pin a thread to it, and have that thread spinning hot on some condition (maybe network card, maybe IPC), for optimal latency in responding to a particular event.
In that scenario, you wouldn’t want the GC to run on that core, for sure.
- snarfy 4y agoYep it's about latency. Even 1ms can be unacceptable depending on the situation. We currently don't have an option for that other than going native.
- KMag 4y ago>> wouldn’t want the GC to run on that core > We currently don't have an option for that other than going native. You mean going manual memory management, I believe. Technically, JIT vs. ahead-of-time native vs. interpreted is orthogonal to manual memory management vs. GC. Yes, most JITs also have garbage collectors, but it's not necessarily the case. But, your point does stand that when you really care about 99.99th percentile latency, you almost certainly need to go with static native compilation and manual memory management. The long-tail on JIT and/or most GCs is just too high.
- kaba0 4y agoThe OS can cause 1ms pauses easily, so I would be wary whether 1ms is truly unacceptable. To consistently achieve it, not even native may be enough in itself, you have to circumvent the OS kernel as well.
- caffeine 4y agoYou can enforce that the Linux kernel not do this with nohz cmdline options
- KMag 4y ago> In that scenario, you wouldn’t want the GC to run on that core, for sure. In that scenario, you wouldn't want the GC to run at all. If it runs on another core but touches the GC header words on objects used by the application code, it's going to end up stealing cache lines from the application core's L1 cache, resulting in pipeline stalls on the application core. I've talked with folks who just disabled the garbage collector, were careful to keep down the number of objects allocated per trade, and ran their trading engines on boxes with huge amounts of RAM so that they were very unlikely to run out of memory before the close of the market. Now, I could see some GC-specific modifications to the cache coherency hardware. For one, it would be useful to have message to set or unset a single bit on a cache line owned by another core, and report the previous value of the bit on the cross-core interconnect, without changing the ownership of the cache line or moving the whole line across the bus. You'd probably also want a variant on the MESIF cache coherency protocol where another core can get a "notify" bit set on a cache line, so when a cache line in the F state with its notify bit set gets naturally evicted from the cache, its contents and ownership would be transferred to the requesting core. I imagine you'd have a small programmable push-down automaton similar to the ESP32's ULP core sitting in the L1 cache. Its stack would be that core's share of the grey set, and it would have a small array of addresses it was waiting on to mark in the background as their cache lines became un-contended. It would undoubtedly be complex, but it seems the cleanest way to do garbage collection in the background without stealing cache lines from cores executing application code. Really, I think you're best off having Erlang/Elixir/BEAM-like concurrency (but using ahead-of-time native compilation) where most objects are kept in per-actor GC arenas, and maybe only have some types capable of being passed between actors, and allocating those types in separate arenas to keep their cache line contention from harming the happy path non-shared objects. (Yes, last I checked, BEAM had moved from per-actor ("process") GC arenas to shared arenas, but that's largely because they had no way of accurately determining which allocations would later be passed across actors.) Also, ideally, you'd have static inference of region-based garbage collection to reduce the amount of scanning done. Static inference of Rust-like lifetimes is undecidable, so you'd need a conservative inference algorithm that would leave some performance on the table, but you'd still have your garbage collector as a fall back for objects where lifetime inference fails. I guess you'd short-circuit lifetime inference for most dev builds to keep build times sane.