4 ms·
Skipping to the good stuff from the patch notes: > Wangyang and Arjan reported a bottleneck in the networking code related to struct dst_entry::__refcnt. Perfo
by chaorace 4y ago
Skipping to the good stuff from the patch notes:
> Wangyang and Arjan reported a bottleneck in the networking code related to struct dst_entry::__refcnt. Performance tanks massively when concurrency on a dst_entry increases.
> This happens when there are a large amount of connections to or from the same IP address. The memtier benchmark when run on the same host as memcached amplifies this massively. But even over real network connections this issue can be observed at an obviously smaller scale (due to the network bandwith limitations in my setup, i.e. 1Gb).
Basically, processes that produce a lot of threads (i.e.: 100+) suffer performance degradation. This impacted network performance by up to 5% on 1 Gb/s connections and showed up to 3.2x performance degradation in synthetic benchmarks.
Independent testing verifies these improvements in the Memcached benchmark. Simulating a more realistic workload using Cockroach DB shows a more modest performance improvement of ~5% in terms of ops/s.
- BenoitP 4y ago> struct dst_entry::__refcnt. Performance tanks massively when concurrency on a dst_entry increases Maybe we could just stop tracking references to these dst_entry. This way there would be no contention when tracking reference counts. So no performance loss. Some would no longer be used, but hey RAM is cheap. And to cleanup memory we'd scan the memory for these references from time to time. Probably beginning by the only way (transitive) references can occur from their roots: the thread stacks. We'd quickly pause threads, scan their stack, unpause them while ordering them to notify us if they would write new content to memory. We'd do that at the compiler level so this can be general to all memory management. This way we'd have a concurrent scanning. We could also copy memory as we scan it, removing the gaps and compacting it for better CPU cache usage. But that would require a memory ordering model, threading, and compiler behavior to go hand-in-hand. We'd call it an abstract machine, dare I say virtual? This system would help us automatically collect these memory objects that have become garbage...
- secondcoming 4y agoYou want a device driver written in Java?
- Cthulhu_ 4y agoI'm 99% confident the post is intended to be sarcastically suggesting memory management is replaced with a garbage collector.
- msla 4y agoI mean, congratulations on pointing out why garbage collection is not acceptable in most parts of an OS kernel, especially a general-purpose one as opposed to one only designed to run on specialized high-end hardware, but was that really something that needed to be said?
- otabdeveloper4 4y ago> We'd quickly pause threads Yeah, no. Not a chance.