4 ms·
Yes, I found many different bottlenecks. I was mostly running on AWS. In terms of hardware, for small-packets loadtests, most systems are constrained on throug
by romange 4y ago
Yes, I found many different bottlenecks.
I was mostly running on AWS. In terms of hardware, for small-packets loadtests, most systems are constrained on throughput, i.e. number of packets per second. Some instances saturate on interrupts reaching 100% CPU on all cores and some can not even saturate the CPU and you will see that CPU is at 60% but you can not go beyond in throughput. The best systems network-wise are c6gn family types. They are also better than instances that other cloud provide. btw, you mentioned hypervisors... About 8 months ago I opened a bug on AWS Graviton team https://github.com/amzn/amzn-drivers/issues/195 https://github.com/amzn/amzn-drivers/issues/195 - about performance issue they had on their instances at high throughput. Recently they issued the fix. I suspect it was in their hypervisor.
In terms of my software I found many performance bugs at those speeds. For example, using a default allocator is a big no. I use mimalloc for uncontended allocations. In general, you can not use mutexes and spinlocks at those speeds. Those will just cripple the system. Sometimes it can be very annoying since you can not rely on a 3rd party library without carefully analyzing its design. For example, I could not use openmetrics c++ library because it was not performant enough. Even to implement a simple counter, say to gather statistics for INFO command becomes an interesting engineering problem:
With share nothing architecture, I use a lot of thread-local counters that I aggregate only when stats are pulled.
As a general note, I expect that Dragonfly will stay very performant with the tailwinds from recent hardware advancements. For example, c7g (Graviton 3) is much better than c6g and DF shows it.
- amluto 4y agoI'm curious: since you're targeting Graviton, and AFAICT Graviton 2 and up support ARM LSE, have you tried directly using LSE for metrics? ARM LSE offers STADD to add to a 64-bit memory location without reading the contents or ordering the access. (I think LDADD with XZR as the destination is identical, but I could be missing some subtlety. I'm not an ARM expert.) You might be able to get good performance by using a global counter and just STADDing to it. x86 will perform terribly if you use XADD because it has no corresponding optimized forms.
- romange 4y agoWow, its beyond my knowledge:) I use my software skills to write a performant software but rarely i go to the assembly level. And i have zero knowledge about specifics if each command on every architecture. The rare exception is a code in dashtable. the authors of the paper designed it to allow vectorizing of the find() operation.
- reflexe 4y agoInteresting. Have you considered/tried to use fdio/user space networking? In my experience, it greatly improves throughout (simple ip forwarding can be more than 10mpps per cortex a53 on some platforms). Fdio also has a so that you can preload in order to use its ip stack in your app (instead of Linux's). See https://s3-docs.fd.io/vpp/22.06/developer/extras/vcl_ldpreload.html https://s3-docs.fd.io/vpp/22.06/developer/extras/vcl_ldprelo...