Y
HN Search
Hacker News Search
new
|
comments
|
top
|
jobs
rigtorp
searching PlanetScale…
1.
▲
2.
▲
3.
▲
4.
▲
5.
▲
6.
▲
14 ms
·
31.
▲
by
rigtorp
6y ago
With the right cooling setup I've been able to get Xeons to run permanently in turbo mode, kind of a back door overclock. You would have to experiment. For lowest latency applications I void avoid using RT priorities. Better to run eac
32.
▲
by
rigtorp
6y ago
That would work with a heat pump setup. Liquid nitrogen cooling is usually done by evaporating into the air directly on the processor package (as far as I know). So you would always have condensation inside the server case. Hmm, I guess a i
33.
▲
by
rigtorp
6y ago
Disabling swap only prevents major page faults on anonymous memory. You want to avoid all page faults by using mlockall ( https://linux.die.net/man/2/mlockall ). At that point swap settings doesn't matter. But
34.
▲
by
rigtorp
6y ago
SMT sibling threads can definitely impact each other. It works great for common workloads. If you have a highly tuned workload with high IPC or want to trade off throughput for latency, disabling SMT can be a win. Disabling SMT also increas
35.
▲
by
rigtorp
6y ago
HT share most execution units in the core. If your workload stalls a lot due to branch misprediction or memory access (low IPC) these units can be shared effectively. The Linux perf tool can be used to check IPC.
36.
▲
by
rigtorp
6y ago
You might be interested in https://rigtorp.se/virtual-memory/ where I look deeper at the VM subsystem.
37.
▲
by
rigtorp
6y ago
I don't consider that file backed since they pull memory from the same pool as anonymous memory and not the page cache. The Linux kernel docs makes the distinction between file backed and anonymous memory. I think a better term would p
38.
▲
by
rigtorp
6y ago
Yes and also subambient using heat pumps, but I have never seen it deployed in a data center. How would you deal with condensation?
39.
▲
by
rigtorp
6y ago
I've always been using kernel bypass for low latency networking. You can also use SO_BUSY_POLL with the Linux stack. I should at least mention this in the guide.
40.
▲
by
rigtorp
6y ago
With busy polling you basically halve the SMT sibling thread's memory bandwidth. But yeah it might work well for a specific usecase anyway.
41.
▲
by
rigtorp
6y ago
Most of these tips apply to all architectures. The only x86 specific parts are regarding CPU power management and turboboost.
42.
▲
by
rigtorp
6y ago
RHEL, but all the tools are open source.
43.
▲
by
rigtorp
6y ago
Many organisations run water cooled overclocked servers in production. I have not yet heard of any production use of sub-ambient cooling, but that would be awesome!
44.
▲
by
rigtorp
6y ago
I have another article on virtual memory: https://rigtorp.se/virtual-memory/
45.
▲
by
rigtorp
6y ago
For lowest latency applications I void avoid using RT priorities. Better to run each core 100% with busy waiting and if you do so with RT prio you can prevent the kernel from running tasks such as vmstat leading to lockup issues. Out of the
46.
▲
by
rigtorp
6y ago
This guide pretty much tells you how to make the Linux kernel interfere as little as possible with your application. How to instrument and what to measure would depend on the application. I agree that measuring queuing delay and processing
47.
▲
by
rigtorp
6y ago
For a truly lowest latency in software application you need to avoid all context switches. Using interrupt driven IO adds to much overhead. You need to use polling and busy waiting. I'm working on a guide for this type of application d
48.
▲
by
rigtorp
6y ago
Yes, definitely turn off HT/SMT and use a single app thread per core with busy waiting. I'm working on a low latency application design guide exploring this more in depth.
49.
▲
by
rigtorp
6y ago
Yeah, split locks have a huge performance impact, especially with high core counts: https://rigtorp.se/split-locks/ Fortunately the Linux kernel can kill offenders with SIGBUS: https://www.phoronix.com/