4 ms·
The way I look at it, the article has two main messages. The first message is that real systems comprise "core" components that are expected to work in all cas
by MatteoFrigo 5y ago
The way I look at it, the article has two main messages.
The first message is that real systems comprise "core" components that are expected to work in all cases (e.g., databases, or CPU arithmetic units and DRAM memory) and "uncore" components that are just performance optimizations (caches, prefetchers, load estimators, whatever) and are expected to be not effective in certain cases. The point is that, while "uncore" components may well provide 100x performance improvement, the designer must make sure that the dynamics of the system are stable even if these components fail completely. A related point is that you cannot infer the dynamics of the core by looking at production metrics alone: you must either have a deep analysis of the system, or a provably robust algorithm that works in all cases, or test cases that explicitly disable caches or whatever other uncore is relevant.
The second message is that real systems have metastable states, which surprises people. "Metastable" means that the system runs fine for a while, then something bad happens, the system goes into a bad state, and stays in the bad state after the bad stimulus is removed. Historically one of the first examples is the infamous Internet congestion collapse of 1986, which prompted the implementation of TCP congestion control. The feedback loop in that case had nothing to do with caches. What happened is that routers drop packets when the rate is too high. Packet drop caused TCP to retransmit, causing more packet drop. This feedback loop is self-sustaining and persists even if you remove the initial overload that caused the initial packet drop. See https://ee.lbl.gov/papers/congavoid.pdf https://ee.lbl.gov/papers/congavoid.pdf for ways to avoid this problem, which are more broadly applicable than just TCP.
Specifically w.r.t. congestion collapse, my experience is that many smart engineers have never heard of it, and I stopped counting how many outages have been caused by congestion-related problems. So there is something to be learned here.
- jacquesm 5y agoThe first is absolutely impossible: the fact that performance changes is in many situations already a problem, even if the change is an improvement because it can lead to security issues. Caching is extremely subtle and due to non-local effect has all kinds of implications that are often poorly understood or only come out after a long time (sometimes decades, for instance in the case of the Spectre type bugs which affected CPUs of many generations). Caching is hard. The benefits are too large to ignore it but the downsides are extensive and can bite you when you least expect it. As for that particular interesting tidbit, see also: thundering herd and exponential back-off.