4 ms·
So, they optimized "Round" which took 18 seconds, and by that reduced total cpu time from 53 to 23 seconds. Did I get that right?
by shepik 11y ago
So, they optimized "Round" which took 18 seconds, and by that reduced total cpu time from 53 to 23 seconds. Did I get that right?
- smoyer 11y agoThey also removed two rate-limiters that someone had (presumably intentionally) put into the code. You always have to wonder when you see "protective code" like that - why was it added in the first place? Was it premature optimization? A guess? Often if a weak component is strengthened, you end up with "forgotten" code like this and it's always an effort to find and remove it.
- hongchaodeng 11y agoUnleashing rate limiters is like releasing a water gate -- now you know which part is too weak to stop the flood. By doing this, we can have more insights of system performance and find out which part is too "weak". There is no reason that we couldn't increase the rate limiter if we keep improving the system in overall.
- philips 11y agoYes, essentially. Of course the real work is in creating all of the benchmarks and harnesses to start measuring and investigating; as it always is in these non-trivial systems. The full story summary: First, hit a rate limiter, removing it partially solved the performance issue at low scale. Second, there was a slightly expensive "initialization operation" that should be only done once, made sure it only happened once. Finally, found the expensive round math operation with micro-benchmarks and improving the math operation resulted in doubled throughput.
- deleted 11y ago[deleted]
- hongchaodeng 11y agoThanks for reading it so carefully. Let me explain it one by one: > they optimized "Round" which took 18 seconds This is CPU profile. The overall sampled time is 79 seconds. The Round() method took 18 seconds. This has no direct indication to scheduling latency. > by that reduced total cpu time from 53 to 23 seconds This is about scheduling latency. In this experiment, we schedule 1000 pods on 1000 nodes and see how much time it took to schedule a pod, aka scheduling latency. Before the optimization, the latency is 53 seconds; after, 23 seconds. Of course, this is just one of the steps we kept optimizing.