2 ms·
At first I thought it is a well-written post-mortem with proper root cause analysis. After reading it for the second time though, it doesn't sound like the root
by throwdbaaway 5y ago
At first I thought it is a well-written post-mortem with proper root cause analysis. After reading it for the second time though, it doesn't sound like the root cause has been identified? At one point, they disabled streaming across the board, and the consul cluster started to become sort of stable. Is streaming to be blamed here? Why would streaming, an enhancement over the existing blocking query, which is read-only, end up causing "elevated write latency"? Why did some voter nodes encounter the boltdb freelist issue, while some other voter nodes didn't?
And there is still no satisfying explanation for this:
> The system had worked well with streaming at this level for a day before the incident started, so it wasn’t initially clear why it’s performance had changed.
But I totally agree with you that the first thing they should look into is to rollback the 2 changes made to the traffic routing service the day before, as soon as they discovered that the consul cluster had become unhealthy.