4 ms·
The system design choice was to make data visible to queries as soon as possible after being pushed to Monarch, to satisfy alerting guarantees. Thus there was
by gttalbot 4y ago
The system design choice was to make data visible to queries as soon as possible after being pushed to Monarch, to satisfy alerting guarantees.
Thus there was no queue like a pubsub or Kafka in front of Monarch.
At scale this required a "smoothness of flow". What I mean by this is that at the scale the system was operating the extent and shape of the latency long tail began to matter. If there are many many many many thousands of RPCs flowing through servers in the intermediate routing layers, any pauses at that layer or at the leaf layer below that extended even a few seconds could cause queueing problems at the routing layer that could impact flows to leaf instances that were not delayed. This would impact quality of collection.
Even something as simple as updating a range map table at the routing layer had to be done carefully to avoid contention during the update so as to not disturb the flow, which in practice could mean updating two copies of the data structure in a manner analogous to a blue green deployment.
At the leaf backends this required decoupling--to make eventual--many ancillary data structure updates for data structures that were consulted in the ingest path, and to eventually get to the point where queries and ingest shared no locks.