19 ms·
server health is necessarily a function of the actual production traffic it receives, determined at the application layer, as observed by a specific observer i
by preseinger 3y ago
server health is necessarily a function of the actual production traffic it receives, determined at the application layer, as observed by a specific observer
it can't be known by the server itself, as (among many other reasons) the server can't know about network issues between itself and any upstream caller
it can't be determined by out-of-band health check queries, because those queries don't represent actual traffic, the simplifying assumption that they _do_ introduces many common failure modes that any seasoned engineer can speak at length about
health checks can be a nice additional signal on top of monitoring actual prod traffic, but they can't be used by themselves, they just don't capture enough relevant information