3 ms·
> Really, setting the interval balances speed of detection/cost of slow detection vs cost of reacting to a momentary interruption. Another option is dynamicall
by schmichael 11mo ago
> Really, setting the interval balances speed of detection/cost of slow detection vs cost of reacting to a momentary interruption.
Another option is dynamically adjusting heartbeat interval based on cluster-size to ensure processing heartbeats has a fixed cost. That's what Nomad does and in my 10 year fuzzy memory heartbeating has never caused resource constraints on the schedulers: https://developer.hashicorp.com/nomad/docs/configuration/server#max_heartbeats_per_second https://developer.hashicorp.com/nomad/docs/configuration/ser... For reference clusters are commonly over 10k nodes and to my knowledge peak between 20k-30k. At least if anyone is running Nomad larger than that I'd love to hear from them!
That being said the default of 50/s is probably too low, and the liveness tradeoff we force on users is probably not articulated clearly enough.
As an off-the-shelf scheduler we can't encode liveness costs for our users unfortunately, but we try to offer the right knobs to adjust it including per-workload parameters for what to do when heartbeats fail: https://developer.hashicorp.com/nomad/docs/job-specification/disconnect https://developer.hashicorp.com/nomad/docs/job-specification...
(Disclaimer: I'm on the Nomad team)