3 ms·
This line from the k8s CronJob docs has always made me nervous about adopting them: > The scheduling is approximate because there are certain circumstances whe
by spiffytech 3y ago
This line from the k8s CronJob docs has always made me nervous about adopting them:
> The scheduling is approximate because there are certain circumstances where two Jobs might be created, or no Job might be created. Kubernetes tries to avoid those situations, but does not completely prevent them.
- ants_everywhere 3y agoThis Stack Overflow [0] answer makes it sound like that's just stating a triviality about distributed systems. For example, if the task that runs your cron is down when your cron is supposed to run, then it won't run. The slack blog says they did some tinkering like preventing nodes from going down at the top of a minute because that's when they think cron jobs are most likely to run. But at scale things are going to break when they break, and you have to weigh the pros and cons of designing the jobs to be robust to failure vs trying to organize failures to correspond to the needs of your jobs. So I think there is space for solutions that make different tradeoffs. But it does seem vastly easier to tune an existing solution that someone else is maintaining than to build your own solution on top of Kafka. [0] https://stackoverflow.com/questions/47691278/why-in-kubernetes-cron-job-two-jobs-might-be-created-or-no-job-might-be-created https://stackoverflow.com/questions/47691278/why-in-kubernet...
- josephg 3y agoMost distributed systems at least promise they do something at least once or at most once. You can often achieve exactly once in practice with a combination of idempotent APIs and a shared database. For a task runner, there are a lot of different behaviours you might want if the system crashes. Maybe the runner should “catch up” after coming back online. That’s easy enough to achieve if you move away from cron and track which tasks have been run in a small data store somewhere.
- marcosdumay 3y agoLooks like they made a cron on top of an eventual-consistent database. Yeah, I'd avoid that too.