5 ms·
The author uses a systemd timer to schedule their backups. For backups going to a remote host I prefer adding a little bit of variance to the execution time to
by TimWolla 6y ago
The author uses a systemd timer to schedule their backups. For backups going to a remote host I prefer adding a little bit of variance to the execution time to avoid consistently hitting some hotspot.
From the timer I use to backup my server using Borg to rsync.net:
[Timer]
OnUnitActiveSec=24h
RandomizedDelaySec=1h
This will run the backup script every 24 hours with a random delay of up to 1 hour, so every 24.5 hours on average. This causes the job to nicely rotate around the day.
- corytheboyd 6y agoThat’s really such a nice solution to the problem, nice. Can you imagine not reading the docs to discover those options. So you spin up a database to save state about runs to implement the delay. And you need a dashboard to monitor the various parts of the system for debugging. Or you read the docs
- coldtea 6y agoOr you prefix your command / script with sleep and a randomlly generated 1h value!
- TimWolla 6y agoThis works, but compared to using systemd it has the drawback that the range of possible times is anchored to the configured time in cron. The systemd timer example I gave causes the next cycle to start when the previous job finished. So if it initially runs the script between 0:00 and 1:00 and the script takes 1.5 hours to finish then the next run will be between 1:30 and 2:30 the next day, instead of 0:00 and 1:00.
- MayeulC 6y ago> So if it initially runs the script between 0:00 and 1:00 and the script takes 1.5 hours to finish then the next run will be between 1:30 and 2:30 the next day Shouldn't it be between 1:30 ans 3:30? I'm just nitpicking, of course, that's a nice solution.
- TimWolla 6y agoYes, you are correct. The script ends no later than 2:30 and then there's a delay of up to 1 hour.
- gerdesj 6y agoWhenever I use a scheduler I always use prime numbers wherever possible.
- 0xdeadb00f 6y agoMay I ask why?
- sk5t 6y agoTo appease the cicadas.
- dimtion 6y agoNot OP, but one good reason to do that is to reduce probability and frequency of collision with other periodically running jobs. Let's say you run your job on the hour, and you have a job running every 4h and another every 24h, without planning, because 4 divides 24, you have one in four chance of having them collide and having the 24h job running at the same time than the 4h job. If you add more 4h jobs, the probability that one of those 4h jobs collide each time with the 24h job increases. The more jobs you have, the higher the probability that some will be divisors of others. Using prime number for scheduling reduce the probability at a given time that those jobs collides. If you create a job every 5h and another every 23h, the 23h will collide with the 5h job every 115h. PPCM(5, 23) = 115. Interestingly, this technique is used in nature by cicadas who developed long, prime-numbered, periodical life cycles to avoid gaining a predator that can sync up with the cicadas[0]. [0] https://www.cicadamania.com/cicadas/cicadas-and-prime-numbers/ https://www.cicadamania.com/cicadas/cicadas-and-prime-number...
- MayeulC 6y agoAnd the number of teeth on two interacting gears is usually coprime, to distribute wear more evenly.
- cbhl 6y agoIf you have a bunch of computers each on a fixed timer, and their clocks are synchronized (say, with NTP), then on the least common multiple of all of those timers you'll get a stampede of requests from all of the computers. If you're on a sufficiently large network, that surge can cause failures. And a fixed retry policy will just cause the same stampede to recur on the retry intervals; you want to add jitter to ensure that you spread the load out.
- ComputerGuru 6y agoA loner straightforward solution - albeit not supported by schedulers like cron with start times denoted by fixed, absolute values - is to use non-recurring intervals. Eg run a task at intervals of 86413 seconds
- CameronNemo 6y agoThe snooze scheduler* has a random delay feature. * https://github.com/leahneukirchen/snooze https://github.com/leahneukirchen/snooze
- klodolph 6y agoIt's why you see a bunch of cron jobs that start off with a random sleep. For example, certbot's cron: perl -e 'sleep int(rand(43200))' && certbot -q renew
- seanwilson 6y agoSeems a bad idea in terms of the accuracy of your logs e.g. so you might not notice if the command you want to run is starting to run unusually long some days because of some error.
- shirakawasuna 6y agoWhat's the downside of consistently hitting a hotspot?
- pathseeker 6y agoUnless you have a whole fleet doing the same thing, nothing.
- TimWolla 6y agoIt possibly makes the process unnecessarily slow. People tend to choose “round” numbers for their cronjobs. Probably most commonly minute 0 of a given hour for an hourly or daily job. Thus on e.g. 0:00 UTC there might be hundreds of clients running their backups. I don't have a strict need to run my backups at a fixed point in time (e.g. within the night hours). By not hitting a hotspot I have a better chance of having a larger percentage of the targets bandwidth for my needs (both network as well as disk IO). The random delay ensures that the job runs at a different point in time every day, with most of these points in time being expected to have a light load. If it accidentally hits a hotspot on one day it will be fine the next.
- andyfleming 6y agoI would guess that most workloads have some sort of slow time when it’s appropriate to schedule backups. Wouldn’t having backups rotate through the day potentially cause slowness during a more active time for users?