8 ms·
Replacing cron jobs with a centralized task scheduler
- _wire_ 1y agoNext thing you know you'll have systemd.
- datadrivenangel 1y agoOr worse, airflow!
- UltraSane 1y agoAirflow can be frustrating but when it works it is so satisfying.
- globular-toast 1y agoI think mistaking Airflow for a mere "task scheduler" is part of that frustration.
- flakes 1y agoAfter using Argo Workflows, I don't think I will ever return to Airflow. Kubernetes is not an easy system to manage, but managing an Airflow setup is somehow worse. The story around disaster recovery and scheduler redundancy was an absolute nightmare for me.
- datadrivenangel 1y agoArgo workflows is much more painful for data processing than Airflow in my experience.
- flakes 1y agoIt’s a tradeoff. Ease of modeling the pipelines vs ease of managing the infrastructure. Im not really a fan of either syntax for defining DAGs, but they're the best options out there imo.
- d00mB0t 1y agoYou forgot D-Bus.
- emchammer 1y ago[flagged]
- gnat 1y agoI find the best comments here to be ones where people use their knowledge and experience to discuss the relative strengths and weaknesses of the technology in the post. I see a bunch of short single-sentence comments here that add no value. For my part, I see this pattern repeatedly at different places. The raw tools in the platforms are too codey and the third-party frameworks like Temporal seem overkill, so you build a scheduler and need to solve the problems OP did: only run once, know if it errored, etc. But it's amazing how "it's firing off a basic action!" becomes a script, then becomes a script composed of reusable actions that can pick up where they left off in case of errors ... Over time your "it's just enough for us!" feature creeps towards the framework's functionality. I'd be curious to know how long the OP's solution stays simple before it submits to the feature creep demands. (Long may complexity be fought off, though! Every day you can live without the complexity of full workflows is a blessing)
- ants_everywhere 1y agoCloud companies also provide globe-scale cronjobs that work a lot like a Unix cronjob. Arguably less mental overhead than adopting a separate framework. And such a service provides reliability guarantees. If I have to do a reliable periodic service, my go-to is a kubernetes cronjob, which is like a baby version of a cloud cronjob. I'd be reluctant to adopt some sort of task queue framework because of the complexity of the mental model plus the complexity of keeping one more thing running reliably. K8s is already running reliably, I might as well use that.
- aloha2436 1y agoMaybe I'm just lucky to work at a place with good tools, but in my experience Temporal isn't super heavyweight to use compared to building your own even-very-simple scheduler. And it's worth it because now you have Temporal, which is the bees knees as far as I'm concerned. I will gladly sing praises of any tool that saves me getting paged, and Temporal has that in spades.
- booi 1y agosecond temporal. plus it gives you more freedom to write jobs in different languages... not that you would or should in most cases but there's definitely good reasons
- deleted 1y ago[deleted]
- meatmanek 1y agoWhy use a 1 minute cron job to run the tasks, instead of a continuously-running queue worker (or several)?
- o11c 1y agoBack in the day, the reason I had 1-minute cron jobs (with flock of course) was because "what if the bespoke daemon gets killed somehow?" We also used screen/tmux a lot, but only for stuff that could afford to wait until somebody poked it (often, because if it repeatedly crashed the cause was likely novel and would need investigation). Systemd has been a game-changer for small-scale deployments.
- entropie 1y ago> Systemd has been a game-changer for small-scale deployments. The deep integration into nixos made me feel the same. You sound like you could enjoy a bit nix too.
- o11c 1y agoI dabbled a little with Nixos a while back (e.g. I think I reported the bug that broke the entire point of /etc/os-release for chroots, as well as commented on how to do a container install from scratch at a point when nobody documented it), but there were 3 things that really pushed me away: 1. Nix has clear advantages for *deployment* (including end-user deployment) but really gets in the way for new *development*. Maybe flakes fix this? Maybe not though. 2. The "Nix on other Linux" install scripts were hostile in attacking startup scripts, rather than allowing opt-in isolation. 3. The Nix language (and library?) is not sane. Nobody actually understands it, only copy-pastes pieces of existing package scripts and hopes the changes work.
- MadnessASAP 1y ago> 3. The Nix language (and library?) is not sane. Nobody actually understands it, only copy-pastes pieces of existing package scripts and hopes the changes work. Perhaps Nix is "Wonko the Sane" and it is in fact the rest of us who are in the asylum? Nix, the language, is a little strange at first but really does make sense. Nixpkgs, the "standard library", is a little stranger and sometimes makes an odd default choice. The nice thing though is that using Nix you can coerce Nixpkgs into just about any shape that suits you.
- shawn_w 1y agoIsn't a "centralized task scheduler" pretty much what cron is?
- somat 1y agoI was going to guess the author needed something that unified the task scheduling across a distributed system of computers. But that requirement is never mentioned in the article. And they still use cron to call their new scheduler... So unless I am missing something they did not replace cron al all, they just rewrote their scheduled jobs to use a common library and have more robust error handling.
- UltraSane 1y agocentralized for many computers.
- eschaton 1y agoIt’s not even a centralized task scheduler on its native UNIX: iI’s a centralized *userspace* task scheduler. Mainframe and minicomputer operating systems support scheduling in the operating system itself, as part of their process/thread scheduler; their native queuing systems are built on top of the primitives their scheduler offers, for proper accounting and maximum resource utilization (including prioritization). Only UNIX would just provide a way to run processes at a specified time or interval and call the job done.
- JdeBP 1y agoAlthough you're right that Unix never really reached having the full three-level scheduling mechanisms of the mainframe operating systems, cron is not the actual Unix parallel of the high-level scheduler that keeps the running jobs list fed. That is in fact batch (and atrun, although that's considered an implementation detail). * https://pubs.opengroup.org/onlinepubs/9799919799/utilities/batch.html https://pubs.opengroup.org/onlinepubs/9799919799/utilities/b... Most implementations flesh out the "implementation-defined algorithms" stuff to be calculations based upon load averages, as on NetBSD. * https://man.netbsd.org/batch.1 https://man.netbsd.org/batch.1 * https://man.netbsd.org/atrun.8 https://man.netbsd.org/atrun.8 Or fairly primitive parallelism limits as on Illumos. * https://illumos.org/man/1/batch https://illumos.org/man/1/batch * https://illumos.org/man/5/queuedefs https://illumos.org/man/5/queuedefs Not quite JECL, is it? (-:
- burnt-resistor 1y agoJobs that need retries, atomicity, monitoring, rescheduling, ad hoc scheduling, and flexibility probably aren't suited to most cron servers. Beanstalkd, cronicle, agenda, sidekiq, faktory, celery, etc. are the usual suspects. What is often missing is HA of the controller service process.
- sfortis 1y agoChronicle is a lifesaver. HA, clustering, API, clean UI, it's doing everything right. I'm using this also as an API wrapper for Bash and Python scripts. https://github.com/jhuckaby/Cronicle/blob/master/docs/Setup.md https://github.com/jhuckaby/Cronicle/blob/master/docs/Setup....
- mrweasel 1y agoI'd probably even add systemd timers to that list. It does most of what you list, minus the retries (but I think you could handle that in the service definition)
- burnt-resistor 1y agosystemd doesn't scale beyond one system or have high availability.
- sgarland 1y agoDo you know how many timers you could run on a single instance? An absurd amount.
- sontek 1y agoI love this solution, I've implemented a very similar task scheduler at many companies. I do think the best solution for this is still RabbitMQ. It has the ability to push tasks in the queue and tell it to run at a very specific time called "Delayed Messages" and then it just processes them at that time.
- pokstad 1y agoTemporal.io is made for this
- halamadrid 1y agoUnmeshed.io is another alternative. You don’t even need to write code for your schedules
- theshrike79 1y agoPaying $500 a month for cron just seems wrong. And adds an external dependency for something very essential.
- pokstad 1y agoYou can run it yourself for free
- pjmlp 1y agoIf they are using AWS, why not use what AWS already has, battle tested for task scheduling functions?
- nevon 1y agoI've built something similar as a service to be used by developers at a large-ish enterprise. Granted, it was based on functionality offered by AWS, but the users didn't really know that. The reason we built it, despite the fact that developers could very well have deployed a CloudWatch EventBridge schedule + SQS + lambda or similar, is because they never did. They would consistently choose to build it into their existing services, which were rarely if ever handling things like limiting concurrency if a task took too long, emitting metrics on success/failure/duration, audit logging for when a task had to be manually triggered for some reason. If I had to guess, I think the reason was because it allowed them to piggyback on existing change controls and "just write application code" instead of having to think about additional pieces of infrastructure. If I could do it again, I would probably have reached for something like Temporal, even though it seemed overkill for what we initially set out to do. It took about a week before people started asking for locking and retries.
- guappa 1y agoSo that they can drop AWS
- UltraSane 1y agoThe Windows Task Scheduler is actually very nice and powerful. One cool trick is to have a task triggered by a windows event.
- jiggunjer 1y agoAka workflow orchestrator, pipeline manager, process runner, automation tool. It's not clear if they used a product or DIY solution. The nice thing many existing products offer is a web UI and a database.
- Felk 1y agoI see that the author took a 'heuristical' approach for retrying tasks (having a predetermined amount of time a task is expected to take, and consider it failed if it wasn't updated in time) and uses SQS. If the solution is homemade anyway, I can only recommend leveraging your database's transactionality for this, which is a common pattern I have often seen recommend and also successfully used myself: - At processing start, update the schedule entry to 'executing', then open a new transansaction and lock it, while skipping already locked tasks (`SELECT FOR UPDATE ... SKIP LOCKED`). - At the end of processing, set it to 'COMPLETED' and commit. This also releases the lock. This has the following nice characteristics: - You can have parallel processors polling tasks directly from the database without another queueing mechanism like SQS, and have no risk of them picking the same task. - If you find an unlocked task in 'executing', you know the processor died for sure. No heuristic needed
- alex5207 1y agoThis is exactly what we're doing. Works like a charm.
- diarrhea 1y agoThis introduces long-running transactions, which at least in Postgres should be avoided.
- danielheath 1y agoDepends what else you’re running on it; it’s a little expensive, but not prohibitively so.
- dthedavid 1y agoGreat work. Did you consider buying instead of building? I’ve worked at organizations that built similar systems, but what was often lacking was developer experience, observability, and scalability, basically everything outside of core functionality; essentially the stuff that you're trying to tack on as you improve your system. Now that I'm building on my own, I’ve thought about building as well, but I’ve found that off-the-shelf systems handle all of this far better (and they are opensourced too), ie trigger-dot-dev and many others.
- dmitry-vsl 1y ago> We had createScheduledPosts.ts that would run every 15 minutes, scan our table of scheduled posts and create any that needed to be published. Why not set the publication_date when you create a post and have a function getPublishedPosts that fetches a list of posts, filtering out those with a publication_date earlier than the current date? With this approach, you don't need cron jobs at all.
- jon-wood 1y agoMaybe there's a bunch of other actions that need to take place when a post is published, such as sending notification emails, or posting stuff to social media. They could of course be scheduled jobs in their own right, but you haven't really saved yourself any effort there, and now if the publishing time changes you've got to reschedule all those individual jobs.
- thunderfork 1y ago[dead]
- shireboy 1y agoOne gotcha with roll your own task scheduler is if you want to run it across multiple machines. If you need 5 machines running different scheduled tasks, you need a locking mechanism to ensure only one machine is processing the task. In the author’s approach this is handled by the queue, but in my read the scheduler can only happen on one machine or you get multiple of the same task in the queue. Retry can get more complicated- depending on the failure you may want an exponential backoff, retrying N times and waiting longer durations between. A nice dashboard to see the status of everything is helpful also. In .NET world I use Hangfire for this. In Node (I assume what this is) I tinkered with Bull, but not sure what best in class is there.
- cyberpunk 1y agoOban enters the chat… :)
- sunshine-o 1y agoIs there a cool lightweight alternative to cron for (at least) a single host? To illustrate what I am looking for, I often end up using supervisord [0] (but I also like immortal [1]) for process control when not on a systemd enabled system. In my experience they are reliable, lightweight and a pleasure to work with. I am looking for something similar for scheduled jobs. - [0] https://supervisord.org/ https://supervisord.org/ - [1] https://immortal.run/ https://immortal.run/
- isp 1y agoSupercronic: https://github.com/aptible/supercronic https://github.com/aptible/supercronic Designed to run in a container, but should equally well work on a single host. However, no option for "high availability" running, where multiple hosts coordinate.
- justusthane 1y agoTake a look at this comment for some options: https://news.ycombinator.com/item?id=44752548 https://news.ycombinator.com/item?id=44752548
- majkinetor 1y agoI find Rundeck is great for this. Using it with hundreeds of jobs for a decade, with a bunch of users accessing it and checking logs, having retries, notifications and all enterprise thingies for free. Providing easy way to have GUI for scripts.
- pinko 1y agoHTCondor is always an option. Lacks shiny tinfoil, but works like a tank.
- rashidae 1y agoWhat happens when the DB gets large? How do you handle idempotency? (What if SQS delivers twice?) The cron job is still a single point of failure...
- qianli_cs 1y agoManaging complex scheduled workflows at scale comes with a lot of nuances. This is exactly why we're building DBOS (shameless plug! https://github.com/dbos-inc https://github.com/dbos-inc), which provides durable cron jobs and exactly-once workflow triggering. Since it's just a library on top of Postgres, it doesn't require a centralized scheduler (well, think of Postgres as the coordinator). One challenge is to guarantee exactly-once processing across software upgrades. DBOS uses the cron-scheduled time as an idempotency key, and tags each workflow execution with a version. We also use the database transactions to guard against conflicting concurrent updates.
- 83457 1y agoI looked around years ago and found Rundeck to be a good system for scheduled tasks.
- chriscbr 1y agoOn my current team we run a centralized task scheduler used by other products in our company that manages on the order of around ~30M schedules. To that end, it's a home-grown distributed system that's built on top of Postgres and Cassandra with a whole control plane and data plane. It's been pretty fun to work on. There are two main differences between our system and the one in the post: - In our scheduler, the actual cron (aka recurrence rule) is stored along with the task information. That is, you specify a period (like "every 5 minutes" or "every second Tuesday at 2am") and the task will run according that schedule. We try to support most of the RRule specification. [1] If you want a task to just run one time in the future, you can totally do that too, but that's not our most common use case internally. - Our scheduler doesn't perform a wide variety of tasks. To maximize flexibility and system throughput, it does just one thing: when a schedule is "due", it puts a message onto a queue. (Internally we have two queueing systems it interops with -- an older one built on top of Redis, and a newer one built on PG + S3). Other team consume from those queues and do real work (sending emails, generating reports, etc). The queueing systems offer a number of delivery options (delayed messages, TTLs, retries, dead-letter queues) so the scheduling system doesn't have to handle it. Ironically, because supporting a high throughput of scheduled jobs has been our biggest priority, visibility into individual task executions is a bit limited in our system today. For example, our API doesn't expose data about when a schedule last ran, but it's something on our longer term roadmap. [1] https://icalendar.org/iCalendar-RFC-5545/3-8-5-3-recurrence-rule.html https://icalendar.org/iCalendar-RFC-5545/3-8-5-3-recurrence-...
- jusonchan81 1y agoUnmeshed.io is a newer startup in the space - and works like a charm. Temporal seems like more targeting durable executions, but scheduling a different game. It starts with crons but soon you got to deal with holidays, adhoc skips and holds and more especially during maintenance and upgrades. Unmeshed has all of these, managing holiday calendars etc and makes it super easy. It even has agents for AS400 server commands if that is still a thing you need.
- jiehong 1y agoI think BMW used to use a paid product named Control-M to handle this (from BMC, still exists). It contained what people quickly need to reach for: - schedule a job in UTC or local time zone for a particular place; - schedule a job but only if another job ran beforehand; - semaphore-like resource limits on jobs. It did this with job generating resource tokens and other jobs stating a token as a condition for being scheduled. It ended up being a not so nice system to debug to be honest, but worked fine. For simple job, I’d reach for systemd timers on a single machine, a kubernetes cronjob on a given platform, or something external altogether otherwise (for geo-distributed scheduled jobs).