5 ms·
I've tried to use Airflow but was way more complicated than I expected. I just want to run a couple hundred jobs, why do I need a database? Surely a few files w
by rb808 6y ago
I've tried to use Airflow but was way more complicated than I expected. I just want to run a couple hundred jobs, why do I need a database? Surely a few files would suffice.
It seems a big gap in the market. I can't rely on cron as its a single point of failure. I have my own hardware so dont want to use AWS Batch or GCP Cloud Scheduler, any other ideas?
- BiteCode_dev 6y agoWait, you only have a couble hundred jobs, but don't want a single point of failure, but think a database is too much, but talk about cloud hosting? This all seems contradictory. Personnaly, using Python, I go for Celery (www.celeryproject.org): it's a persistant daemon that can run tasks, provide queues and shedule work like cron . A lot of people prefer Python-RQ, as it seems simpler, but the truth is you can start using celery with just the file system for storing tasks and result: https://www.distributedpython.com/2018/07/03/simple-celery-setup/ https://www.distributedpython.com/2018/07/03/simple-celery-s... If your needs grow, you can plug it to redis, rabbit MQ and/or a database later. It can expose an API so that other languages can talk to it and trigger tasks or retrieve results (but not write tasks, they must be in python).
- ramraj07 6y agoCelery for scheduled jobs seem to not be a supported design pattern at all, and any job that starts to potentially come close to the 1 hour timeout seems to get annoying to work with in celery. It seems primarily designed to send emails in response to web requests, which is not the use case most people are discussing here.
- BiteCode_dev 6y agoI don't see how you came to this idea. The jobs can be as long as you want, you can have retry, persistant queues, priorities, and dependancies. Of course, I would advice to put a dedicated queue for very long running tasks, and set worker_prefetch_multiplier to 1 as the doc recommand for long running tasks: https://docs.celeryproject.org/en/stable/userguide/optimizing.html https://docs.celeryproject.org/en/stable/userguide/optimizin... With flowers (https://flower.readthedocs.io/en/latest/ https://flower.readthedocs.io/en/latest/), you can even monitor the whole thing or deal with it manually. I assume your comment is reporting on other comments, but not direct experience?
- ramraj07 6y agoDirect experience very fresh in memory :) The issue with long running tasks is that you have to change the timeout to longer than the default value of one hour (otherwise the scheduler assumes the job is lost and requeues it). But this is a global parameter across all queues so this means we essentially loose the one good feature of celery for small tasks which is retrying lost tasks within some acceptable timeframe. Further flower seems weird - half the panels don't work when connecting through our servers; our vpc settings are a bit bespoke but not completely out there, so it's not fully useful. Also flower only keeps track of tasks queued after you start the dashboard (but then it accumulates a laundry list of dead workers across deployments if you keep it running continuosly). We were also excited to use it's chaining and chord features but went into a series of bugs we couldn't dig ourselves out of when tasks crashed inside a chord (went into permanent loops). I just declared bankruptcy on these features and we implemented chaining ourselves. Point is, I'm sure we got some parameters wrong, but I and another engineers spent WEEKS wrangling with celery to at least get it running somewhat acceptably. That seems a bit too much. We are not L10 Google engineers for sure but we aren't stupid either. The only stupid decision we made was probably choosing celery from what I can see. In the end we still keep celery for the on demand async tasks that run in a few minutes. For scheduled tasks that run weekly, we just implemented our own scheduler (that runs in the background in our webservers in the same elastic beanstalk deployment) that uses regular rdbms backend and does things as we want. Turns out it's just a few hundred lines of simple python.
- sillycube 6y agoYes, I feel the same after spending several weeks to deal with celery, redis, docker compose config, flower, set up celery workflow, rate limit, worker, etc. Testing for a really long time... When there is a workflow to play with chord & chain, it became unintuitive to find out the issue. I was stuck in an issue and finally I posted on SO to ask for help. Luckily I got an answer I thought it's my problem due to no experience in scheduling stuffs. I hope there is something which is simpler
- BiteCode_dev 6y agoFair enough and very honest. > But this is a global parameter across all queues so this means we essentially loose the one good feature of celery for small tasks which is retrying lost tasks within some acceptable timeframe. Oh, for this you just setup two celery deamon, each one with their own queues and config. I usually don't want my long running task on the same instance than the short ones anyway. > We were also excited to use it's chaining and chord features but went into a series of bugs we couldn't dig ourselves out of when tasks crashed inside a chord (went into permanent loops). I just declared bankruptcy on these features and we implemented chaining ourselves. Granted on that one, they not the best part of celery. Just out of curiosity, which broken and result backend did you use for celery? I mostly use Redis as I had plenty of problems with Rabbit MQ, and wonder if you didn't have those because of it.
- somurzakov 6y agoWindows Task Scheduler, it is way more powerful and robust than many people think
- rb808 6y agoAgreed its very good, but has same problem as Cron with a single point of failure. I want to add I can handle a single point of failure if all job definitions are in git or a directory of flat files.
- somurzakov 6y agoi dont really understand concept of SPoF. My windows servers have never really failed me. In case machine reboots in the middle of your batch - Task scheduler can restart your job. In case the job fails, it can retry after a certain interval. you can configure your job to retry unless it returns success - and unless someone nukes your windows machine - it will get executed. if you are afraid that your windows machine will get nuked - then you can use SQL Server Agent on High Availability Cluster - and it will do the job. if you want your jobs as a code - you can code them in Powershell and store in git/azure devops repo. You can deploy your jobs with the same powershell, or even use a CI/CD pipeline to do that.
- ramraj07 6y agoIt's probably no more safe to assume a single server will never fail when we're talking about AWS or GCP. Honestly maybe we can, given I see ec2 instances running for four years without even a restart, but it's still not in the general dev philosophy of cloud at the least.
- ramraj07 6y agoSomeone else who has the same problem! Looks like airflow, celery and every other workflow orchestrator doesn't want to deal with it and just asks you to deal with it using shit like k8s. I decided to write a simple scheduler that runs in a separate thread in the background of our webapp, so that it can be parceled into our existing elastic beanstalk deployment. I use two database tables to pick up tasks and run them, and have some amount of failover. Just need to be a bit careful around deadlocks, but thats a cakewalk compared to the dumpster fire that is configuring these Babylon tower frameworks. If you're interested I can write up some sample scheduler code and publish it which shouldn't be more than a few hundred lines and do what we're looking for.