4 ms·
Workers hanging is a thing that happens everywhere. (You might not have enough resources if it's happening often, though.) You should design the workers so tha
by gervu 7y ago
Workers hanging is a thing that happens everywhere. (You might not have enough resources if it's happening often, though.)
You should design the workers so that what needs to happen still happens in the event of expected failures, or so that it at least fails gracefully and with a useful paper trail. Failures happen, good engineering anticipates and plans around them.
For example, you could schedule up to three attempts spaced at least five minutes apart, set a timeout on jobs so they don't stay open indefinitely (appearing to hang), have jobs that still fail get routed to a dead queue, and make sure worker code behaves appropriately in response to internal errors and improper input data (such as getting an HTTP error or unexpected MIME type) while logging any unexpected states for later review. Most of the point of a library like Celery is that it makes common strategies like these easier to implement.
You mentioned in a reply that the jobs are requests to external websites. The rate of errors from that is going to be like a thousand times all other sources of jobs not completing as expected unless something is hella weird with your setup.
- sharmi 7y agoThank you for taking the time to respond. That is only part of the problem. I am quite aware of that there are quite a number of external issues that can affect a job. Data extraction is something I have been working in for more than a decade. I have a few other pet peeves with Celery. I run scheduled tasks using CeleryBeat but those tasks cannot be tracked from flower. Signals don't work. I also would like something language agnostic so I can write memory/processing intensive tasks in something more performant (Go or Rust).