4 ms·
The difference between a one off script and a production-grade system with documentation, instrumentation, and monitoring is pretty big. Here's less than a min
by daotoad 6y ago
The difference between a one off script and a production-grade system with documentation, instrumentation, and monitoring is pretty big.
Here's less than a minute's worth of questions I would want answered for a professional version of the log fetching example:
* Is that list of a thousand servers static?
* How is it updated?
* What format is it in?
* What do you do when the resource isn't available?
* What other systems are impacted when the resource can't be contacted?
* Is it acceptable to use a cached server list until it the resource can be updated?
* How do you handle servers that don't respond when polled?
If we retry connections, what scheme do we use? Linear or exponential back off?
* After a service interruption, how do we resume data collection? Do we keep rolling forward or do we request older data? How quickly do we retrieve backed up data? Is service resumption plan different for the case when a single system is down vs when the log fetching system has an outage?