3 ms·
As a capacity planner I tried to argue in favor of tools like Hystrix being built into our middleware services because when we had large IPPV events cascading f
by Sylamore 3y ago
As a capacity planner I tried to argue in favor of tools like Hystrix being built into our middleware services because when we had large IPPV events cascading failures was the biggest risk to our availability and it happened because 99% of the time our services could process any queues before downstream timeouts occured but during high volume events the queues would grow due to nearly instant demand occurring (normally within 5 minutes of the PPV event start time) and causing the queues to go deeper than our timeouts. Combine that with the queues not being durable if a process restart was needed and things got real ugly real fast under extreme load. Automatic retries + deep queues + short timeouts = service issues that take hours to unwind often requiring a coordinated cold restart of the entire middleware pipeline and millions in lost revenue.
To compensate we had to scale our systems for the absolute instantaneous peak demand because being legacy systems (pre containers) with a lot of rigid plumbing in them we couldn't just scale on demand. Things did not degrade gracefully once you hit the timeout limits in the composite API calls.