3 ms·
I've been bit by the same issue on ECS. Some stream processing applications restarted and they perform a lot of reads at startup to recover their in-memory sta
by hashhar 6y ago
I've been bit by the same issue on ECS.
Some stream processing applications restarted and they perform a lot of reads at startup to recover their in-memory state. All other containers on that instance eventually also restarted and got migrated to another EC2 instance which also got IOPS depleted soon enough.
And the cycle continues. The issue was there was no proper monitoring set up and getting a timeout from Docker isn't very helpful error message.
Since then I've made sure to build in checks to prevent bouncing all services on a machine at once and spreading out applications that use disk across machines instead of binpacking.
- 3np 6y agoThe icing on the cake here is that those IOPS from the docker agent is outside of your control. Before having to dive deeper into it, I would have assumed that only IOPS stemming from the workloads themselves would count against quotas. Using these abstractions of abstractions of abstractions that all end up leaking fatal failure modes you have to deal with yourself makes me start questioning the fundamental value proposal. The one major thing you get away from is setup costs, but the total time investment gets amortized.