6 ms·
Hi! Author of the article here. The core concern is not about the capabilities of the compute abstraction being used (bare metal, containers or functions) or
by neo01124 6y ago
Hi!
Author of the article here.
The core concern is not about the capabilities of the compute abstraction being used (bare metal, containers or functions) or testing OS capabilities. The aim is to validate mitigations which are in place to counter turbulent scenarios (For example: massive spike in traffic, network outage, dependency is down, etc). These scenarios generally originate outside the given system.
These kind of questions should be asked and systematically validated (quoting the article):
* Have you tested how the system behaves when the underlying instances have a sustained CPU spike?
* Is the system behavior understood under different stress?
* Is there sufficient monitoring?
* Have the alarms been validated?
* Are there any countermeasures implemented? For example, is auto-scaling set up, and does it behave as expected? Are timeouts and retries appropriate?
- fxtentacle 6y agoI believe we just have a rather different approach here. "Have you tested how the system behaves when the underlying instances have a sustained CPU spike?" Since dedicated boxes are cheap, I'd just buy 5x the CPU resources that I reasonably need and call it a day. If there ever is a more than 5x traffic spike, then docker will prevent it from being a noisy neighbor, so the affected services will just become slower than usual. But even a 10x traffic multiplier would just produce a 2x slowdown, which should be tolerable for most users. I agree that on clouds you want to save costs by only booking what you need. But bare metal, you can usually afford to keep spare capacity around all the time. As such, I wouldn't plan for the system to behave well under stress. I'd try to always have enough resources around so that stress never happens. At the end of the day, this seems like a developer time vs. resource costs trade-off and for most companies, developers are sparse and resources are plentiful, so they'll have a very different trade-off from big FAANG companies. "For example, is auto-scaling set up, and does it behave as expected?" If your system is usually 90% idle, I wonder if you'll ever need that auto-scaling. Also, I'd say my customers can endure it if page load time goes up from 100ms to 200ms. So in my opinion, there is little need for auto-scaling for most companies.
- ses1984 6y ago>"Have you tested how the system behaves when the underlying instances have a sustained CPU spike?" You didn't really address this question, you addressed a different question, which is a traffic spike. >Also, I'd say my customers can endure it if page load time goes up from 100ms to 200ms. So in my opinion, there is little need for auto-scaling for most companies. 100ms to 200ms average? What about the tail? Your app might go from P99 - 500ms to P95 - timeout. That's when you'll lose customers.
- fxtentacle 6y agoIf the underlying hardware is a bare metal server, it won't magically turn slow and have a CPU spike. That problem is caused noisy neighbor and kind of exclusive to clouds. Well, with the 2x example, my app might get from a 1s P99 to a 2s P99 which feels slow, but is still doable. Again, those timeouts are usually introduced by cloud infrastructure. For example, if you use nginx outside of Heroku, it won't have a 30s timeout for file downloads.
- ses1984 6y agoYour own instances can have an unexpected CPU spike. Even if you're running on bare metal I find it hard to believe you don't have a layer with short timeouts between your front and backend.
- fxtentacle 6y agoWhy would I? I have redundant 1GBit LAN cables between front end, back end, and database servers.
- ses1984 6y agoBecause it's bad ux for your users to see a spinning loading icon forever.
- 6y ago