6 ms·
(Disclaimer, I worked at Google over a decade ago, for about half a decade; as an SRE for about 2 of those years) Yes, it absolutely makes sense. You never kno
by clavoie 6y ago
(Disclaimer, I worked at Google over a decade ago, for about half a decade; as an SRE for about 2 of those years)
Yes, it absolutely makes sense. You never know when a typo in a config file or bug in your job-management code will go bonkers and try to take over all resources. Same think as disk quotas on computers with tons of users: you want to limit the damage that a user can accidentally (or not...) do.
Good quota systems saved my team's bacon way back when a few times, when fewer people were at the company; I can only imagine how useful they are at what, 10-20x the size?
- byecomputer 6y agoGood thing quota systems can't go bonkers and take over all the resources, that would be a nightmare!
- mav3rick 6y agoWhat's the need for the snark ? Bugs can happen anywhere. Doesn't mean we don't try ?
- byecomputer 6y agoI guess I thought the irony was amusing. Computers doing unexpected things is a universal experience, so I don't think there's any need to take it personally. You're also not the person I responded to, so I take it you have some sort of distaste for verbal irony no matter where it happens? Or you work for Google
- mav3rick 6y agoYou seemingly refuted a good solution. Rather than guessing my background or character, introspect about your comment.
- deleted 6y ago[deleted]
- byecomputer 6y agoI made guesses because you took an incongruously dark read of my comment that seemed more reflective of internal disquiet than anything I said. I'll respectfully ask that you refrain from being hostile towards others based on assumptions regarding their intent
- ChrisRR 6y agoAnd that's precisely why you implement checks at multiple levels. You can't rely on one level to always be well behaved and so you never have to worry about issues at another level. Not sure why you had to be snarky about it
- antonvs 6y agoStories like this remind me of stories like the one about avoiding death: https://www.k-state.edu/english/baker/english320/Maugham-AS.htm https://www.k-state.edu/english/baker/english320/Maugham-AS....
- dang 6y ago"Don't be snarky." "Please don't post shallow dismissals, especially of other people's work. A good critical comment teaches us something." https://news.ycombinator.com/newsguidelines.html https://news.ycombinator.com/newsguidelines.html
- byecomputer 6y agoThat's injecting a lot of malicious intent into what I said. I don't think making an unreasonably dark assumption about someone's intent is a good reason to treat them with hostility.
- dang 6y agoI believe you about your intent, but the problem is that we have to judge these things by the effect that they produce in threads. Intent doesn't communicate itself, unfortunately, and if a comment takes a flamey or trollish form, that's the sort of effect it's likely to produce. So the burden is on each of us to disambiguate our intent. It doesn't happen automatically. Probably the most reliable way to do that is to include enough markers of good intent in a comment to shift its pH a bit. Here are some previous explanations in case they're helpful: https://hn.algolia.com/?dateRange=all&page=0&prefix=true&sort=byDate&type=comment&query=troll%20effects%20by:dang https://hn.algolia.com/?dateRange=all&page=0&prefix=true&sor... https://hn.algolia.com/?dateRange=all&page=0&prefix=false&sort=byDate&type=comment&query=disambiguate%20burden%20by:dang https://hn.algolia.com/?dateRange=all&page=0&prefix=false&so... https://hn.algolia.com/?dateRange=all&page=0&prefix=true&sort=byDate&type=comment&query=%22expected%20value%22%20by:dang https://hn.algolia.com/?dateRange=all&page=0&prefix=true&sor...
- byecomputer 6y agoFair enough, I got the impression that the user who called me snarky was gaming the rules to get my comment deleted, so I was a bit annoyed. After the first page, their comment history is mostly sarcastic and demeaning remarks, so it was clear to me that they simply didn't like what I said for personal reasons. I realize it doesn't make a difference and perhaps you reached your conclusion independently of their remark, but I was understandably a little miffed and felt like I got played. Regardless, I'll remain conscious of what you said.
- darig 6y agoI think google has killed more products than it has added in the last decade (ignoring rebrandings), so I'd consider them less than 1x the size with 10-20x the liabilities. Why wouldn't you want to limit the damage that instilling distrust in your biggest profit centers can do?
- zadokshi 6y agoWell here’s the thing though. Google are removing quotas from things like appengine. How can it be simultaneously helpful to have quotas but at the same time google wants to remove them?
- alanfranz 6y agoI think you're conflating very different things. Quotas may be removed if you've got a virtually unlimited amount of a certain resource (and even there they can be useful to prevent overbilling), but disk quotas are necessary because disk space is not unlimited. Without those, a noisy neighbour can break your application.
- anthony_r 6y agoIn particular you cannot provide an internal SLA on write requests to low-level storage solutions if you can't guarantee that you actually have the spare disks. The SLA would have to be something like "we guarantee 99.999% write availability unless team A or team B or team C makes a mistake", so now you have to look up what team A and B and C are up to and what do they guarantee. Instead of the more sensible "we guarantee 99.999% write availability unless you are out of quota".
- clavoie 6y agoExactly. At some point, someone somewhere will have a bad day / sneeze on a keyboard at just the perfectly wrong time -- at Google's scale, that's statistically a daily (hourly?) occurrence. Quota systems for the win -- if only as a shared agreement on who can do how much damage / take how much resources before they're automatically stopped and more resources must be justified through some review process (likely involving budgetary concerns).
- mrighele 6y agoShouldn't there be a soft limit that start sending out warnings before hitting the hard one ? Unless some service started eating the quota so fast that it reached both limits in quick succession, the upper limit should never be reached.
- erhk 6y agoFrom the article capacity was reduced
- mrighele 6y agoYou're right. Still, a better approach could have been adopted, like blocking a change that would cause some service to go over quota, or have the automatic system change the quotas gradually, thus giving time for alerts and human intervention
- cranekam 6y agoI'm sure that this Google outage was the combination of several independent issues. Perhaps there are the kind of safeguards you mentioned in place, but they returned an incorrect view of the world because of a bug exposed by a network partition? The point is that with huge distributed systems the things take them down (at least once they're mature) are typically the result of several failures that compound and interact in a way that wasn't foreseen. IMO it's more likely that the suggestions you made (blocks and/or automation) are already in place but failed in some more complex way. Source: worked at a FAANG for a decade, saw many big incidents. Almost all were the confluence of several smaller issues that, had they happened alone, wouldn't have been newsworthy.
- clavoie 6y agoI can't comment on the current incident, as I've been gone for more than a decade. But my recollection is that most SRE teams would have had exactly such things -- either as really crazy borgcfg (sorry, I mean kubectl) code, or as a borgmon (sorry, prometheus) alert. The former were a real PITA to work with (borgcfg hadn't been designed with that in mind), and the latter were only reactive (and thus unable to warn about things until they were fast becoming a real problem). So most likely those safeguards were somehow disabled or made ineffective by the exact circumstances of the problem, possibly the speed at which things developed.