5 ms·
My favorite quote: "prevent fast implementation of global changes." This is how large organizations become slower compared to small "nimble" companies. Everyone
by physicsgraph 6y ago
My favorite quote: "prevent fast implementation of global changes." This is how large organizations become slower compared to small "nimble" companies. Everyone fails, but the noticeability of failure incentivizes large organizations to be more careful.
- Xorlev 6y agoSlowing down actuation of prod changes to be over hours vs. seconds is a far cry from the large org / small org problem. Ultimately, when the world depends on you, limiting the blast radius to X% of the world vs. 100% of it is a significant improvement.
- inopinatus 6y agoIt doesn’t have to be this way, but that’s partly a matter of culture. By aspiring to present/think/act as a monoplatform, Google risks substantially increasing the blast radius of individual component failure. A global quota system mediating every other service sounds both totally on brand, and also the antithesis of everything I learned about public cloud scaling at AWS. There we made jokes, that weren’t jokes, about service teams essentially DoS’ing each other, and this being the natural order of things that every service must simply be resilient to and scale for. Having been impressed upon by that mindset, my design reflex is instead to aim for elimination of global dependencies entirely, rather than globally rate-limiting the impact of a global rate-limiter. I’m not saying either is a right answer, but that there are consequences to being true to your philosophy. There are upsides, too, with Google’s integrated approach, notable particularly when you build end-to-end systems from public cloud service portfolios and benefit from consistency in product design, something AWS eschews in favour of sometimes radical diversity. I see these emergent properties of each as an inevitability, a kind of generalised Conway’s Law.
- simon1573 6y agoDo you blog? I really enjoyed reading that.
- inopinatus 6y agoThanks. I'm more a forum-dweller when it comes to self-expression. There's an obvious .org but you'll be sorely disappointed, unless you're looking for arcane and infrequent Ruby/Rails tips.
- ridaj 6y agoWell, GCP followed that principle as their service account auth mechanisms did not fall over. So if you were to compare with AWS, it looks like something similar was happening. The infrastructure running the auth service, like any service, is going to have a quota system, whether it's global or not. The lesson learned might be different if it wasn't global ("prevent fast changes to the quota system for the auth service") but the conclusion would be substantially similar - there is usually no good reason, and plenty of danger, for routine adjustments to large infrastructure to take place in a brusque manner. That doesn't mean that non-infrastructure service needs to abide by the same rules...
- inopinatus 6y ago> The infrastructure running the auth service, like any service, is going to have a quota system, whether it's global or not. Don’t assume that is the case. That’s exactly the kind of cultural assumption I’m speaking of. Case in point, I routinely run services without quotas or caps and what have you, scale out for load, and alarm on runaway usage, not service unavailable or quota exceeded. I’d rather take the hit than inconvenience my customers with an outage. In this frame of mind, quotas are a moral hazard, a safety barrier with a perverse disincentive. That principle of “just run it deep” works even at global scale, right up until you run out of hardware to allocate, which is why growth logistics are a rarely-discussed but critical aspect of running a public cloud. The core learnings become, how to factor out such situations at all. That might be through some kind of asynchronous processing (event-driven services, queues, tuplespaces), or co-operative backpressure (a la TCP/IP) and so on. Synchronous request/response state machines are absolute murder to scalability, so HTTP, especially when misappropriated as an RPC substrate, has a lot to answer for.
- ridaj 6y ago> Don’t assume that is the case. What I mean is, it's going to have limits of some sort, right? The world is finite...
- inopinatus 6y ago
- felixhuttmann 6y agoI often hear 'aim for elimination of global dependencies', but the reality is that there is no way around global dependencies. AWS STS or IAM is just as global as google's. The difference is that google more often builds with some form of guaranteed read-after-write consistency, while AWS is more often 'fail open'. For example, if you remove a permission from a user in GCP, you are guaranteed consistency within 7 minutes [1], while with AWS IAM, your permissions may be arbitrarily stale. This means that when the GCP IAM database leader fails, all operations will globally fail after 7 minutes, while with AWS IAM, everything continues to work when the leader fails, but as an AWS customer, you can never be sure that some policy change has actually become effective. In general, AWS more often shifts the harder parts of global distributed systems onto their customers, rather than solving them for their customers, like GCP does. For example, GCP cloud storage (s3 equivalent) and datastore (nosql database) provide strongly consistent operations in multi-region configurations, while dynamodb and s3 have only eventually consistent replication across regions; and google's VPCs, message queues, console VM listings, and loadbalancers are global, while AWS's are regional. [1] https://cloud.google.com/iam/docs/faq#access_revoke https://cloud.google.com/iam/docs/faq#access_revoke
- silentsea90 6y agoS3 is strongly consistent. https://aws.amazon.com/s3/consistency/ https://aws.amazon.com/s3/consistency/ Which of Google's nosql db provides strong consistency - bigtable? Just confirming
- felixhuttmann 6y agoS3 is newly strongly consistent within a single region since last reinvent or so (google cloud storage has been strongly consistent for much longer). However, the cross-region replication for s3 is based on 'copying' [1] so presumably async and not strongly consistent. GCP datastore and firestore are strongly consistent nosql databases that are available in multi-region configurations [2]. [1] https://docs.aws.amazon.com/AmazonS3/latest/dev/replication.html https://docs.aws.amazon.com/AmazonS3/latest/dev/replication.... [2] https://cloud.google.com/datastore/docs/locations https://cloud.google.com/datastore/docs/locations
- erhk 6y agoIf Google is flaky you use yahoo. If Facebook/Twitter/instagram is flaky you wait until it isn't and then post that update.
- evil-olive 6y agoSlow rollouts can be a double-edged sword, too: > a change was made in October to register the User ID Service with the new quota system, but parts of the previous quota system were left in place which incorrectly reported the usage for the User ID Service as 0. An existing grace period on enforcing quota restrictions delayed the impact, which eventually expired, triggering automated quota systems to decrease the quota allowed for the User ID service and triggering this incident. Grace period on enforcement of a major policy change is an excellent practice...but it also means months can go by between the introduction of a problem and when the problem actually surfaces. That can lead to increased time-to-resolution because many engineers won't have that months-old change at the front of their mind while debugging.
- azornathogron 6y ago"Slow" isn't really a precise enough descriptor. You need gradual rollouts. In particular, you need rollouts where the behavior of your system changes gradually as you apply your change to more of your instances/zones/whatever-rollout-unit. And the right speed is whatever speed gives you enough time to detect a problem and stop the rollout while the damage is still small enough to be "acceptable". With "acceptable" determined by the needs of your service (but if you say "no damage is ever acceptable" then I have some bad news for you). Grace periods don't give you gradual rollouts like this; that's not their purpose. And I agree, grace periods can be a double edged sword for the reason you mention.
- ffggvv 6y agoits a paradox because they should be slower when they have so many things depending on them and they can't afford to fail. so being quicker would actually be dumb of them.
- mulcahey 6y ago"Move fast and break things!" ... "Move fast! ...with stable infrastructure!" [1] [1] https://www.cnet.com/news/zuckerberg-move-fast-and-break-things-isnt-how-we-operate-anymore/ https://www.cnet.com/news/zuckerberg-move-fast-and-break-thi...
- notacoward 6y agoHad the same thought. Saw this same scenario play out many times between services at FB, and I'm still really not sure there's a good "one size fits all" answer either there or at peer companies like Google. For every "just do X" I've seen here I could probably identify the incident where that fix led to or exacerbated a different outage. Sometimes teams don't collaborate well, and that requires a specific fix beyond outsiders' view instead of more platitudes.
- KajMagnus 6y agoI'm pretty certain that that's a misunderstanding, and rollouts at Google still happen faster than at most other companies incl small & mid size. It's just that Google won't rollout to the whole world in one instant step, instead, it depends: can be hours, days or weeks. Or sometimes a first small step during less than a minute, and then a more large scale deployment — e.g. if there's an urgent security bugfix. It's in the SRE book (well as far as I remember), I think you'll find it if you search for "seconds" or "minutes" https://static.googleusercontent.com/media/sre.google/en//static/pdf/building_secure_and_reliable_systems.pdf https://static.googleusercontent.com/media/sre.google/en//st... But yes definitely there are other things that slow down big companies.