5 ms·
I suspect most business logic can handle 25ms for authz and that’s the right trade off. I think Google’s Zanzibar is also centralized but leverages extreme cach
by leoqa 3y ago
I suspect most business logic can handle 25ms for authz and that’s the right trade off. I think Google’s Zanzibar is also centralized but leverages extreme caching to get lower latencies?
I work on an IAM system that is sub-ms p99 for our authz checks, with policies and keys pushed to each network edge instead of running a centralized system. The biggest perf hits are crypto verification and logging to the fs. We fail-closed to last known policy state when we have partitions, data loss would imply the application service or datastore proxy is lost. We measure policy deploy times in minutes though, and it’s eventually consistent.
- prpl 3y agonot sure if it applies but depending on instance type I usually see pings in the .55ms range in a single AZ in AWS, cross-AZ pings higher (implying it is hard to be sub ms for many types of durable applications, especially if disk/S3 is involved)
- magden 3y agoWhen discussing GCP, the latency between AZs within the same region is approximately 5 ms. Thus, if you have a 3-node database cluster spanning 3 AZs (all within the same region), a transaction can be committed in the 5-10ms range using Raft. CockroachDB should handle this seamlessly. However, if you're considering a multi-region setup, the latency will depend on the distance between the regions.That's why usually you define a preferred region (that stores primary copy of the records) or deploy in a geo-partitioned mode (when data is automatically pinned to configured regions).
- jzelinskie 3y ago>I think Google’s Zanzibar is also centralized but leverages extreme caching to get lower latencies? That's correct. In a Zanzibar-like model, you have a global storage, but individual clusters in each datacenter/edge providing consistency-aware caching. This means p99 can be something like 25ms, but p95 or p50 is often FAR lower. Disclosure: I'm a co-creator and maintainer of SpiceDB[0] [0]: https://github.com/authzed/spicedb https://github.com/authzed/spicedb
- leoqa 3y agoI’ve watched y’all’s Papers We Love talk about Zanzibar and have recommended authzed to organizations bootstrapping permission modeling. It’s been awhile, is the gist that Spanner’s coordinated clocks allow tighter consensus (i.e. faster writes) and caching provides read-my-write consistency?
- jzelinskie 3y agoThanks for watching our presentation and recommending our solution. Unfortunately, nothing is ever simple; comparing Spanner and CockroachDB is comparing apples to oranges. Two years ago, we wrote an article that details exactly how the differences matter in terms of a Zanzibar implementation[0], but I can give as short of a summary as possible: Spanner is linearizable and CockroachDB only guarantees external consistency for transactions that share rows. The post outlines how we workaround this and we've also more recently talked about how we've managed to scale that to 1M requests per second[1]. Our team focuses a lot on CockroachDB because we offer a permission systems that can span not only regions within a single cloud, but across various cloud providers. However, if you're all in on GCP, SpiceDB itself supports Cloud Spanner (which we also use in production for our GCP-only customers). [0]: https://authzed.com/blog/prevent-newenemy-cockroachdb https://authzed.com/blog/prevent-newenemy-cockroachdb [1]: https://authzed.com/blog/maximizing-cockroachdb-performance https://authzed.com/blog/maximizing-cockroachdb-performance
- audioheavy 3y agoI can attest to that statement: comparing Spanner and CockroachDB is difficult. Spanner and Fauna (where I work) are more comparable (Fauna is based on Calvin, see [0]) since they both support strict serializability (in different ways). The article referenced here is excellent, and it highlights what we've seen from some customers: CockroachDB is (to say the least) a challenge to learn and adequately deploy, I've seen a few others that have a similar lessons-learned result. I'm glad, however, that highly consistent distributed databases provide value in these implementations. Although not OSS, Fauna is comparable and more turnkey (read: much less ops) than these options. [0]: https://fauna.com/blog/distributed-consistency-at-scale-spanner-vs-calvin https://fauna.com/blog/distributed-consistency-at-scale-span...
- aeneas_ory 3y agoOne of Ory’s core competencies is permissions. We built the first Google Zanzibar implementation in the world and it’s part of Ory Network‘s global multi-region platform (https://github.com/ory/keto https://github.com/ory/keto) A push model is also valid if you’re heavy on policies and can accept eventual consistency. We will investigate how to generally push things to the edge (like we did with Ory Edge Sessions) or to cryptographic verification wherever staleness is acceptable. By solving the primitives correctly in the beginning (with a multi region architecture) that job does become a lot easier, which is what we decided doing at Ory :)
- leoqa 3y agoMy intuition is that offering this as a service you’re targeting business logic that can handle 25ms authz. We’re on the core path in a latency sensitive industry, and end up running many permissions checks at various layers for a single api call.
- aeneas_ory 3y agoAbsolutely, having P99 of sub-ms is of course way more attractive than 25ms - with a SaaS offering you always have the network latency to the provider in the path, which is why multi region capabilities are so important for this case. But you’ll never beat systems where the decision can be made locally. Have you any documentation on your approach publicly available? I‘d love to get some education and insights from other large scale authz systems! We have a couple of ideas such as running a local replica in our customer’s stack but nothing concrete yet.
- leoqa 3y agoIt’s quite similar mechanically to this blog post about Uber’s policy framework: https://www.uber.com/blog/attribute-based-access-control-at-uber/ https://www.uber.com/blog/attribute-based-access-control-at-.... We have an additional scaling dimension though, as our permission model is richer and mutable by end users, therefore our policies are not uniform. For special hot-path services, we use symmetric keys to reduce latency further but that makes rotations complicated.
- KRAKRISMOTT 3y agoWhat does GitHub use for their Authz? After inviting a user, there is no perceptible sync delay and they can start cloning the repo immediately.
- tonyhb 3y agoNot sure, but I'd guess they're using Vitesse (ie. Planetscale) which is honestly really fast and durable.