6 ms·
Blueprint for a distributed multi-region IAM with Go and CockroachDB
- nosequel 3y agoThis isn't the typical 1000 word, "here's how we did it, now use our thing" company fluff blog post. What a great writeup. Sometimes reading docs, it is hard to figure out the fine details when making a decision. Your comparison of Regional Tables, Regional By Row Tables, and Global Tables is a really nice summary of the pros & cons of each. Well done.
- aeneas_ory 3y agoThank you, I appreciate that feedback because this was the explicit goal of writing the article: Informing that multi region is no longer just a vision for companies that aren’t Google; Sharing how difficult it is; And some of the learnings made along the way! Personally, I am extremely proud of the work. I believe that in a year or two, most companies will adopt multi region IAM (hopefully from Ory as we’re currently the only ones capable of this). :) And what could be better than hearing these kind words from the critical readers on HN :) Cheers!
- maxpert 3y agoOne of the reasons I started writing Marmot (https://maxpert.github.io/marmot/ https://maxpert.github.io/marmot/) was for replicating bunch of tables across regions that were read heavy. I even used it for cache replication (because who cares if it’s a cache miss, but a hit will save me time and money). It’s hard to make such blue prints in early days of product, and by the time you hit a true growth almost everyone builds a custom solution for multi-region IAM.
- aeneas_ory 3y agoThat is true! And the reason why we decided to build this and offer it to everyone - building multi-region IAM is incredibly difficult and expensive and typically not the core competency of an average software company. Also, very interesting project. I love SQLite and what the community is contributing to it - yours included!
- leoqa 3y agoI suspect most business logic can handle 25ms for authz and that’s the right trade off. I think Google’s Zanzibar is also centralized but leverages extreme caching to get lower latencies? I work on an IAM system that is sub-ms p99 for our authz checks, with policies and keys pushed to each network edge instead of running a centralized system. The biggest perf hits are crypto verification and logging to the fs. We fail-closed to last known policy state when we have partitions, data loss would imply the application service or datastore proxy is lost. We measure policy deploy times in minutes though, and it’s eventually consistent.
- prpl 3y agonot sure if it applies but depending on instance type I usually see pings in the .55ms range in a single AZ in AWS, cross-AZ pings higher (implying it is hard to be sub ms for many types of durable applications, especially if disk/S3 is involved)
- magden 3y agoWhen discussing GCP, the latency between AZs within the same region is approximately 5 ms. Thus, if you have a 3-node database cluster spanning 3 AZs (all within the same region), a transaction can be committed in the 5-10ms range using Raft. CockroachDB should handle this seamlessly. However, if you're considering a multi-region setup, the latency will depend on the distance between the regions.That's why usually you define a preferred region (that stores primary copy of the records) or deploy in a geo-partitioned mode (when data is automatically pinned to configured regions).
- jzelinskie 3y ago>I think Google’s Zanzibar is also centralized but leverages extreme caching to get lower latencies? That's correct. In a Zanzibar-like model, you have a global storage, but individual clusters in each datacenter/edge providing consistency-aware caching. This means p99 can be something like 25ms, but p95 or p50 is often FAR lower. Disclosure: I'm a co-creator and maintainer of SpiceDB[0] [0]: https://github.com/authzed/spicedb https://github.com/authzed/spicedb
- plexicle 3y agoAwesome post, really. One of the best I've read in a while! Total side question, if anyone knows -- what tool (if any?) was used for the graphics in this article? The dot matrix looking map style stuff? I really dig it.
- aeneas_ory 3y agoThank you! I appreciate that a lot! Our designers will love that feedback! Unfortunately it’s not a shelf product but they used Figma to design the graphs.
- endisneigh 3y agoSeems that cockroachdb saved the day here with its multi region capabilities. Did the other vendors have the same capabilities? Specifically things like the regional tables and columns.
- aeneas_ory 3y agoWe made the decision to choose Cockroach in 2018 and back then no product had these capabilities. We stuck to CRDB because they delivered on their product vision and as far as our research went they have the most advanced solution. Even Google Cloud Spanner (NOT the same as Google Spanner - the internal DB) lacks a couple of things we needed for data homing.
- ushakov 3y agoIf you’d like to deploy a containers or even Ory itself to multi-region cloud, you should check out EdgeNode (https://edgenode.com https://edgenode.com), which I helped build
- hiatus 3y agoA friendly note: when I visited your site, I immediately clicked away when I saw that learning more about the deployment process, pricing, etc required me to sign up.
- ushakov 3y agoThanks for the feedback. This helps us a lot! We’re currently in a early stage, but more info will be available to public in the next couple of weeks. You can fill out the form, if you want to be in the known: https://tally.so/r/w2ajRb https://tally.so/r/w2ajRb Thanks again, appreciate your honest feedback.
- proleisuretour 3y agoThis is a great article about building global apps that require multi-region deployments. Thanks for sharing. Curious about the transaction retry errors for UPDATE that required 2 days to resolve. Probably could of been avoided using a distributed SQL database that supports a read committed isolation level ¯\_(ツ)_/¯ For those going down this path, maybe check out open source YugabyteDB. There is a great doc about how to build Global Apps using various application design patterns: https://docs.yugabyte.com/preview/develop/build-global-apps/ https://docs.yugabyte.com/preview/develop/build-global-apps/
- Randis 3y agoFlexibility in isolation levels is coming in CRDB :)
- Sytten 3y agoGood post, side remark our experience with kratos have been mixed while self hosting the solution. You can feel OSS is second class for them (lots of PR never getting merged, endless debates and little progress in code), it's OK its a business and they are not doing support contracts. Just know what you are getting into. Just my experience, might be different with other products.
- vinckr 3y agoHey Sytten, I work in the community team for Ory. OSS is actually very important for us and definitely not second class. (see also https://www.ory.sh/docs/open-source/commitment https://www.ory.sh/docs/open-source/commitment) The assumption that Ory does not offer support contracts for self-hosted Ory is wrong (although we did not in the past, when the team was smaller). We are doing contracts for companies using our software self-hosted: See here and contact us if you are interested! https://www.ory.sh/support/ https://www.ory.sh/support/ This way we can assign engineers to your case and work on any issues you encounter or work on any contributions or features required. Ory releases all features for free for everyone to use. What is not free however is our time and work. To merge a PR/add a new feature/etc. a significant amount of time is needed to make sure the code lives up to standards, passes all tests, any security implications, etc. This depends on the feature of course, but the one you are alluding to is probably one of those. See the Code of Conduct on OSS support as well: https://github.com/ory/hydra/blob/master/CODE_OF_CONDUCT.md https://github.com/ory/hydra/blob/master/CODE_OF_CONDUCT.md I hope that makes it clearer and feel free to reach out to me directly in the Ory Community on github or slack.
- aeneas_ory 3y agoSorry to hear that this has been your experience! What exactly was the issue for you? It’s true that there are lots of open PRs. We’re a small team and often busy with customer requirements which doesn’t allow us to get some community PRs over the finishing line (finish tests, refactor code, fix remaining bugs, do security reviews, …). Sometimes, PRs are not aligning with an architecture or API principle which is when they often go stale. This is why we generally require design documents for changes or additions to APIs. Saying that the open source is second class is a false accusation in my view: - Over 1500 PRs merged in Ory Kratos alone: https://github.com/ory/kratos/pulls https://github.com/ory/kratos/pulls - Very active contributor and commit frequency: https://github.com/ory/kratos/graphs/contributors?from=2018-05-27&to=2023-08-08&type=c https://github.com/ory/kratos/graphs/contributors?from=2018-... - A growing community and footprint Also, we do offer support contracts for self hosted environments - this is relatively new though: https://www.ory.dev/support/ https://www.ory.dev/support/ It is true though that have to balance open source work and things that people pay us for. It’s the only way to ensure that Ory open source, for which we have a deep commitment, continues for a long time. Hope this makes sense!
- magden 3y agoI can only concur that multi-region apps are becoming the new normal. Deploying app instances across distant locations was never an issue. However, databases used to be the bottleneck. I'm glad to see that changing, thanks to CockroachDB and YugabyteDB. My favorite multi-region deployment mode is geo-partitioned deployment. This is when a database automatically pins user data to specific locations, ensuring low latency for both reads and writes, regardless of user location. One-minute demo how it works: https://www.youtube.com/watch?v=9ESTXEa9QZY&list=PL8Z3vt4qJTkIAYWaUOE_CIntxTHho_pBh&index=14 https://www.youtube.com/watch?v=9ESTXEa9QZY&list=PL8Z3vt4qJT...