4 ms·
1. On a short time horizon, not as sure, back of the napkin, it took ~12 dev months (5 months with 2.5 people average on it). However, our cost per 1000 msgs/se
by addisonj 7y ago
1. On a short time horizon, not as sure, back of the napkin, it took ~12 dev months (5 months with 2.5 people average on it). However, our cost per 1000 msgs/sec is much lower (like 1/4 the cost of Kinesis) so we fully expect that investment to pay off over time assuming that adoption by the rest of the org continues and we don't find a ton of issues.
1a. You are correct we didn't require geo-replication for existing use cases, however, initially, we saw geo-replication as an easy way to improve DR and we have an internal requirement for a DR zone in another region. Now that we have done the work, we are starting to see multiple places where we can simplify some things with geo-replication, so we think long term the feature will be really valuable
2. We split up auth into two main components: auth of users (where we use Okta) and auth of services. For okta, we just wrote a small webapp that users can log into via OKta and generate credentials. For apps/services, we already had hashicorp in place and wanted to just piggyback of our existing form of identity (IAM roles). Essentially, a user just associates an IAM role with a pulsar role and we generate and drop off credentials into a per-role unique shared location in vault that any IAM role can access (across multiple AWS accounts)
3. Once again, geo-replication wasn't really a hard requirement initially but more of something that we really like now that we have. I think the biggest reason why not postgres is that we have combined message rates (not everything is migrated yet) on the order of 300k msgs/sec across a few dozen services. Pulsar is designed to scale horizontally and also has really great organizational primitives as well as an ecosystem of tools. While I think you could maybe figure that out with some PG solution, having something purpose built really can pay big dividends for when you are trying to make a solution that can easily integrate into a complex ecosystems of many teams and many different apps/use cases
- ignoramous 7y agoAgreed. One more: For replication across regions, do you peer VPCs via Transit Gateways or some such, or do it over the public Internet? I ask because a lot of folks complain about exorbitant AWS bandwidth charges for cross-AZ and cross-region communication (esp over the Internet versus over AWS' backbone): At 300k msgs/sec, the bandwidth costs might add up quickly? Consequently, maintaining a multi-region, multi-AZ VPC peering might have been complicated without Transit Gateway, so I'm curious how the network side of things held up for you.
- addisonj 7y agoIn this case, we use just straight VPC peering with a full mesh of all our regions. We may eventually migrate to being built on our VPN based mesh (we do that in other places) Bandwidth is certainly a concern and that is one of the nice bits about Pulsar is not everything is replicated. You mark a namespace by adding additional clusters it should replicate to. We don't expect to replicate everything, just the things teams care about. When we did this, Transit Gateway was just within the same region. At re:invent they announce the cross region transit gateway which we will look at moving to as well, but for now, it is just a full mesh of VPC peers, which for 8 regions isn't bad... but certainly gets worse with each new region we need to add. For exposing the service into other VPCs in the same region we use private-link endpoints as to avoid needing to do even more peering.