4 ms·
Recipe: Create a global Cassandra cluster with regional datacenters. Use one keyspace per region Use per-keyspace replication to only replicate that region's
by tupshin 10y ago
Recipe:
Create a global Cassandra cluster with regional datacenters.
Use one keyspace per region
Use per-keyspace replication to only replicate that region's data locally, and to one or more additional datacenters
Have stateless app servers colocated with Cassandra in each DC handling all local traffic
Run spark on top of Cassandra to do analytics, or to do the etl to a dedicated analytics system
Optionally have a single "master" DC, with replicas of all data from all locations, that doesn't serve end user traffic, but is to allow efficient cross region analytics.
Profit (optional step)
And yes, the company I work for (Datastax) has a product and services to help make it simple.
- scaleout1 10y agoThats an interesting approach. One question though, in our use case user often travel from city to city and country to country. How do you model that if you are only using local DC and local replications?
- methehack 10y agoI think the regional keyspaces would be have to be caches -- denormalize it, basically. Pop/re-fresh people into the geo-based caches as they moved around. Truth sits behind it, centralized (perhaps partitioned in some way that makes sense globally but is sub-optimal from a regional cache perspective). Might not be worth it -- hard to know from here. :)
- ajslater 10y agoafaik, this is how Facebook does it, but with regional sources of truth. If you signed up for FB in Paris and move to San Francisco, your master profile lives in Europe in perpetuity and you'll use your regional cache forever in the USA. The number of people moving far away from their home DC's should be a reasonably small fraction of the total for it not to matter.