3 ms·
I think it's just saying that it's willing to place two of the three replicas in a datacenter, for example if one of the three datacenters is down. This has do
by voidmain 8y ago
I think it's just saying that it's willing to place two of the three replicas in a datacenter, for example if one of the three datacenters is down. This has downsides, since losing a datacenter will make it aggressively fill up disks, but mitigates against subsequent failures causing data loss.
Most of the people who have run FoundationDB at scale have, for performance reasons, used configurations other than the "datacenter aware" mode for their inter region replication, so they may not be the strongest thing operationally.
There is some work that from what I can see in the code is still in progress to build a new, almost magical inter-region replication mode that I am very excited about, which combines synchronous replication to a "satellite" datacenter within region with asynchronous replication between regions and recovery logic that will finish replication and fail over in case of a partial failure of a region. You get fast transaction commits (much less than the inter region ping time), can fail over to a secondary region automatically and safely (without losing any committed transactions) in the vast majority of circumstances, and in the worst case you can (manually, because you are accepting data loss!) give up very recently committed transactions to fail over.
- wll 8y agoHow would FoundationDB stay externally consistent with asynchronous cross-region replication? Thank you for your time and FoundationDB—along with @nlavezzo, and team(s)!
- voidmain 8y agoThe satellite mode that I described is an active/passive mode. One region is accepting reads and writes; the other is just replicating everything. When it looks like the active region is in trouble, the asynchronous replication is "finished up" before switching over to the other region. The multiple datacenters in each region ensure that usually a regional failure will be "slow enough" that this automatic process (which after all only takes hundreds of milliseconds to seconds) can usually complete before a region goes away. And this will be handled pretty transparently by the datastore. If a region is blown up instantly by an orbital laser cannon, then the database will go down and you will have to manually tell it to recover ACI in the other region, sacrificing the durability of whatever committed transactions in the lost region were destroyed by the laser cannon.
- wll 8y agoWhile that’s perfect to shield from orbital laser cannons, is active/active geo-independent replication possibile?
- voidmain 8y agoWell, if you want ACID then you are going to have to pay for at least one geographic round trip per committed transaction. (So why not go active/passive, and have at least one of your datacenters be fast?) But what if you have different pieces of data and you want them to be fast in different datacenters? I think a great solution to this can be layered on top of multiple FoundationDB clusters, each using the satellite mode, but this is one thing that I at least haven't been able to think of a way to provide properly at the data model agnostic key/value store layer - the details about what to put where seem fundamentally dependent on your data model.
- wll 8y ago> So why not go active/passive, and have at least one of your datacenters be fast? While local writes would stay fast, wouldn’t active/passive see higher-latency non-local writes than Spanner or Fauna’s (assuming a NAM-EUR-ASIA topology)? I agree with and do appreciate the multiple FoundationDB clusters suggestion.
- voidmain 8y agoI'm speculating, but I think in this mode, from the "slow" datacenters you would see one round trip time to start a transaction, then reads will be fast (they can be done safely from your local datacenter because of MVCC), and then one round trip time to commit the transaction. I think that's as good as Spanner does with the same geography, but I'm not sure. I think you could get rid of the first round trip time even without any clock synchronization nonsense, by speculating on a read version for read/write transactions. And 1xRTT is obviously as fast as physically possible for ACID.
- 8y ago