3 ms·
Re: SQLite -- Yes, Bedrock is just the replication layer, SQLite does all the SQL and storage. Re: Paxos -- It's our own implementation. Split brain is preven
by quinthar 10y ago
Re: SQLite -- Yes, Bedrock is just the replication layer, SQLite does all the SQL and storage.
Re: Paxos -- It's our own implementation. Split brain is prevented by the master refusing to stand up unless a majority of configured (not just active) slaves approve its standup request. So in a 6 node deployment, 1 master and 5 slaves, 3 of the slaves would need to approve. This means one "half" would have 4 nodes, and the other would have 2 -- and the half with 2 would recognize it doesn't have quorum and thus stay idle. In a scenario where there is a split down the middle, nobody would do anything because nobody has quorum. (This is precisely why you shouldn't deploy in just 2 datacenters -- three is the magic number.)
Re: Sticky sessions -- True multi-machine consistency is impossible to guarantee and has no real practical application. The only consistency that matters is from the perspective of an observer, ensuring that it always sees a world whose time arrow progresses linearly forward. If that is what "linearizable" means, then perhaps. Sorry, my terminology could be off. Thanks for the correction!
Re: SQLite on SSD -- Sorry, I didn't mean to claim Bedrock was somehow faster than SQLite. Indeed, Bedrock adds overhead to SQLite (to do all the networking and such). My main claim is that Bedrock is faster than other databases, in particular when using C++ stored procedures for complex operations.
Thanks for the questions!
- toolslive 10y agoIs the set of servers that participate a paxos value too (ie determined using the consensus algorithm), or is it configured ? Can you grow a cluster without bringing it down ?
- quinthar 10y agoUnfortunately you can't add new nodes without reconfiguring the cluster, and currently that requires restarting each server (though so long as you don't do it all at once, server restarts are normal maintenance that causes no downtime to the end user). This could likely be added without too much effort, but also the real-world use case of this is uncertain. The consensus algorithm is only used to elect a master -- once the master is identified, it coordinates all the distributed transactions. If the master dies, everyone who remains elects a new master and re-escalates all unprocessed write transactions to the new master. Read transactions are processed locally by each node, so only writes need to be escalated.
- msackman 10y agoPaxos: I've been trying to look for that. Having cloned the code and grepped for paxos I'm getting no hits. Where is the paxos implementation?
- quinthar 10y agoIt's not called out specifically (actually, when writing it I didn't even know what Paxos was, and only realized I had implemented it years later). However, the logic is here: https://github.com/Expensify/Bedrock/blob/master/sqlitecluster/SQLiteNode.cpp https://github.com/Expensify/Bedrock/blob/master/sqliteclust...
- justinclift 10y agoWith these: https://github.com/Expensify/Bedrock/blob/ecda922dc279e06fdad23b5173d5b0d207eba8e7/sqlitecluster/SQLiteNode.cpp#L14-L32 https://github.com/Expensify/Bedrock/blob/ecda922dc279e06fda... Do you use the Issue Tracker to keep on top of things like that, and/or prioritise?
- quinthar 10y agoWe're using GitHub Issues: https://github.com/Expensify/Bedrock/issues https://github.com/Expensify/Bedrock/issues However, honestly those specific issues aren't on the list. In general we focus less on what could happen, and more on what actually does happen. Those specific issues haven't ever occurred, and thus never got "fixed" because they never became real problems. But PRs welcome!!