4 ms·
I should get better stats. 4.5M users, 6 nodes split between 3 datacenters. Each node has 16 CPUs. 2158341079 total write transactions (over 8 years); not su
by quinthar 10y ago
I should get better stats. 4.5M users, 6 nodes split between 3 datacenters. Each node has 16 CPUs. 2158341079 total write transactions (over 8 years); not sure how many read (10-100x more?). I'll try to get better stats on peak read/write transactions per second. Not sure the total number of rows of the largest table (that actually takes a long time to count).
Counting "failures" is difficult -- if it breaks, it's because of some bug in our application logic (eg, our stored procedure). Most of our restarts are due to normal maintenance and upgrades. In the history of the company there were a handful of core problems to the logic (generally as we encountered some weird edge case for the first time), but I can't remember the most recent.
Regardless, I'll try to get better data on this. Thanks for asking!
- ngrilly 10y ago> 6 nodes split between 3 datacenters This makes 2 nodes per datacenter. Do the 2 nodes cooperate in some way, or are they fully independent from each other (like shards)?
- ngrilly 10y agoJust found the answer to my question: > In practice, we deploy two servers in each datacenter (across three geographically-distributed datacenters, so six nodes total in the cluster -- with one configured as a "permaslave" that doesn't participate in quorum). Given this, the node that's in the same datacenter generally gets most if not all of the master's transactions in a real world crash scenario. Source: http://p2p-hackers.709552.n3.nabble.com/p2p-hackers-Advice-on-concurrent-relational-database-writes-td4025323.html http://p2p-hackers.709552.n3.nabble.com/p2p-hackers-Advice-o...
- spotman 10y ago2 billion writes over 8 years? Cool if so, but many database systems both MySQL, and others are capable of getting this in a day. I know this because I maintain some high availability MySQL systems that see about 4 billion writes per day across a similar amount of hardware. Have you run this through Jespen or done any actual load testing or deep testing for failures related to machines dying, network partitions or other? I encourage you to do so, because (not trying to sound rude) this kind of reads like "hey my homemade car has 25,000 miles on it and I still use it". But you can build a performance and failure testing framework to really put it through its paces. Anyhow , have fun and good luck~~
- Jweb_Guru 10y agoWhile I agree that we're not talking about huge write volume in the grand scheme of things, and I definitely agree with you that Bedrock should pay aphyr to run this through Jepsen (particularly with their new replication strategy--see some of my concerns below), Expensify's overall transaction volume is probably close to the limit of what 95% of businesses are ever going to see, they have contractual latency and availability requirements they have to fulfill that are also probably more stringent than 95% of businesses are ever going to see, and they've been doing it with a single master database (no partitioning) and a single writer on that master. I think it's useful information that SQLite works for them, because it means SQLite would work fine for most people and (as is evidenced by this thread) a lot of people don't think SQLite can work for a site like Expensify. And seeing actual numbers are important, too. People often act like in terms of write volume you're either shared hosting fodder, or you're Facebook, but there's a lot of room in between, and a lot of HA workloads are very read-dominated. It's very hard to find information about the businesses fall in the middle of that spectrum because usually no technology company talks about its traffic unless it's bragging about it (and then will release only vague information that could be spun in lots of different ways--for instance, not to call you out, I have no clue what "HA" means in your context, whether you need distributed transactions, whether the writes are cleanly partitionable, what isolation level you need, what your latency requirements are, or what the read:write ratio looks like). So I'd rather encourage people to release their numbers than poo-poo them because they aren't as big as whatever the biggest system you've worked on is.
- spotman 10y agoMy point is that if they are advertising rock solid distributed data, I don't think they have achieved that without putting it through some more paces. I agree, actual numbers are great. We also agree, jespen is great. But I am saying, another path, to a similar level of confidence jespen provides, is some real critical stress testing. Sure, it will be out of the scope of what most people need, but it will give them the confidence to say "rock solid distributed data".
- justinclift 10y agoSome of these FIXME's in Bedrock's er... Paxos implementation look like they could be important: https://github.com/Expensify/Bedrock/blob/ecda922dc279e06fdad23b5173d5b0d207eba8e7/sqlitecluster/SQLiteNode.cpp#L14-L32 https://github.com/Expensify/Bedrock/blob/ecda922dc279e06fda...