4 ms·
I'm not surprised that Cassandra doesn't perform ideally with only three nodes, considering the scales it's intended for. Does anybody know how many nodes are r
by apike 16y ago
I'm not surprised that Cassandra doesn't perform ideally with only three nodes, considering the scales it's intended for. Does anybody know how many nodes are required for its resiliency safeguards to work properly?
- jbellis 16y agoIt wasn't that the resiliency safeguards didn't work, so much as reddit was simply underprovisioned. I have no idea what TFA meant by "at only three nodes we weren't able to take advantage of most of Cassandra's safeguards for ending up in this situation," except in the sense that adding an extra 20k ops/s on a 30 node cluster will be 10% of capacity instead of 100%. (Just picking reasonable ballpark numbers, I don't know what reddit's are.)
- stephenjudkins 16y agoIf "to work properly" you mean that a cluster can suffer the loss of one node with no data loss, you only need two nodes. The consistency guarantees Cassandra gives you are based upon configuration values (ReplicationFactor) and a ConsistencyLevel provided at runtime when performing an operation. A perfectly available, partition-tolerant and perfectly consistent system is impossible so adjusting these settings lets you specify the tradeoffs you desire. If for all operations you specify a consistency level of ALL, you are guaranteed consistency. However, you lose availability if one server (for a given part of a keyspace) goes down. Conversely, you can also use Cassandra in a way that offers few consistency guarantees by specifying a ConsistencyLevel of ANY, or ONE. Further, you can perform reads with ConsistencyLevel ONE that don't give you a guarantee that the information you read is consistent. As long as no servers go down, you have consistency, and as long as you have one server left, you have complete availability. The QUORUM consistency level guarantees at least (ReplicationFactor/2 + 1) nodes agree on a write and gets the latest timestamp from a majority of servers on read. This offers both consistency and availability. For three servers with a ReplicationFactor of 2 this is the same as ALL; if your ReplicationFactor is 1 you lose effective consistency guarantees. So, to be able to use the full range of options in the consistency/availability tradeoff space one needs a cluster of at least four servers. Someone let me know if my reasoning is incorrect. See http://wiki.apache.org/cassandra/API http://wiki.apache.org/cassandra/API for more info.
- megablast 16y agoIt seems odd that it is configured to look up a key-pair, when that key-pair is no longer needed. Surely it would better to have it no longer cache queries that are no longer needed.
- jbellis 16y agoIt's like how in postgresql, if a client runs "select * from lots", and you kill the client, the query keeps going even though there's nobody to hand the answer to. That said, there are ways we can mitigate this, primarily in https://issues.apache.org/jira/browse/CASSANDRA-685 https://issues.apache.org/jira/browse/CASSANDRA-685