4 ms·
> Raft is a consensus algorithm that is designed to be easy to understand. > Consensus typically arises in the context of replicated state machines, a general
by 63 3y ago
> Raft is a consensus algorithm that is designed to be easy to understand.
> Consensus typically arises in the context of replicated state machines, a general approach to building fault-tolerant systems.
I recognize that I'm not the intended audience but I do think I would be a lot more capable of understanding this article if it used less jargon or at least defined what it meant. I'm only mentioning this because ease of understanding is an explicit goal.
Can someone give a real world example of where this would be used in a production app? I'm sure it's very practical but I'm getting caught up in trying to understand what it's saying
- galenmarchetti 3y agoConsensus algorithms underly distributed databases like Cassandra, or distributed caches like Redis. (in fact Redis does use a version of Raft). Simple use case, you're running a database across three servers, so that users can ping any one of the servers and it still works (fault tolerance). User X pings server A to write something to a table. User Y pings server B to read from that table. How do you make sure that User Y reads what User X wrote, so that User X and User Y are on the same page with what happened? That's the consensus algo
- couchand 3y agoNot to be too pedantic, but the step of actually shipping those changes across servers is usually outside the scope of consensus algorithms. Generally they are limited to picking a single server to act as a short-term leader. That server then is responsible for managing the data itself. Though you can conceive of a system where all data flows through the consensus algorithm, practically speaking that would introduce significant overhead at a granularity where it isn't adding value. There isn't neceasarily one dictator for the whole cluster, but rather usually they are scoped to each domain that requires serialization.
- adunk 3y agoWikipedia has a nice list of software using the algorithm: https://en.wikipedia.org/wiki/Raft_(algorithm)#Production_use_of_Raft https://en.wikipedia.org/wiki/Raft_(algorithm)#Production_us...
- oriettaxx 3y agoapparently there is no reference to Docker in swarm mode, where it is used to decide what node should be the Leader
- uw_rob 3y agoThink databases which run across many different machines. Distributed databases are often conceptually modeled as a state machine. Writes are then mutations on the state machine. With a starting state (empty database), if everyone agrees that on a fixed list of mutations which are executed in a rigidity defined ordering, you will get the same final state. Which makes sense, right? If you run the following commands on an empty database, you would expect the same final state: 1. CREATE TABLE FOO (Columns = A, B) 1. INSERT INTO FOO (1, 2) 1. INSERT INTO FOO (3, 4) Which would be: ``` FOO: |A|B| |-|-| |1|2| |3|4| ``` So where does "consensus" come into play? Consensus is needed to determine `mutation 4`. If the user can send a request to HOST 1 saying 'Mutation 4. should be `INSERT INTO FOO (5, 6)`' then HOST 1 will need to coordinate together with all of the other hosts and hopefully all of the hosts can agree that this is the 4th mutation and then enact that change on their local replica. This ordering of mutations is called the transaction log. So, why is this such a hard problem? Most of the reasons are in [Fallacies of distributed computing](https://en.wikipedia.org/wiki/Fallacies_of_distributed_computing https://en.wikipedia.org/wiki/Fallacies_of_distributed_compu...) but the tl;dr is that everything in distributed computing is hard because hardware is unreliable and anything that can go wrong will go wrong. Also, because multiple things can be happening at the same time in multiple places so it's hard to figure out who came first, etc. RAFT is such an algorithm to let all of these hosts coordinate together in a fault tolerant way to figure out the 4th mutation. Disclaimer: The above is just one use of RAFT. Another way RAFT is used in distributed databases is as a mechanism for the hosts to coordinate a hierarchy of communicate among themselves and when a host in the hierarchy is having problems RAFT can be used again to figure out another hierarchy. (Think consensus is reached on leader election to improve throughput)
- dcchuck 3y agoSimplest terms - your app needs several different services or instances of a service to agree upon a value. There are a lot of reasons you can't use things like the system clock to agree upon when something happened for example - this is where RAFT steps in. You'll see "fault-tolerant" and "replicated state machines" often alongside them. Let's break those down in this context. For "fault-tolerance" - think production environments where I need to plan for hardware failure. If one of my services goes down, I want to be able to continue operating - so we run a few copies of the app, and when one goes down, another operating instance will step up. In that case - how do we pick what's in charge? How do all copies agree on things while everything is working smoothly? Raft. For "replicated state machines" - let's stay in this world of fault-tolerance, where we have multiple instances of our app running. In each service, could reside a state machine. The state machine promises to take any ordered series of events, and always arrive at the same value. Meaning - if all of our instances get the same events in the same order, they will all have the same state. Determinism. This is where it all comes together, and why I think the jargon becomes tightly coupled to an "easy to understand" definition. You will reach for replicated state machines when you need deterministic state across multiple service instances. But the replicated state machines need a way to agree on order and messages received. That's the contract - if you give everything all the messages in the same order, everything will be in the same state. But how do we agree on order? How do we agree on what messages were actually received? Just because "Client A" sends a messages "1", and "2', in a specific order does not guaranteed it is delivered at all, let alone in that order. Raft creates "consensus" around these values. It allows the copies to settle on which messages were actually received and when. So, you could use other approaches to manage "all your service copies getting along" but a replicated state machine is a nice approach. That replicated state machine architecture needs some way to agree on order, and Raft is a great choice for that.
- bjornasm 3y agoMade a similar comment as this. If it is for a wider audience etc you would think it would be beneficial to explain it in a way that the wider audience understands.