4 ms·
"The Byzantine generals' algorithm is used to handle situations where the computers do not agree. That situation could come about because of a radiation event c
by martinced 14y ago
"The Byzantine generals' algorithm is used to handle situations where the computers do not agree. That situation could come about because of a radiation event changing memory or register values, for example."
This always got me wondering: what happens if the algorithm used for handling that is itself affected by a radiation event? Is that just too unlikely because such few code is executed? Or what if the computer in charge of verifying that the other computers do come up with a correct answer is itself dying, isn't it a SPOF?
(curious mind wants to know)
- whatshisface 14y agoI think the point is that all the redundant computers check on eachother, and it is very unlikely for the majority of computers to have the exact same fault at the exact same time.
- chickopozo 14y agoThe same hardware, in the same location, running the same software, same power source, next to each other. It's quite likely actually. Do you do your backups to an identical* computer right next to your main one? (* factor in wear)
- shabble 14y agoMulti-version programming[1] (independent implementations of the same specification) is one of the classic solutions to this problem. Likewise for power, location, etc. If you really care about these failure modes, you'll have N different designs of PSU & hardware fed via redundantly pathed links, etc. Aside from the (huge) cost/dev time, the biggest issue is that you still can't protect against logic errors in the specification, and the difficulty in testing every sequence of failure modes across implementations. [1] https://en.wikipedia.org/wiki/N-version_programming https://en.wikipedia.org/wiki/N-version_programming
- atsaloli 14y agoYou can protect against design errors through formal logic verification of the model. Www.spinroot.com
- ersii 14y agoAnd for those who think the link is unrelated spam, it's not. Description provided on spinroot.com: "Spin is a popular open-source software tool, used by thousands of people worldwide, that can be used for the formal verification of distributed software systems. The tool was developed at Bell Labs in the original Unix group of the Computing Sciences Research Center, starting in 1980."
- atsaloli 14y agoThank you, ersii. I was wondering why my karma ticked down.
- wyager 14y agoGenerally speaking, BGA solutions to fault-tolerance work by distributing everything across multiple computers. So every computer compares results with every other computer, and if a minority of computers find themselves in disagreement with a majority, the computers in the minority adopt the stance of the majority.
- DanielRibeiro 14y agoA good discussion about Fault Tolerant algorithms can be found on The Next 700 BFT Protocols[1] [1] http://citeseerx.ist.psu.edu/viewdoc/download?doi=10.1.1.164.2304&rep=rep1&type=pdf http://citeseerx.ist.psu.edu/viewdoc/download?doi=10.1.1.164...
- SeanDav 14y agoI don't believe this addresses the point the OP was trying to make. There is an algorithm or just a simple piece of code which does the distributing, what safeguards this algorithm? Another way of putting it is "who watches the watchers?".
- wyager 14y agoThat's the whole point. There is no central authority; all computers are on the same level. Each computer does the calculations independently, and then they compare results. The computers in the majority either get the minority computers to accept the results or they somehow override the minority computers. I'm not sure how they negotiate, for example, low-level hardware access, but that's just an implementation issue. Take a look at the Bitcoin protocol. It's a very impressive example of many different computers agreeing upon something precisely, despite the fact that many parties have a vested interest in disrupting the Bitcoin network.