4 ms·
It's hard to understand what tigerbeetle is about. Can anyone ELI5 it for me? As far as I can tell, it's some kind of a library/system geared at distributed tra
by rosetremiere 4y ago
It's hard to understand what tigerbeetle is about. Can anyone ELI5 it for me?
As far as I can tell, it's some kind of a library/system geared at distributed transactions? But is it a blockchain, a db, a program ? (I did look at the website)
- deleted 4y ago[deleted]
- laserbeam 4y agoIt's a special purpose DB. No relation to blockchains.
- eatonphil 4y agoHey thanks for the feedback! We've got concrete code samples in the README as well [0] that might be more clear? It's a distributed database for tracking accounts and transfers of amounts of "thing"s between accounts (currency is one example of a "thing"). You might also be interested in our FAQ on why someone would want this [1]. [0] https://github.com/tigerbeetledb/tigerbeetle#quickstart https://github.com/tigerbeetledb/tigerbeetle#quickstart [1] https://docs.tigerbeetle.com/FAQ/#why-would-i-want-a-dedicated-distributed-database-for-accounting https://docs.tigerbeetle.com/FAQ/#why-would-i-want-a-dedicat...
- rosetremiere 4y agoThe faq helped, thanks! So, an example of typical use would be, say, as the internal ledger for a company like (transfer)wise, with lots of money moving around between accounts? But I understand it's meant to be used internally to an entity, with all nodes in your system trusted, and not as a mean to deal with transactions from one party to another, right?
- eatonphil 4y agoYes that's a good example! And you can model external accounts that have their own confirmation process using our two-phase transfer support. https://docs.tigerbeetle.com/FAQ#what-is-two-phase-commit https://docs.tigerbeetle.com/FAQ#what-is-two-phase-commit
- jorangreef 4y agoGreat to hear! Joran from TigerBeetle here. Yes, exactly. You can think of TigerBeetle as your internal ledger database, where perhaps in the past you might have had to DIY your own ledger with 10 KLOC around SQL. And to add to what Phil said, you can also use TigerBeetle to track transactions with other parties, since we validate all user data in the transaction—there are only a handful of fields when it comes to double-entry and two-phase transfers between entities running different tech stacks. The TigerBeetle account/transfer format is meant to be simple to parse, and if you can find user data that would break our state machine, then it's a bug. Happy to answer more questions!
- ngrilly 4y agoIt's a distributed database for financial transactions, using double entry accounting, written in Zig, and with a very innovative design: - LMAX inspired - Static memory allocation - Zero copy with Direct I/O - Zero syscalls with io_uring - Zero deserialization - Storage fault tolerance - Viewstamped Replication consensus protocol - Flexible Quorums - Deterministic simulation like FoundationDB
- Yoric 4y agoZero deserialization? That sounds rather scary. This means absolute trust in data read from disk or received from other nodes?
- rom-antics 4y agoWhat is the threat model you're worried about? If an attacker can write data to your disk or authenticate to your cluster, aren't you already screwed?
- Yoric 4y agoYes, these are exactly my threats. First, because I'm a strong believer in defense-in-depth. Secondly because both disk corruption and network packet corruption happen. Alarmingly often, in fact, if you're operating at large scale.
- jorangreef 4y agoOurs too! For example, our deterministic simulation testing does storage fault corruption up to the theoretical limit of f according to our consensus protocol. Details in our other reply to you.
- jorangreef 4y agoGreat question! Joran from TigerBeetle here. "This means absolute trust in data read from disk or received from other nodes?" TigerBeetle places zero trust in data read from the disk or network. In fact, we're a little more paranoid here than most. For example, where most databases will have a network fault model, TigerBeetle also has a storage fault model (https://github.com/tigerbeetledb/tigerbeetle/blob/main/docs/DESIGN.md#safety https://github.com/tigerbeetledb/tigerbeetle/blob/main/docs/...). This means that we fully expect the disk to be what we call “near-Byzantine”, i.e. to cause bitrot, or to misdirect or silently ignore read/write I/O, or to simply have faulty hardware or firmware. Where Jepsen will break most databases with network fault injection, we test TigerBeetle with high levels of storage faults on the read/write path, probably beyond what most systems, or write ahead log designs, or even consensus protocols such as RAFT (cf. “Protocol-Aware Recovery for Consensus-Based Storage” and its analysis of LogCabin), can handle. For example, most implementations of RAFT and Paxos can fail badly if your disk loses a prepare, because then the stable storage guarantees, that the proofs for these protocols assume, is undermined. Instead, TigerBeetle runs Viewstamped Replication, along with UW-Madison's CTRL protocol (Corruption-Tolerant Replication) and we test our consensus protocol's correctness in the face of unreliable stable storage, using deterministic simulation testing (ala FoundationDB). Finally, in terms of network fault model, we do end-to-end cryptographic checksumming, because we don't trust TCP checksums with their limited guarantees. So this is all at the physical storage and network layers. "Zero deserialization? That sounds rather scary." At the wire protocol layer, we: * assume a non-Byzantine fault model (that consensus nodes are not malicious), * run with runtime bounds-checking (and checked arithmetic!) enabled as a fail-safe, plus * protocol-level checks to ignore invalid data, and * we only work with fixed-size structs. At the application layer, we: * have a simple data model (account and transfer structs), * validate all fields for semantic errors so that we don't process bad data, * for example, here's how we validate transfers between accounts: https://github.com/tigerbeetledb/tigerbeetle/blob/d2bd4a6fc240aefe046251382102b9b4f5384b05/src/state_machine.zig#L867-L952. No matter the deserialization format you use, you always need to validate user data. In our experience, zero-deserialization using fixed-size structs the way we do in TigerBeetle, is simpler than variable length formats, which can be more complicated (imagine a JSON codec), if not more scary.