4 ms·
>> We ensure that only one copy of every record is read from disk and delivered over the network by including the copy set in the header of every record copy. A
by mleonard 9y ago
>> We ensure that only one copy of every record is read from disk and delivered over the network by including the copy set in the header of every record copy. A simple server-side filtering scheme based on copy sets coupled with a dense copy set index guarantees that in steady state only one node in the copy set would read and delivery a copy of the record to a particular reader.
Can someone explain/expand on the above please? I've read the article a couple of times and tried to understand the above paragraph in context but I don't get it.
- adrianratnapala 9y agoThey've neglected to define the term "copy set". I assumed it was some kind of record ID (although I don't get why that should be distinct from the Log Sequence Number).
- martincmartin 9y agoI'm on the LogDevice team at Facebook. The "copy set" is the subset of server machines that a given record is stored on. For example, if you have a cluster with 10 machines numbered 0 through 9, then the copy set for record 1 might be 3, 6, 8; for record 2 might be 0, 2, 5, etc. Copy set can change with each record. The node set (the list of 10 machines) is fixed for the length of an epoch.
- adriencrth 9y agoAdrien here from the LogDevice team. A copy set is the list of storage node ids that the sequencer chose as recipients for a record. This piece of metadata is stored on storage nodes alongside record payloads in an index that maps the sequence number to copy set mapping. Storage nodes consult the index to filter out payloads they shall not send to readers. The filtering logic consists in sending the copy if the storage node sees itself as the first recipient in the copy set after it's been shuffled using the client id as a seed. This results in readers receiving exactly one copy of the payload.