5 ms·
Yeah. I understand the problem that TCP is trying to solve. My thinking was that there is a whole class of network applications that don't work great under TC
by Nrsolis 11y ago
Yeah. I understand the problem that TCP is trying to solve.
My thinking was that there is a whole class of network applications that don't work great under TCP because the retransmission mechanism drives latency super high when congestion is experienced in the network. TCP works to back off transmission that might be the cause of packet loss, but those packets still have to be transmitted again and the bandwidth timeslot consumed by the dropped packet amounts to wasted resources in the network. In a datacenter microburst situation, retransmissions just tickle the problem you're trying to avoid.
FEC over a lossy network would allow for a slight increase in the baseline resource utilization with the side benefit of eliminating retransmissions which might be beneficial in a microburst environment. I don't really know, I'm just thinking aloud.
- raattgift 11y agoConsider three conditions: - high link-layer integrity; packet loss is mainly due to congestion - low packet loss; most flows unaffected - bursty packet loss; individual flows, where they see any loss at all, have a non-negligible probability of multiple back-to-back packets lost (e.g. in FIFO queue tail drop, or a dynamic routing transient like a black hole or brief (~ ttl) forwarding loop). Consider the limit in which the number of simultaneous flows becomes large. In that limit, there is a lot of additional FEC data that has been generated, successfully transported across a network, and ultimately discarded as redundant. There is also FEC data that fails to help in the reconstruction of flows which experience a multi-packet loss too great for the erasure code being used. We can calculate wasted power in this limit, and in some situations that waste can be significant. One situation is when the FEC-generator is sending a large number of flows that stay near steady-state equilibrium amongst themselves and with other long-lived flows, which is reasonable to expect when multi-minute adaptive streaming video is the dominant component of the traffic. Conversely a transmission that is sufficiently short that it never enters into the moral equivalent of congestion avoidance, and where the application is also latency-sensitive, small amounts of FEC in the early parts of conversations might be a win over retransmissions, depending on how often the FEC forms part of loss recovery. In the modern Internet, the cost-benefit analysis of the extra power the extra FEC requires is probably strongly time-dependent, and it may be counter-intuitive. (e.g., if FEC doubles initial traffic, it may exacerbate any burstiness, which may in turn lead to greater burst synchronization; that is, in a microburst environment the additional FEC data, unless it's somehow congestion-avoiding (hard to do for very short flows) it might be the opposite of beneficial). It's probably also worth considering differences in FEC receivers, for instance whether they are drawing power from a small battery or not. Consider a toy question like, "What drains more battery energy: two hours of receiving FEC or two and a bit hours of receiving non-FEC?" "A bit" may matter. It might also be dynamically calculable: lots of retransmits? Transition to sending FEC proactively. No FEC use? Transition to reactive packet loss recovery. Anyway, it's cool to get even a little look at what Google's learned from its experiments. Everyone should encourage the sharing of these sorts of details, and we should probably encourage other large scale entities with different types of traffic (e.g. non-YouTube Google; Facebook; Twitter) to do likewise.