5 ms·
Facebook has a similar system, unsurprisingly. An agent runs on every host to sample 1 in N packets, accumulating counts by source/dest host/cluster/service/con
by cranekam 5y ago
Facebook has a similar system, unsurprisingly. An agent runs on every host to sample 1 in N packets, accumulating counts by source/dest host/cluster/service/container/port and so on before sending the aggregate data to Scuba [0] for ad-hoc analysis. This tool was really useful — in a matter of seconds you could see traffic types and volumes broken down by almost any dimension. Did service X see a huge jump in traffic last week? From where? Which container or service? How much bandwidth did your compression changes save? And so on. It also had some really neat stuff to identify whether or not a flow was TLSed or not, which was crucial for working out what still needed to be encrypted in light of the Snowden revelations.
TCP retransmits were also sampled in detail. Being able to see at a glance if e.g. every host in a rack is the source of (or target for) retransmits made troubleshooting much faster.
These systems were really awesome and a good example of what could be built with relatively little effort (the teams involved were small) when you already have great components for reliable log transmission and flexible data analysis.
[0] https://research.fb.com/publications/scuba-diving-into-data-at-facebook/ https://research.fb.com/publications/scuba-diving-into-data-...
- hintymad 5y ago> TCP retransmits were also sampled in detail. Hosts usually collects SNMP metrics, which includes TCP retransmits and more. Do you know what SNMP was lacking compared to eBPF? What I can think of is more dimensions in eBPF's case.
- ikiris 5y agoyou're comparing apples and pumpkins. SNMP is a query mechanism, eBPF is the sampling mechanism.
- hintymad 5y agoAha! I was too removed from the infrastructure then. All I knew was that SNMP metrics showed up in our telemetry system, and we got a set of standard metrics to look into. Our platform team took care of having sampling agent in place. I didn't know the protocol was about querying instead of sampling.
- cranekam 5y agoCorrect me if I'm wrong but the SNMP retransmit counters are just that: a count of retransmits the host sent. Raw retransmit counts are often just a vague indication that something's up — the host is retransmiting, either because there's loss on the path or the receiver is overloaded. But given a host can talk over many paths to many other hosts a raw count isn't specific enough to be useful. The system Facebook (which predated eBPF via a custom ftrace event) produced, effectively, tuples of `(src_ip, src_port, dst_ip, dst_port, src_container, dst_container, ...)` and aggregated them over all hosts. This allowed counting retransmits by, say, receiving host. If there's one host that has a bad cable and is receiving retransmits from 1000 clients we may not see that signal in simple counters on the clients — for them it's just a tiny bump in the overall retransmit rate. But if we aggregate by receiving host the bad guy will stand out like a sore thumb. Same thing for all the hosts in a rack, or all hosts reachable over a given router interface, or whatever else you want. One of my common workflows when facing a bump in general errors (e.g. timeouts to the cache layer) was to quickly try grouping retransmits by a few dimensions to see if one particular combination of hosts stood out as sending or receiving more retransmits. tl;dr: the SNMP data is one-dimensional. FB's system allowed aggregating and querying by many dimensions. This is really useful when there are thousands of machines talking to each other over many network paths.
- paulfurtado 5y agoDoing it with eBPF gives you a hook for each retransmit, this makes it possible to know the exact connection, process, and network interface that hit the retransmit and allows you measure things like "100s of hosts are hitting 1000s of retransmits to 10.0.0.56:443", which 10.0.0.56's netstat metrics may not clearly indicate. It gets more interesting if you break things down by VM hosts, racks, rows, data centers, etc. If you go deeper with the eBPF tracing, you can also determine which code path the retransmit occurred on, which may or many not be interesting.
- Hikikomori 5y agoeBPF can give detailed stats and TCP state information on a per connection basis (flow), much more powerful than TCP stats you can grab with SNMP that are aggregated.
- takeda 5y agoeBPF similarly to DTrace allows you to view internals of your OS. With SNMP you have fixed number of metrics that you can view, while with eBPF you can create new ones. You could implement SNMP daemon that uses eBPF to get the data, and perhaps that will happen in the future if it already didn't.
- hintymad 5y agoThanks for all the answers. I learned so much!