4 ms·
Excuse me while I shill for my employer but we're indeed big fans of BPF at Facebook. Our L4 load balancer is implemented entirely in BPF byte code emitting C+
by bdd 7y ago
Excuse me while I shill for my employer but we're indeed big fans of BPF at Facebook.
Our L4 load balancer is implemented entirely in BPF byte code emitting C++ and relies on XDP for "blazing fast" (comms approved totally scientific replacement for gbps and pps figures...) packet forwarding. It's open source and was discussed here at HN before https://news.ycombinator.com/item?id=17199921 https://news.ycombinator.com/item?id=17199921.
We discussed how we use eBPF for traffic shaping in our internal networks at Linux Plumber's Conference http://vger.kernel.org/lpc-bpf2018.html#session-9 http://vger.kernel.org/lpc-bpf2018.html#session-9
We presented how we enforce network traffic encryption, catch and terminate cleartext communication, again, you guessed, with BPF at Networking@Scale https://atscaleconference.com/events/networking-scale-3/ https://atscaleconference.com/events/networking-scale-3/ (video coming soon, I think.)
Firewalls with BPF? Sure we have 'em. http://vger.kernel.org/lpc_net2018_talks/ebpf-firewall-LPC.pdf http://vger.kernel.org/lpc_net2018_talks/ebpf-firewall-LPC.p...
In addition to all these nice applications we heavily rely on fleet wide tooling constructed with eBPF to monitor:
- performance (why is it slow? why does it allocate this much?)
- correctness (collect evidence it's doing its job like counters and logs. this should never happen, catch if it does!)
...in our systems.
- davemarchevsky 7y ago> - performance (why is it slow? why does it allocate this much?) One of the pieces of fleet-wide tooling that heavily uses eBPF is PyPerf, which we talked about publicly at Systems@Scale in September ("Service Efficiency at Instagram Scale" - https://atscaleconference.com/events/systems-scale-2/ https://atscaleconference.com/events/systems-scale-2/ - video also coming soon, I think).
- stephen999 7y agoIs there any public code for these perf tools?
- bdd 7y agohttps://github.com/iovisor/bcc/tree/master/examples/cpp/pyperf https://github.com/iovisor/bcc/tree/master/examples/cpp/pype...
- lathiat 7y agowoah I hadn't seen that.. sounds like PySpy but implemented in BPF. That's crazy and cool: https://github.com/benfred/py-spy https://github.com/benfred/py-spy
- dmix 7y agoThe article mentions the servers run up to 100 different BPF related programs on an individual server. I get the firewall and traffic shaping stuff, but are there any unusual use-cases or hacks you could share?
- bdd 7y agoThe number increases because there are various monitoring tools. Stuff like something that extracts more data from TCP retransmits so we have a better understanding of congestions and errors in the network path. ...or things that simply collect certain system events for security event detection. Imagine, for every certain event, what is injected is counted as a separate program. On top of these which is common for every machine, service owners can deploy their own BPF programs for specific use cases. In fact our self service tracing tooling is also a BPF program. We talked about it back in 2014 when it did not use BPF https://tracingsummit.org/w/images/6/6f/TracingSummit2014-Tracing-at-Facebook-Scale.pdf https://tracingsummit.org/w/images/6/6f/TracingSummit2014-Tr...
- aloknnikhil 7y agoHas there been any comparison on if XDP-eBPF packet processing is faster/slower than pure user-space packet processors like Cisco's VPP, especially the ones with support for DPDK and zero-copy? I ask because user space packet processors are extensively used as virtual switches in container deployments. Assuming XDP-eBPF is faster, I wonder if, in combination with namespaces, there could be an efficient "virtual switch" implemented completely in eBPF.
- bdd 7y agoPersonal experience spanning across my time at Twitter and Facebook, DPDK is definitely faster but it is (/was) a pain in the butt to program and share the NIC with the host. XDP allows you to reuse very many packet parsing / handling capabilities in the kernel while in DPDK world you’re shipping a tiny stack with your app. Re: Virtual Switch implemented in BPF: see Cilium’s work to connect containers with their BPF based connectors.
- shaklee3 7y agoThat depends on the nic. Mellanox's bifurcated driver is very easy to share with the kernel. Intel's, not so much. Do you mind mentioning which NICs are used at FB? Also for GP, is anyone using vpp in production?
- aloknnikhil 7y agoPredominantly, some internal teams at Cisco. VPP and its CNI (Ligato Contiv) seem to be picking up steam off late. This is going by activity in the mailing lists. I know Yahoo Japan uses it. http://events19.linuxfoundation.org/wp-content/uploads/2018/07/ONS.NA_.2019.VPP_LB_public.pdf http://events19.linuxfoundation.org/wp-content/uploads/2018/...
- fdee 7y agoThere is not even need for a bifurcated driver. The upstream kernel has AF_XDP which Intel and Mellanox NICs support in their drivers, and DPDK has official integration for it as well: https://doc.dpdk.org/guides/nics/af_xdp.html https://doc.dpdk.org/guides/nics/af_xdp.html This will make deploying DPDK significantly easier for those that need/want to use it and allows to share the same driver for pushing packets up to DPDK and into the normal kernel stack with very close to "native" (as in user space driver) DPDK performance (target is ~90-95%).