8 ms·
How Netflix Accurately Attributes eBPF Flow Logs
- butlike 1y agoAll that logging and they cant figure out why people are going to other streaming services
- neogodless 1y agoI mean, I haven't been a subscriber for a couple years but are they losing a lot of subscribers? https://backlinko.com/netflix-users https://backlinko.com/netflix-users I see some slowing growth particularly across 2021/2022 but as of this report (April 2024) they were still growing through 2023. 260M subscribers. They aren't exactly hurting.
- EwanToo 1y agoIt's now over 300 million subscribers, as you say they're doing OK. https://ir.netflix.net/investor-news-and-events/financial-releases/press-release-details/2025/Netflix-to-Announce-First-Quarter-2025-Financial-Results/default.aspx https://ir.netflix.net/investor-news-and-events/financial-re...
- ASalazarMX 1y agoDespite their awful UX, I'm always impressed with how reliable their service is, technically speaking. Video is always good and responsive even on less-than-stellar connections, you can leave a show paused for hours, and resume it almost instantly. Their fast.com speed test is always much faster than your regular internet access, I guess thanks to their Open Connect Appliances. It must be great to work for them in infrastructure and backends.
- yuters 1y agoI have an old fire tv and never tried to stop automatic updates on it, it has become so slow and unresponsive that I'm barely able to switch inputs to use something else. Netflix is the only app that still works on that tv.
- ciupicri 1y agoI also have an old TV and guess what? Netflix stopped working last year. The application is not supported anymore. Beats me why.
- acdha 1y agoOften it’s CA certificates expiring. My old Toshiba had app rot set in like that, where after about 5 years none of the built in apps worked any more and the errors appeared to be TLS related. I suspect that was due to pinned certs to prevent MITM pirating.
- faitswulff 1y agoI can only remember one major outage from them in the past ~10 years (in the 2020s, not the 2012 outage), and if I recall correctly, it was fixed in short order...and they never released a postmortem
- silisili 1y agoGenerally I'd agree, but you must not have seen their attempts at live events :(.
- toomuchtodo 1y agoThey also performed work to ensure it performed well for Starlink customers. A Global Perspective on the Past, Present, and Future of Video Streaming over Starlink - https://dl.acm.org/doi/10.1145/3700412 https://dl.acm.org/doi/10.1145/3700412 | https://doi.org/10.1145/3700412 https://doi.org/10.1145/3700412
- ndriscoll 1y agoNot that this detracts from the wider point, but I'd expect unpause to just work unless you go out of your way to make it not work. Even if you drop the connection at some point, afaik they use ~15 Mb/s as their "premium" bitrate, so e.g. a 30 s buffer takes less than 64 MB. That gives plenty of time to re-establish streaming after an unpause. It's not like the computer forgets what it was doing if you leave it alone.
- ASalazarMX 1y agoCounterpoint: Plex and Jellyfin free resources if you leave your video paused too long, and it will take a noticeably amount of time to resume streaming, much more if it needs transcoding. They're not going out of their way to annoy us, they try to be efficient with the finite resources a home server has. Netflix is going out of their way to make it smooth no matter what you do, even if they have to pool a bit of their own resources for it.
- ndriscoll 1y agoBuffering would be on the client though. Assuming it has a couple dozen MB of memory, it should be able to buffer like 30 seconds. I realize their resume has more to do, but e.g. my jellyfin server can initiate playback or seek within maybe 100-250 ms (it's just barely a noticable pause after a random seek). So a 30 s buffer should be more than sufficient for unpausing without any stutter.
- seneca 1y agoNo one I know uses Netflix anymore, and I haven't for a while, but from what I've seen their subscriber numbers are actually doing quite well.
- temp0826 1y agoI'm not sure I know anyone that _doesn't_ have it
- steve_adams_86 1y agoLike seneca, I also don’t know anyone who uses it. That’s interesting. I haven’t used it for around 2 years. I wonder who uses it when I see their reports, because I don’t know them. I probably have a weird group of friends. I’m sure if I asked, some of my coworkers use it. My kids would love if we used it. Perhaps it’s big among younger people.
- temp0826 1y agoMaybe that's part of it?...my sibling has a toddler and they use it for a lot of children's shows. Conversely, my mother watches it often too (not for kids shows!)
- itishappy 1y agoDo you use something different? Anecdotally, I find Disney+ to be a major divider. Friends of mine have kids and Disney+ and/or they're still using a decade-old Netflix subscription.
- steve_adams_86 1y agoWe use Apple TV a small amount (for Silo and Severance most recently), and Disney+ slightly more for the kids. We occasionally use YouTube. We got a promotion for Disney+ and probably won't renew it. By the time it's over, it seems like there won't be much worth watching left on there. The kids already have a hard time finding anything they're interested in watching.
- blinded 1y agolol wat? They make 10 billion+ a <b>quarter</b>.
- thewisenerd 1y agoso they didn't want to pay for AWS CloudWatch [1]; decided to roll their in-house network flow log collection; and had to re-implement attribution? i wonder how many hundreds of thousands of dollars network flow logs cost them; obviously at some point it is going to be cheaper to re-implement monitoring in-house. [1]: https://youtu.be/8C9xNVYbCVk?feature=shared&t=1685 https://youtu.be/8C9xNVYbCVk?feature=shared&t=1685
- Hikikomori 1y agoBecause vanilla flowlogs that you get from VPC/TGW are nearly useless outside the most basic use cases. All you get is how many bytes and which tcp flags were seen per connection per 10 minutes. Then you need to attribute ip addresses to actual resources yourself separately, which isn't simple when you have containers or k8s service networking. Doing it with eBPF on end hosts you can get the same data, but you can attribute it directly as you know which container it originates from, snoop dns, then you can get extremely useful metrics like per tcp connection ack delay and retransmissions, etc. AWS recently released Cloudwatch Network Monitoring that also uses an agent with eBPF, but its almost like a children's toy compared to something like Datadog NPM. I was working on a solution similar to Netflix's when NPM was released, was no point after that.
- nptr 1y agoThis is spot on. The AWS logs can also be orders of magnitude more expensive.
- DadBase 1y agoI recall a time when we managed network flows by manually parsing /proc/net/tcp and correlating PIDs with netstat outputs. eBPF? Sounds like a fancy way to avoid good old-fashioned elbow grease.
- nikolay_sivko 1y agoAt Coroot, we solve the same problem, but in a slightly different way. The traffic source is always a container (Kubernetes pod, systemd slice, etc.). The destination is initially identified as an IP:PORT pair, which, in the case of Kubernetes services, is often not the final destination. To address this, our agent also determines the actual destination by accessing the conntrack table at the eBPF level. Then, at the UI level, we match the actual destination with metadata about TCP listening sockets, effectively converting raw connections into container-to-container communications. The agent repo: https://github.com/coroot/coroot-node-agent https://github.com/coroot/coroot-node-agent
- zX41ZdbW 1y agoIf you are interested in network monitoring in Kubernetes, it's worth looking at Kubenetmon: https://github.com/ClickHouse/kubenetmon https://github.com/ClickHouse/kubenetmon - an open-source eBPF-based implementation from ClickHouse.
- thewisenerd 1y agoi mean.. from your blog post linked in the repo; this isn't eBPF based? https://clickhouse.com/blog/kubenetmon-open-sourced https://clickhouse.com/blog/kubenetmon-open-sourced the data collection method says: "conntrack with nf_conntrack_acct"
- nimbius 1y agoi refuse to believe a company that wasted $320 million dollars on "the electric state" could ever manage to do anything correctly. stripe the parking lot? stock the breakroom? clean the toilets? simply not possible.
- ZeWaka 1y agoBeautiful art and book, but what a unfaithful travesty of a production that absolutely trodded on the original work.
- autoexec 1y agoBadly managed companies often hire some very good talent which can allow them to sometimes do very impressive things in spite of themselves.
- jandrese 1y agoThe scriptwriters aren't managing the network flows across the backend of their infrastructure. Netflix's trouble with scripts doesn't affect their ability to move bits around.
- slt2021 1y agoQuestion to the Netflix folks: I saw a lot of in-house developed tools being quoted, do you guys have service mesh like linkerd ? Have you guys evaluated vendors like Kentik? I would love to get more insight into what do you guys actually do with flow logs? for example if I store 1 TB of flow logs, what value can I actually derive from them that justify the cost of collection, processing, and storage.
- retiredpapaya 1y agoI think Netflix does use an Envoy-based Service Mesh [1], and they roll their own control plane. https://netflixtechblog.com/zero-configuration-service-mesh-with-on-demand-cluster-discovery-ac6483b52a51 https://netflixtechblog.com/zero-configuration-service-mesh-...
- slt2021 1y agoIf the goal of gathering and attributing VPC flows is to have a workload granularity flow logs, then imho gathering mesh level logs is more direct and atraight forward approach, because mesh(and workload orchestrator) are uniquely qualified to know when workload A is running on a host X and is trying to connect to workload B. Looking at Envoy access logs for example is more straightforward and simple aplroach, than running distributed ebpf and memory intensive large spark streaming job
- nptr 1y agoThe blog post mentioned that "The eBPF flow logs provide a comprehensive view of service topology and network health across Netflix’s extensive microservices fleet, regardless of the programming language, RPC mechanism, or application-layer protocol used by individual workloads." Service mesh may have restrictions on the network protocols and may not cover all network traffic (like connections to Kafka and databases).
- madduci 1y agoExactly my thought. Maybe it's the "not invented here" syndrome? We use Istio as Service Mesh and get the same result, using the same architecture as shown in the blog post (especially the part where each workload has a sidecar container running Flow).
- r3tr0 1y agowe are working on a similar product that is eBPF powered and can extract flow logs: https://yeet.cx https://yeet.cx
- __turbobrew__ 1y agoMaybe Im missing something but can’t you run workloads in separate network namespaces and then attach a bpf probe to the veth interface in the namespace? At that point you know all flows on that veth are from a specific workload as long as you keep track of what is running in which network namespaces? I wonder if it is possible with ipv6 to never (or you roll through the addresses so reuse is temporally distant) re use addresses which removes the problems with staleness and false attribution.
- VaiTheJice 1y agoI think thats pretty reasonable tbf and probably at a more 'simpler' scale and i use simple loosely because Netflix’s container runtime is Titus, which is more bare metal oriented than, say, Kubernetes. It doesn’t always isolate workloads as cleanly in separate netns per container, especially for network optimisation purposes like IPv6-to-IPv4 sharing. "I wonder if it is possible with ipv6 to never... re use addresses which removes the problems with staleness and false attribution." Most VPCs (also AWS) don’t currently support "true" IPv6 scaleout behavior. Buttt!! if IPs were truly immutable and unique per workload, attribution becomes trivial. It’s just not yet realistic... maybe something to explore with the lads?
- __turbobrew__ 1y agoMakes sense, I have worked in and around CNI stuff for k8s and generally netns+veth is how most of them work. That being said we run k8s on bare metal, there isn’t any reason why running things on bare metal excludes netns usage. > Most VPCs (also AWS) don’t currently support "true" IPv6 scaleout behavior. Thats a shame. > if IPs were truly immutable and unique per workload, attribution becomes trivial I would like to see that. IPAM for multi-tenant workloads always felt like a kludge. You need the network to understand how to route to a workloads, but the network when running on ipv4 has many more workloads than addresses. If you assign immutable addresses per workload (or say it takes you a month to chew through your ipv6 address space) it makes it so the network natively knows how to route to workloads without the need to kludge with IP reassignments. I have had to deal with IP address pools being exhausted due to high pod churn in EC2 a number of times and it is always a pain.
- mmckeen 1y agohttps://retina.sh/ https://retina.sh/ is a similar open source tool for Kubernetes. It's early and has some bugs but seems promising.
- roboben 1y agoTried it, had some issue, opened a bug report, no response. I think it is dead.
- meltyness 1y agoI wonder how much of Netflix infra is on AWS. Feels like building a castle on someone else's kingdom at that scale; in light of the Prime Video investment, and I guess twitch too.
- r3trohack3r 1y agoNetflix serves nearly all of its video from a server down the street from you via its OpenConnect infrastructure. AWS only hosts its microservice graph that does stuff like determining which videos and qualities you should be offered. That being said, its core product has been nearly comoditized. When Netflix entered the market, delivering long form high quality video over the public internet was nascent. Now everyone and their grandma can spin up a video streaming service from a vendor.
- meltyness 1y agoI was aware of this, I think a talk about a performance regression on the BSD variant these appliances run was up here recently. I mean I guess a similar argument holds for colocating with telecoms that would have recently been cable or IPTV providers. It'd be pretty tough to design around though if any of their caching infrastructure reveals viewership, or engagement data to intermediaries.
- scyzoryk_xyz 1y agoAny chance you would be able to point me to a good source or article describing/explaining the first half of your comment? I.e. someone getting into the nuts and bolts? Netflix is a platform - their strategic advantage is in their content sourcing and development pipeline which is fed the unique insights on audience preferences. This is distributed with recommendation algorithms and UX. It could be argued, like someone also already pointed out, that this infra aspect is a commodity at this point.
- miyuru 1y agohttps://openconnect.netflix.com/en/ https://openconnect.netflix.com/en/ there are attempts to serve high bandwidth throughput, I think the last update was below. https://news.ycombinator.com/item?id=40329303 https://news.ycombinator.com/item?id=40329303
- philsnow 1y agoIs it necessary to rely on ip address attribution? If FlowExporter uses ebpf and tcp tracepoints, could each workloads be placed in its own cgroup and could FlowExporter directly introspect which cgroup (and thus, workload) a given tcp socket event should be attributed to?
- nptr 1y agoThat may help identify the local IPs but not the remote IPs.