3 ms·
This is a wonderful article. Thanks for sharing. As always, Cloudflare blog posts do not disappoint. It’s very interesting that they are essentially treating I
by uvdn7 4y ago
This is a wonderful article. Thanks for sharing. As always, Cloudflare blog posts do not disappoint.
It’s very interesting that they are essentially treating IP addresses as “data”. Once looking at the problem from a distributed system lens, the solution here can be mapped to distributed systems almost perfectly.
- Replicating a piece of data on every host in the fleet is expensive, but fast and reliable. The compromise is usually to keep one replica in a region; same as how they share a single /32 IP address in a region.
- “sending datagram to IP X” is no different than “fetching data X from a distributed system”. This is essentially the underlying philosophy of the soft-unicast. Just like data lives in a distributed system/cloud, you no longer know where is an IP address located.
It’s ingenious.
They said they don’t like stateful NAT, which is understandable. But the load balancer has to be stateful still to perform the routing correctly. It would be an interesting follow up blog post talking about how they coordinate port/data movements (moving a port from server A to server B), as it’s state management (not very different from moving data in a distributed system again).
- remram 4y agoI have a lot of trouble mapping your comment to the content of the article. It is about the egress addresses, the ones CloudFlare use as source when fetching from origin servers. Those addresses need to be separated by the region of the end-user ("eyeball"/browser) and the CloudFlare service they are using (CDN or WARP). The cost they are working around is the cost of IPv4 addresses, versus the combinatorial explosion in their allocation scheme (they need number of services * number of regions * whatever dimension they add next, because IP addresses are nothing like data). I am not sure where you see data replication in this scheme?
- uvdn7 4y agoIt's not meant to be a perfect analogy. The replication analogy is mostly talking about the tradeoff between performance and cost. So it's less about "replicating" the ip addresses (which is not happening). On that front, maybe distribution would be a better term. Instead of storing a single piece of data on a single host (unicast), they are distributing it to a set of hosts. Overall, it seems like they are treating ip addresses as data essentially, which becomes most obvious when they talk about soft-unicast. Anyway, I just found it interesting to look at this through this lens.
- majke 4y ago"Overall, it seems like they are treating ip addresses as data essentially" Spot on! In past: * /24 per datacenter (BGP), /32 per server (local network) (all 64K ports) New: * /24 per continent (group of colos), /32 per colo, port-slice per server This is totally hierarchical. All we did is build a tech to change the "assignment granularity". Now with this tech we can do... anything we want. We're not tied to BGP, or IP's belonging to servers, or adjacent IP's needing to be nearby. The cost is the memory cost of global topology. We don't want a global shared-state NAT (each 2 or 4-tuple being replicated globally on all servers). We don't want zero-state (a machine knowing nothing about routing, just BGP does the job). We want to select a reasonable mix. Right now it's /32 per datacenter.... but we can change it if we want and be more, or less specific than that.
- ludikalell2 4y agoOnly downside seems more stress on quicksilver ;)
- 0xbkt 4y agoIs this feature fully rolled out yet? I still see unicast IPs connecting from the entry colo and traceroute confirms that.
- nvarsj 4y agoBy stateful NAT they mean connection tracking. In the described solution, the LB/router doesn’t track connections - it simply looks up the server via local mapping from port range to server, and forwards the packet. Incidentally, this is exactly how GCP Cloud NAT works.