7 ms·
Introducing the Fan – simpler container networking
- frequent 11y agoI would want to object and say "There is IPv6 in the cloud!". We have developed re6stnet in 2012. You can use it to create an ipv6 network on top of an existing ipv4 network. It's open source and we are using it ourselves internally and in client implementations since then. I wrote a quick blogpost on it: http://www.nexedi.com/blog/blog-re6stnet.ipv6.since.2012 http://www.nexedi.com/blog/blog-re6stnet.ipv6.since.2012 The repo is here in case anyone is interested: http://git.erp5.org/gitweb/re6stnet.git/tree/refs/heads/master?js=1 http://git.erp5.org/gitweb/re6stnet.git/tree/refs/heads/mast...
- dustinkirkland 11y agoThere's IPv6 in some clouds. https://forums.aws.amazon.com/thread.jspa?messageID=536049 https://forums.aws.amazon.com/thread.jspa?messageID=536049
- teraflop 11y agoFrom the link in the parent comment: > Well, most importantly we have stable IPv6 everywhere - including on IPv4 legacy networks.
- rsync 11y agorsync.net cloud storage has been ipv6 accessible since 2006. I'm surprised that is still interesting/noteworthy. HN Readers discount, since we're on the subject. Just email.
- paulasmuth 11y agoI don't seem to get it. How is this different from just using a non-routed IP per container?
- carapace 11y agoYeah, I'm bemused. It seems like a solution to an issue that should have been designed away in the first place. But I'm speaking out of ignorance here so...
- falcolas 11y agoI don't quite get how your solution is equivalent: How many of IPs do you have available per host in, say an EC2 instance? How do you talk to a container which isn't on the same host? It feels similar to vxlan, only without the centralized repository of IP -> host mappings.
- epistasis 11y agoBenefits: * cross-host container networking * no overlay IP database to sync and maintain * only a single IP used per host They still use encapsulation, like other network overlay technologies, it's just that by using a specific addressing scheme they can eliminate a lot of cross-host communication and all the database lookups.
- maccam94 11y agoTo build on epistasis's comment a bit, this creates a private network for all containers on all hosts that reside in the same /16 network. So if you have a VPC of up to 65k machines, each machine can run up to ~250 containers that can all talk directly to each other by just relying on basic network routing. This is better than your typical private NAT bridge networking because containers on different hosts can talk to each other without having to set up port forwarding or discovering what port that particular application server is running on.
- kbaker 11y agoSorry, I just gotta rant a bit... this is a really bad hack, that I wouldn't trust on a production system. Instead of doubling down and working on better IPv6 support with providers and in software configuration, and defining best practices for working with IPv6, they just kinda gloss over with a 'not supported yet' and develop a whole system that will very likely break things in random ways. > More importantly, we can route to these addresses much more simply, with a single route to the “fan” network on each host, instead of the maze of twisty network tunnels you might have seen with other overlays. Maybe I haven't seen the other overlays (they mention flannel), but how does this not become a series of twisty network tunnels? Except now you have to manually add addresses (static IPv4 addresses!) of the hosts in the route table? I see this as a huge step backwards... now you have to maintain address space routes amongst a bunch of container hosts? Also, they mention having up to 1000s of containers on laptops, but then their solution scales only to 250 before you need to setup another route + multi-homed IP? Or wipe out entire /8s? > If you decide you don’t need to communicate with one of these network blocks, you can use it instead of the 10.0.0.0/8 block used in this document. For instance, you might be willing to give up access to Ford Motor Company (19.0.0.0/8) or Halliburton (34.0.0.0/8). The Future Use range (240.0.0.0/8 through 255.0.0.0/8) is a particularly good set of IP addresses you might use, because most routers won't route it; however, some OSes, such as Windows, won't use it. (from https://wiki.ubuntu.com/FanNetworking https://wiki.ubuntu.com/FanNetworking) Why are they reusing IP address space marked 'not to be used?' Surely there will be some router, firewall, or switch that will drop those packets arbitrarily, resulting in very-hard-to-debug errors. -- This problem is already solved with IPv6. Please, if you have this problem, look into using IPv6. This article has plenty of ways to solve this problem using IPv6: https://docs.docker.com/articles/networking/ https://docs.docker.com/articles/networking/ If your provider doesn't support IPv6, please try to use a tunnel provider to get your very own IPv6 address space. like https://tunnelbroker.net/ https://tunnelbroker.net/ Spend the time to learn IPv6, you won't regret it 5-10 years down the road...
- epistasis 11y agoAnd what about those environments where only IPV4 is available, as specifically addressed in this article? There are lots of overlay networks, are those all hacks too? >Except now you have to manually add addresses (static IPv4 addresses!) of the hosts in the route table? This does not appear to be true at all, based on the configuration that's posted at the bottom of the article.
- ademarre 11y agoI'd like to see a better explanation of how this compares to the various Flannel backends (https://github.com/coreos/flannel#backends https://github.com/coreos/flannel#backends), and also how this would be plugged into a Kubernetes cluster.
- dustinkirkland 11y agoIP-IP is the first encapsulation supported, but the Fan is engineered in a way that any encapsulation scheme can be added easily. We'll be adding support for VXLAN, GRE, STP tunnels as well. My colleagues will have blog posts about how to enable Kubernetes clusters with Fan Newtorking.
- GauntletWizard 11y agoWhy do people keep giving whole IP addresses to every little container? It's a terrible management paradigm compared to service-discovery and using hostports for every address.
- vidarh 11y agoBecause a lot of software stuff is written without intrinsic support for service discovery, which means all kinds of hacks (e.g. proxying; dynamically rewriting config files and reloading) to work around it if the ports may change. With an IP per container, things like low TTL DNS tied to a service registry (e.g. Consul, or SkyDNS with something to update Etcd) is viable and often a much easier alternative. I agree with you from a purity point of view that proper service-discovery is better, but in terms of practicality, IP per container is often a lot simpler to implement.
- toast0 11y agoThis is all for your internal network though, which you should be able to control, right?
- vidarh 11y agoHaving control of your internal network does not fix missing service discovery support in all the applications you're running that likes to assume that port numbers are static and unchanging.
- rubiquity 11y agoBecause service discovery itself is ridiculous. I have to deploy three or more Consul/etcd nodes to run service discovery for my handful of EC2 instances and containers?
- otterley 11y agoOf course not. You can maintain configuration files if you like, but anything beyond a relatively small network is going to be a pain to maintain. It's also not going to react very quickly to node failures or topology changes.
- regularfry 11y agoOr you could go somewhere with IPv6. The number of places with an IPv4-only restriction is only going to drop.
- falcolas 11y agoThat's been the promise for... how many years now? IIRC, it was before EC2 even existed, and we're obviously not there yet. Also, it's worth noting that IPV6 is nowhere near as battle hardened as IPV4; there's too many optimization and security gaps to depend on it in production. I've watched a few network gurus burn themselves out attempting to harden a corporate network against IPV6 attacks while keeping it usable.
- regularfry 11y ago> That's been the promise for... how many years now? IIRC, it was before EC2 even existed, and we're obviously not there yet. For some values of "we". If you're stuck on EC2, yeah, you've got a problem.
- coldtea 11y agoBut if you don't use EC2 but some bizarro provider noone really uses then you can have IPv6...
- betaby 11y agoI've heard OVH is second biggest in the world and they do have IPv6.
- regularfry 11y agoNobody ever got fired for buying IBM, right?
- coldtea 11y agoBecause AWS is some legacy dinosaur and not the world leader in such infrastructure, right?
- tobbyb 11y agoThis looks unbelievably simple, on the lines of why hasn't it been done before. So you ping 10.3.4.16 and your host automatically 'knows' to just send it to 17.16.4.16 where lying in wait, the receiving host simply forwards it to 10.3.4.16. I like it. This is a vexing problem for containers and even VM networking. If they are in a NAT you need to create a mesh of tunnels across hosts, or you create a flat network so they are all on the same subnet. But you can't do this for containers on the cloud with a single IP and limited control of the networking layer. Current solutions include L2 overlays, L3 overlays, a big mishmash of GRE and other type of tunnels, or VXLAN multicast unavailable in most cloud networks, or proprietary unicast implementations. It's a big hassle. Ubuntu have taken a simple approach, no per node database to maintain state and uses commonly used networking tools. And more importantly it seems fast. And it's here and now. That 6gbps suggests this does not compromise performance like a lot of other solutions tend to do. It won't solve all multi-host container networking use cases but will address many.
- epistasis 11y agoYes, the only thing that's lost is live migration of IPs between hosts. Which may or may not be a big thing for containers, depending on things are clustered.
- stephengillie 11y agoWhy does NAT require a mesh of tunnels? Are you trying to separate the containers onto securely-separate networks? Why isn't DHCP used? What does it not do that this service does?
- jpgvm 11y agoYou can use any method to program the VXLAN forwarding table, you don't need to use multicast. This can even be done on the command line using iproute2 utilities: https://www.kernel.org/doc/Documentation/networking/vxlan.txt https://www.kernel.org/doc/Documentation/networking/vxlan.tx... Though you should probably use netlink to do it programatically. Personally I like to combine netlink + Zookeeper or similar to trigger edge updates via watches.
- 11y ago
- rcarmo 11y agoThis is very neat indeed, and I'd love to try it out, but the launchpad links are broken. Anyone know where I can get the package for Ubuntu armhf? Or the source?
- geku 11y agoSeems to be a smart solution but it only works when you have control over the "real" /16 network if I understand it correct? E.g. having multiple nodes on multiple cloud providers with completely different IP addresses not in the same /16 network will not work, correct?
- stephengillie 11y agoI've read the article twice. Did they just reinvent putting DHCP behind a NAT? What does that combination of systems not do that Fan does? *Remap 50 addresses from one range to another. *Dynamically assign those addresses to servers. *Special Something that Fan does. What's the benefit of using a full class A subnet when you are only using 250 addresses?
- dustinkirkland 11y agoSure, DHCP/NAT is used by each container, to get out. But how does it route to another container on some other host elsewhere in your cloud? That's what the Fan addresses.
- stephengillie 11y ago> Sure, DHCP/NAT is used by each container, to get out. So...Each physical host has its software networking pass through a NAT before hitting the physical adapter? And it's already using DHCP to assign addresses to its containers? > But how does it route to another container on some other host elsewhere in your cloud? That's what the Fan addresses. You use DNS to create a lookup table matching container to IP? Since this isn't being done, it must not work. What does Fan do instead? Is there some other complicating factor to which I'm ignorant? Are we talking about having multiple Kubernetes clusters inside containers inside VMs inside a physical host? Also, HOW does Fan address this problem? What does it use instead of one kind of database lookup or another, like a distributed database system such as DNS? Or does Fan use fancy subnet math?
- api 11y agoProbably doesn't matter much here, but 240.0.0.0/4 is hard coded to be unusable on Windows systems. It's in the Windows IP stack somewhere. Packets to/from that network will simply be dropped.
- rspeer 11y agoThat will certainly not be the only obstacle to running your containers on Windows.
- dustinkirkland 11y agoYou can use absolutely any /8 that you want. I used 250.0.0.0/8 in my examples, but the FanNetworking wiki page uses 10.0.0.0/8. You can use any one you want. Have at it ;-)
- api 11y agoOn Windows?
- rsync 11y ago"Also, IPv6 is nowehre to be seen on the clouds, so addresses are more scarce than they need to be in the first place." We've[1] had ipv6 addressable cloud storage since 2006. Currently our US (Denver), Hong Kong (tsuen kwan o) and Zurich locations have working ipv6 addresses. [1] You know who we are.
- nacs 11y agoWhat percent of the traffic is actually flowing through the ipv6 ones compared to v4 -- <10% I'd guess (just curious)?