14 ms·
Weave – The Docker Network
- ferrantim 12y agoThese are the same people who built RabbitMQ.
- netcraft 12y agoas someone who is more developer than ops, I feel like the docker stuff is still changing fast and that the way you would use docker today will be very different a year from now; but that containers seem to be the way of the future - if I have no pressing need to change my server architecture does it make sense to wait for things settle or would it be more beneficial to get in and learn now and experience the changes and why they were necessary?
- weavenetwork 12y agoIt is definitely changing fast, but that is not a reason to wait before trying out the technology in order to learn about it. If nothing else, the community will welcome your feedback :-)
- ipedrazas 12y agoprecisely because you're a developer you whould embrace docker with open arms. Because docker removes all the trouble of running applications that you need for your development: databases, application servers, queues... I love the fact that I can focus on my code and not on all those details that stole so much of my time.
- pas 12y agoI'd like to see a post/writeup (or even an essay!) on how to do the magic "backing services" (http://12factor.net/backing-services http://12factor.net/backing-services) with Docker. And that's where I find Deis and other Docker orchestration systems lacking very much. Sure, you can run MySQL in Docker, but it's a far cry from running it on native xfs with aligned partitions and whatever fancy you feel configuring. And since docker containers are very reusable, whereas backing data is by default should be persistent, my impression is that it's too easy to accidentally remove a docker container.
- vidarh 12y agoNothing is stopping you from running MySQL in Docker on native xfs with aligned partitions: bind-mount whatever partition you want into the container by defining a volume in Docker. This will be persistent, and will survive when you destroy the container. I use this to e.g. share a /home directory between a dozen experimental dev container I use to run my various projects - each container ensures I keep track of exact dependencies for each individual project, while I get to have a nice "comfortable" swiss-army-knife container with my dev tools and all all project files. I also run a number of database containers which use volumes where I bind mount host directories to ensure persistence so I can wipe and rebuild the containers themselves without worrying about touching data.
- wastedhours 12y agoAm intrigued (and, again, showing my current early-stage understanding of LXCs), can you link the same data store container to multiple application containers? As in, have both a beta application and production application pulling data from the same core DB? And do you simply define the container as a volume to ensure it stays persistent? That was the feeling I got from the docs, but again, might just be flagging how little I know at the minute...
- TheDong 12y agodocker run --name=my-data -v /host/data:/container/data data-container docker run --volumes-from=my-data app-beta-container docker run --volumes-from=my-data app-prod-container That would share the data store, however the real way you'd do this would be.... docker run --name=my-data -v /host/data:/container/data data-image docker run --volumes-from my-data --name my-database database-image docker run --link=my-database beta-app docker run --link=my-database prod-app Doing --link will allow those two containers to network-communicate and you should only be communicating with your database over the network anyways.
- Someone1234 12y agoDocker hopefully will be a little different in a year as it will offer solid security isolation (which is not the case now). Right now if you run a Docker container and run something within it as root, if that application gets compromised someone can break out of that container and alter the master (due to the way Docker links container-root and system-root, a root user in a container is effectively a root user on the whole system). Docker are working on allowing containers to run entirely in user-mode (thanks to improvements in LXC). This would mean that you can run a process as root within a container, and if that gets compromised there is near zero chance of leveraging that into damaging the master OS (since it will just have normal user privileges). Here's an article about their progress (to usermode): http://s3hh.wordpress.com/2013/07/19/creating-and-using-containers-without-privilege/ http://s3hh.wordpress.com/2013/07/19/creating-and-using-cont... To quote Docker's own documentation[1]: > However, it has been pointed out that if a kernel vulnerability allows arbitrary code execution, it will probably allow to break out of a container — but not out of a virtual machine. In other words, right now, a root process is likely able to escape a Docker container. You can use SELinux, AppArmor, and similar to somewhat mitigate that when it happens but neither are near as powerful as having that usermode isolation on the master. If Docker is able to get usermode containers working, it will be very difficult for a Docker container to either alter other Docker containers or the master system (other than over the network, maybe). [1]http://blog.docker.com/2013/08/containers-docker-how-secure-are-they/ http://blog.docker.com/2013/08/containers-docker-how-secure-...
- SEJeff 12y agoIt has nothing to do with LXC, but the Linux kernels. LXC simply bundles all of the linux kernel namespace features together into one shiney thing to make containers. Conceptually, docker, via libcontainer, does the exact same thing.
- jpgvm 12y agoI hope all of these Docker overlay networks start using the in-kernel overlay network technologies soon. User-space promiscuous capture is obscenely slow. Take a look at GRE and/or VXLAN and the kernels multiple routing table support. (This is precisely why network namespaces are so badass btw). Feel free to ping me if you are working on one of these and want some pointers on how to go about integrating more deeply with the kernel. It's worth mentioning these protocols also have reasonable hardware offload support, unlike custom protocols implemented on UDP/TCP.
- shykes 12y agoI would love to take you up on that. I want to bake vxlan support into Docker upstream (as an optional plugin, like everything else). Edit: Hi Joseph! Just realized it was you :)
- jpgvm 12y agoHaha for sure Solomon, I have some other ideas now too that I am using Docker in production. I will hit you up to have a chat. :)
- kijiki 12y agoIf you're going down the path of VXLAN support in Docker, I'd love to talk. The company I founded built a Linux distribution for commodity hardware switches that can do VXLAN encap/decap in hardware at 2+ Tbit/sec. The same configuration that works in a Linux container host or a hypervisor works on the switches. nolan@cumulusnetworks.com
- darren0 12y agoYou have to pick either kernel or user space, not both. Either implement it purely in the kernel or purely in user space. In reality pure user space is faster, just look at Snabb switch https://github.com/SnabbCo/snabbswitch/wiki https://github.com/SnabbCo/snabbswitch/wiki
- wmf 12y agoSnabb and DPDK aren't magic though. Because they poll you have to dedicate a whole core to the vSwitch. Containers are a different case than VMs because the packets start in the kernel TCP/IP stack; to get into a userspace vSwitch they'd have to exit the kernel.
- thu 12y agoThis seems very nice. What would be the pros and cons of using Weave instead of Tinc ? I have used Tinc for a while[0] and, the end result looks very similar (i.e. there is not a nice command-line tool dedicated to use Tinc with Docker, but the high level description match). [0]: https://gist.github.com/noteed/11031504 https://gist.github.com/noteed/11031504
- weavenetwork 12y agoJust having a look at Tinc.. one difference is completely non-technical - Tinc is GPL but Weave is Apache licensed so aligns better with the whole Docker ecosystem. More comments to come.
- grkvlt 12y agoThis is really interesting. I've been looking for a way to build in support for networking between Docker hosts in my clocker.io software, to simplify deploying applications into a cloud hosted Docker environment. I'd been young with adding Open vSwitch, but am going to try weave as the network layer in the next release. Will there be any problems running in a cloud where I have limited control over the configuration of the host network interfaces and the traffic they can carry, such as AWS only allowing TCP and UDP between VMs?
- weavenetwork 12y agoTCP and UDP is all you need. (And the UDP can actually be quite broken, just not completely) We've created weave networks spanning hosts on EC2, GCE and local data centres.
- zobzu 12y agoWhat this really means security wise: http://i.imgur.com/Cko02do.png http://i.imgur.com/Cko02do.png
- cschneid 12y agoEach `W` node in that graphic is a different physical host. So 3 kernels.
- saryant 12y agoQuestion for weavenetwork: are containers addressable by hostname from other containers? Is there a good way to do that? I didn't see anything about it in the readme. I suppose service discovery is out-of-scope for this project but having some sort of weave-wide hostsfile would certainly simplify it. Am I misunderstanding the project?
- weavenetwork 12y agoWeave itself does not provide addressability beyond IP. That is the situation now, but this area is very much high on the agenda for us - service discovery is definitely in scope for weave. Meanwhile, two points of note: 1) In weave the IP addresses can be much "stickier" than in other network setups, i.e. a moving a container from one host to another can retain the containers IP. That means it is quite amenable to relatively static name resolution configurations, e.g. via /etc/hosts files. 2) Since weave creates a fully-fledged L2 Ethernet network between app containers, name resolution technologies like mDNS that rely on multicast should work just fine. So, in summary, while weave currently does not have any built-in service discovery, existing solutions and technologies for that should be relatively easy to deploy inside weave application networks, until weave itself grows these capabilities.
- SEJeff 12y agoPerhaps you'd consider looking for the best way to "weave" weave into consul? http://consul.io http://consul.io
- weavenetwork 12y agoWe are certainly aware of consul, and have indeed been thinking of weaving weave into it. Would love to see an experiment along those lines, if there are any volunteers.
- greenimpala 12y agoErr can anyone spot the tests in the repo? I cannot.
- t0mas88 12y agoThis looks like a great idea. For me this was a missing piece two months ago when playing with Docker. However I have strong doubts about the network performance, not only the overhead of the UDP encapsulation (that should be quite small), but mostly the capturing of packets with pcap and then handling them in user-mode. Looks like a lot of context-switches, copying and parsing with non-optimal code paths. Are there any benchmarks available? My feeling is that this will consume large amounts of CPU for moderate network loads and thus be unusable with most NoSQL kind of systems that benefit from clustering across hosts?
- weavenetwork 12y agore benchmarks...publishing some is on the TODO list. See https://github.com/zettio/weave/issues/37 https://github.com/zettio/weave/issues/37. Informally, weave is pretty fast but it's not saturating Gbit Ethernet. As you say, capturing with pcap and handling packets in user space carries an appreciable overhead. We've got some issues filed to look at pcap alternatives and also generally aim to improve performance. re suitability for NoSQL clustering... depends on where the bottlenecks are; if you want to cluster for HA rather than scale, i.e. there aren't any real bottlenecks, then weave will work well. Same if you want to cluster because of CPU or memory bottlenecks. If, otoh, networking is the bottleneck then adding weave into the mix isn't going to improve matters.
- t0mas88 12y agoOk, good to know. I think the challenge in taking another route than pcap is that you would need to do complex tricks with the existing network stack. Because if I understand the way Weave works you would really only need to do processing at the beginning of a connection and for some ARP requests etc while you don't need to do anything to existing TCP streams apart from encapsulating and forwarding?
- weavenetwork 12y ago> complex tricks with the existing network stack To retain the essence of how weave operates, this would likely not just be complex but impossible, short of kernel hackery. > you would really only need to do processing at the beginning of a connection and for some ARP request Weave needs to look at every Ethernet packet. Well, the headers at least. It's a virtual Ethernet switch. It doesn't even really know about IP, let alone TCP streams. See https://github.com/zettio/weave#how-does-it-work https://github.com/zettio/weave#how-does-it-work
- deleted 12y ago[deleted]
- baq 12y agocan you compare this with openvpn, or any other vpn if we're at it?
- brazzledazzle 12y agoHas anyone compared this to rudder (https://coreos.com/blog/introducing-rudder/ https://coreos.com/blog/introducing-rudder/)?
- bboreham 12y agoRudder has a central configuration (via etcd); Weave communicates network changes across all nodes by itself - as long as each new node knows how to contact at least one other node, it will learn of all other nodes and connect to as many as it can. Also Rudder is Layer 3 and Weave is Layer 2 and Weave can encrypt traffic.
- weavenetwork 12y agoThe most significant conceptual difference is that in rudder sub-nets are tied to hosts. So containers on different hosts will always be on different sub-nets. By contrast, in weave containers belonging to the same application reside in the same sub-net, regardless of what host they are running on. In other words, weave makes the network topology fit the application topology, not the other way round.
- GrantNelson 12y agoYikes, this looks scary. Just because you can do something doesn't mean you should. Networks are finicky, perf is king.