6 ms·
Serf: A decentralized solution for service discovery and orchestration
- AsymetricCom 13y agoThis isn't much of a solution when it involves putting a binary on every host. Clearly, the best solution is a service framework at the platform level, not a separate, unmodifiable blob thrown on your hardware.
- WestCoastJustin 13y agoLooks like this might be loosely related to Mitchell Hashimoto [1], who makes Vagrant. [1] https://twitter.com/mitchellh https://twitter.com/mitchellh
- moondowner 13y agoBoth are HashiCorp projects.
- deleted 13y ago[deleted]
- kyyd 13y agomitchellh 237 commits / 16,925 ++ / 9,149 -- from github, so yeah
- mitchellh 13y agoIndeed, I'm one of the creators of the project. See here: http://www.serfdom.io/community.html http://www.serfdom.io/community.html :)
- deleted 13y ago[deleted]
- deleted 13y ago[deleted]
- mitchellh 13y agoI'm jumping on a plane right now (a couple hours) but I'd be happy to answer any questions related to Serf once I land. Just leave them here and I'll give it my best shot! We've dreamt of something like Serf for quite awhile and I'm glad it is now a reality. Some recommended URLs if you're curious what the point is: "What is Serf?" http://www.serfdom.io/intro/index.html http://www.serfdom.io/intro/index.html "Use Cases" http://www.serfdom.io/intro/use-cases.html http://www.serfdom.io/intro/use-cases.html Comparison to Other Software: http://www.serfdom.io/intro/vs-other-sw.html http://www.serfdom.io/intro/vs-other-sw.html For the CS nerds, the internals/protocols/papers behind Serf: http://www.serfdom.io/docs/internals/gossip.html http://www.serfdom.io/docs/internals/gossip.html Also, I apologize the site isn't very mobile friendly right now. Unfortunately I can write a lamport clock implementation, but CSS is just crazytown.
- nwh 13y agoYou need to do some repair on the website for iOS; the front page is white text on a white background.
- mitchellh 13y agoFixed. How does CSS work?
- stormbrew 13y agoThe one thing that feels lacking to me, and maybe there's just something I don't get, is the ability to tag nodes with metadata. From the docs it seems like what you're expected to do is fire an event for a node to, eg., declare itself a webserver, but this seems prone to failure in the long run. If I bring a new load balancer online how does it find out what's a webserver already?
- armon 13y agoThis might be a little unclear, but if you check the documentation for agent configuration (http://www.serfdom.io/docs/agent/options.html http://www.serfdom.io/docs/agent/options.html), there is an option to provide a role. The role is the metadata support currently
- lambda 13y agoThe title needs to be improved. "A decentralized, highly available, fault tolerant solution..." for what? The title should include "for service discovery and orchestration".
- taterbase 13y agoAgreed, these tag lines come off as buzzword soup instead of informative. I would love a small quick scenario describing what serf can help prevent/enable.
- plainOldText 13y agoYou might find this useful: http://www.serfdom.io/intro/use-cases.html http://www.serfdom.io/intro/use-cases.html
- linker3000 13y agoYes, that page helped me, but the front page should be the hook and it left me none the wiser.
- mjohan 13y agoHow does security work in a system like this. If this is used in a shared hosting system, can a user inject false messages into Serf with for example PHP?
- mitchellh 13y agoYes, they can. In the general case its not an issue because usually your nodes are inaccessible by the public, but if you're using a shared hosting environment, this is entirely possible. We're addressing this in the next release by signing/encrypting gossiped messages. See the roadmap: http://www.serfdom.io/docs/roadmap.html http://www.serfdom.io/docs/roadmap.html
- philips 13y agoIt seems serf and etcd are both trying to solve service discovery but are attacking it using different approaches. Which is pretty cool! Serf looks to be eventually consistent and event driven. So you can figure out who is up and send events to members. This gives you a lot of utility for the use cases of propagating information to DNS, load balancers, etc. But, you couldn't use serf for something like master election or locks and would need etcd or Zookeeper for that. Serf and etcd aren't mutually exclusive systems in any sense just solving different problems in different ways. They have a nice write-up on the page here: http://www.serfdom.io/intro/vs-zookeeper.html http://www.serfdom.io/intro/vs-zookeeper.html
- burntsushi 13y agoThe package documentation for the `serf` library[1] looks really exciting. I've been wanting to make a distributed file synchronization tool, and perhaps this would be an excellent library to build it on. Question: as a relative networking idiot, how does NAT traversal fit into all of this? [1] - http://godoc.org/github.com/hashicorp/serf/serf http://godoc.org/github.com/hashicorp/serf/serf
- armon 13y agoWe designed the `serf` library to be able to be easily embedded, so hopefully it can be of some use. Unfortunately Serf does not make use of any sort of NAT traversal currently. We've open sourced the project hoping to get the community involved, and NAT traversal is something we'd gladly work with the community to get implemented.
- philips 13y agoThe first building block, STUN, is implemented over here: https://github.com/ccding/go-stun https://github.com/ccding/go-stun
- dugmartin 13y agoThey should add "from the folks who brought you Vagrant" to the top of the homepage.
- nemothekid 13y agoThis is actually very cool and is something incredibly handy for managing failover/membership. What I unfortunately don't understand is that there doesn't seem to be a library I can use to take advantage of this in my own application. If I have a program (in Go) am I expected to spin up my own serf process then communicate with it via socket? Is there an option for me to have serf live inside my application?
- armon 13y agoThe Serf executable is actually just a wrapper around the `serf` library. That library is designed to be embedded in Go applications. Documentation for the library is available here: http://godoc.org/github.com/hashicorp/serf/serf http://godoc.org/github.com/hashicorp/serf/serf
- smandou 13y agoHighly available?
- armon 13y agoYes, the availability of the system is not tied to any given node(s). Any node (or group of nodes) can continue to operate in the face of failure.
- peterwwillis 13y agoWhat happens to your cluster when your network experiences intermittent packet loss and your random UDP messages get lost? Nodes just start going down and up randomly? (For those of you going "So what, that's normal", this not a quality of an HA system)
- armon 13y agoI'd highly recommend taking a look at this page: http://www.serfdom.io/docs/internals/gossip.html http://www.serfdom.io/docs/internals/gossip.html. One of the great attributes of the gossip protocol is it is very robust to intermittent network failures. Under minimal packet loss conditions (<5%), the rate of false positives should be very low. This is due to a few techniques, one of which is indirect probing, and another is a novel "suspicion" mechanism. In the case of a network partition, the parts of the cluster can run in isolation and will recover when the partition heals. If you are interested, the paper referenced there ("SWIM: Scalable Weakly-consistent Infection-style Process Group Membership Protocol"), is the foundation of Serf. In the paper you can find more details about the behavior of the cluster, false positive rates under packet loss, and partition handling. tl;dr the systems is in fact designed with network errors in mind, as opposed to handling them being an afterthought.
- peterwwillis 13y agoWhat you're saying is it's designed with the knowledge that it's going to cause false positives, and basically doesn't work well under anything more than minimal packet loss. I think this is probably an important factor to note in the description (and I still fail to see how this is considered highly available or fault tolerant, as described in Intro pages)
- armon 13y agoI think we are maybe just working with different definitions. High Availability for Serf means that it can continue to handle changes in topology and deliver user events in the face of node failures and network problems. However, it is inevitable that there will be a degradation in it's performance given network failures. If there are serious packet loss issues, Serf will mark a node as failed. I'm not saying it "won't work well". It works as it is designed to. It will be available for operations, it will automatically heal when the partition recovers, and the state will be resynchronized with the "failed" nodes. The system will be in an eventually consistent state, which is expressly documented and is it's normal mode of operation. If you consider 5% packet loss "minimal", I'm not sure what applications you are running. TCP degrades at over 0.1% packet loss, and most UDP streaming protocols have serious degradation over 5%.
- kevinpet 13y agoThis would really benefit from a "how does this relate to zookeeper". I think this is an entirely new service, with different technical insides, and trying to provide a higher level solution to what people usually cobble together with ZK. But I'd be interested in comments from someone knowledgeable. Edit: I see this is addressed at http://www.serfdom.io/intro/vs-zookeeper.html http://www.serfdom.io/intro/vs-zookeeper.html but it would be nice to have something more "just the facts" rather than arguing the serf is good.
- armon 13y agoIn writing that section, we tried to provide "just the facts". If there is anything that seems wrong or misleading in any way, we'd like to know so that the page can be corrected. It is not our intention to say "Serf is good, ZooKeeper is bad". They are very different tools, and we are just trying to highlight the differences. In fact, we believe that the strongest use cases involve using those tools together.
- kevinpet 13y agoI didn't mean to say that it sounded like a sales pitch. What I meant is that it's talking about relative strengths and weaknesses, where what I really need to understand what Serf is is more along the lines of how the API / model differs from ZKs notion of writing to or waiting on locations in the distributed space.
- totoy 13y agocan we see serf as a kind of riak core but written in go?
- armon 13y agoRiak Core provides a superset of the features of Serf. Riak Core uses gossip to manage membership, but it also provides quorums for coordination, and is based around the notion of a hash ring and virtual nodes. You could instead using Serf to build riak core like technology on top.