7 ms·
Orchestration tools (puppet/chef) are already able to get your infrastructure to a target/desired state, and keep them in that state, and notify when deviations
by amoghe 11y ago
Orchestration tools (puppet/chef) are already able to get your infrastructure to a target/desired state, and keep them in that state, and notify when deviations occur or the target state cannot be achieved. What does StackStorm do that these tools cannot?
- lotyrin 11y agoThis is for the guy who gets a support call because of a disk getting full, and has to ssh into a box and delete old log files because logrotate is fubared by some other team but he finds out that there's a nagios monitoring everything (with yet another team ignoring the noise) so he wants to just have his bash oneliner for deleting old log files run any time the disk monitor hits critical, and all he has to do is sell someone with a purchase card a SaaS app (easy), and doesn't have to sell his entire organization the concept of not being fuckups (hard). It's sad how big the market for this is.
- bigdubs 11y agoI don't like being the voice of the purist in this, but this seems like a bandaid on a bullet wound. For most of the cases where this would seem to be useful there is probably a failure upstream of that usage that should really be fixed.
- doriftoshoes 11y agoGiant disclaimer: I work at StackStorm...but I also have an extensive Ops background. This is really the next step in runbook automation. It gives users a way to express procedures and operational patterns in code (the workflow definitions are in yaml). With any sort of automated remediation there is always the concern of "painting over the mold" but at the same time you don't want to get stuck doing a large number of manual steps when you could be focusing your energy on tracking down the root cause of the issue and resolving that. The more important aspect to me personally is the easy of version controlling these ops patterns. Store the workflow definitions in source control and it is easy to diff the changes in your procedure.
- mst 11y agoIt seems YAML is the new S-expressions for people trying to pretend they didn't actually invent a programming language. This is, I guess, less annoying than executable XML.
- doriftoshoes 11y agoYAML is definitely the hotness right now but it works. No need to invent a full blown language for something like this. Way easier than trying to write json.
- mst 11y ago> Way easier than trying to write json. Which is why ingy and I invented http://p3rl.org/JSONY http://p3rl.org/JSONY for config files.
- eropple 11y agoOr you can use Ruby or Python and Perl and have a scripting language that looks and acts as a scripting language is generally expected to look and act. (And with Ruby in particular it's very easy to provide a flexible and terse DSL that provides significant benefits on its own.) This "invent your own worse programming language" fad is disappointing.
- mst 11y agoFor plain config, I mostly use JSONY now. For executable declarationish things, I like Tcl (usually embedded in perl so I can use Moo(se) for OO rather than the inferior crap in other languages). Hashicorp's HCL is an interesting middle ground.
- eropple 11y agoI can respect Tcl. I'm not super familiar with it aside from hacking on eggdrop, but it's in the same ballpark. For configuration, I just source a shell script and grab env vars wherever I can. It's the most portable (between applications, not necessarily between platforms) option available to me and plays nicely with stuff like jails. At a glance, I don't quite understand the value prop of JSONY over YAML, which (as you see used in something like Rails) allows for a good bit of flexibility with references to more tersely communicate intent. I think HCL is one of the worse decisions made by a somewhat influential software company in quite some time. It's harder to use than YAML (which is funny because the stated reason for its existence is "YAML is harder", and it makes me really, really curious what sort of users they're polling) while being less expressive and they're very careful to not really care about anybody's interop "because you can just use JSON instead", except that, as trying to work with Terraform amply proved, no, you can't use JSON, because that path isn't tested because nobody cares and unit tests are hard. =( HCL is a regular problem in my life. I needed to vent a little.
- mst 11y ago> For most of the cases where this would seem to be useful there is probably a failure upstream of that usage that should really be fixed. Of course. But in the mean time, it would be nice to keep production up.
- ewindisch 11y agoMonitoring exists because software is not perfect and will never be perfect. Software raises errors because it doesn't know what to when those events occur. Operators usually know what needs to be done for their specific environment, even if the software they're running doesn't. If we could eliminate the bullet wounds we wouldn't EVER need logs. As long as we have logs, we need some way to process them and react to their events. For many organizations, for decades, this has been to have an alerts system and an active operations team fighting fires. Those teams maintain knowledge-bases to track institutional knowledge of how to manually react to these events. Solutions such as Stackstorm and IBM ZAware seem to be created to allow that institutional knowledge to be automated. I've also seen (and built) proof-of-concepts using Bayesian filters as part of such systems. It's been a long-time coming and I'm happy solutions are evolving to address this need. Finally, I think that bandaids such as this may be, might be precisely what is needed in some cases. Assuming it's even a legal or technical possibility, the cost of the "right fix" may far exceed the harm of leaving a bug unfixed. Sometimes the right fix has a large time-cost, which automation can help bridge the gap for operations teams while a permanent fix is developed (or more hardware is acquired, etc).
- ascendantlogic 11y agoAnd how do you universally fix the bullet wounds? How do you stop every upstream issue from ever occurring so middle-of-the-night remediation doesn't have to happen? Of course everyone would love the actually solve every problem but real life is messier than that. For every issue that needs to be fixed, there's usually one or more reasons why they can't be fixed the "right way" right this instant: Management bullshit, prioritization, $$$, etc etc.
- spdustin 11y agoYou aren't kidding. I thought it was a clever framework for other kids of "ChatOps" in addition to all that you said. I can imagine some clever Hubot scripts taking advantage of some of the components. Man, I sure do love my Hubot.
- jamesfryman 11y agoDisclaimer: I am an employee at StackStorm Tools like Puppet and Chef are great at managing the state of a single node. However, with any piece of infrastructure that expands beyond a single node, things can get complex pretty quickly. In times where the state of the systems change, there often are a multitude of steps that need to occur to ensure that the intended state is met. However, because tools like Puppet and Chef are node-centric, you often find yourself waiting a long time for eventual convergence as a code is executed for a node, updated state data transferred to some upstream data server (Chef server, PuppetDB), and then other nodes converge with updated data. Depending on the task, this convergence time can be killer. In contrast, StackStorm is an event-driven automation framework that will help perform the incremental tasks across many systems necessary to properly move from one state to another. Common examples include ensuring that Load Balancers are up before advertising network services, ensuring SQL standby servers are alive before enabling replication, and so forth. Each of these steps is going to require an imperative set of steps to transition between states. With StackStorm, we plug into a multitude of tools (including Puppet and Chef!) that provide event updates as actions take place, intercept these triggers, and execute workflows. In many cases, we have clients that heavily use a configuration management tool like Puppet or Chef, but rely on StackStorm to orchestrate the various runs of these tools on different nodes as checkpoints are reached. In this way, decrease the feedback loop as you have StackStorm listening for "run finished" notifications from these CM tools, and then we can go and figure out what is next by kicking off actions or workflows as necessary. There is absolutely a ton more about StackStorm beyond this immediate answer. In addition, we include things like Role Based Access Control, ChatOps support, full audit trails, and more. I encourage you to check it out and provide feedback. We'd absolutely love to help!
- helloiamaperson 11y agoYou mention Puppet and Chef, but how does it compare to more holistic tools like terraform, juju, and bosh?
- dzimine 11y agoterraform/juju/bosh are purpose-build for app and infra deployments. StackStorm is a generic automation platform with no . One can use it to run arbitrary chain of actions on events. As such, it is used in auto-remediation & automating runbooks. We internally use it for variety of things like irc-to-slack relay, zombie ec2 vm periodic clean-up, ChatOps-ing JIRA, etc. Some folks do run complex continuos deployment pipelines and blue/green deployments on StackStorm. Would be good if they can comment here.