10 ms·
Gremlin Free – Run chaos experiments to prevent outages
- lklig 8y agoHey folks, I work at Gremlin and we're super excited to announce this launch. Drop any questions, comments, or concerns, we're happy to help!
- mpls 8y agoThanks for putting this out! I caught a Gremlin talk at a recent conference and was very impressed with how knowledgeable the developers were. How does the Gremlin platform interact with one of my hosts? Do I need to install an agent or something? Does it need root access to my host, hypervisor, cloud console?
- lklig 8y agoSimply install an agent, authenticate with our control plane, and create attacks through our webapp. No root access required. Check out more info at https://www.gremlin.com/docs/infrastructure-layer/installation/ https://www.gremlin.com/docs/infrastructure-layer/installati... .
- nodesocket 8y agoIf root is not required, how does the agent issue a shutdown or restart?
- DuvalKingJabub 8y agoHello! I am a Solutions Architect for Gremlin. Great question! It uses four Linux Capabilities to accomplish that listed here: https://www.gremlin.com/docs/security/overview/#linux-capabilities https://www.gremlin.com/docs/security/overview/#linux-capabi...
- dkersten 8y agoI like the look of this and love that you have released a free version. I am a little dismayed, though, that the two options are $0 and $1000/m (paid annually) with nothing in between. The free version seems great to get started, but I'd really like a lot more of the attacks that the paid version has, but $12,000 is much, much too high a price for a startup or personal project. That's quite a jump in cost.
- TheSpiceIsLife 8y agoI can’t speak for this vendor in particular, but one common reason for pricing like this is the vendor doesn’t want to deal with smaller customers as they often have the highest support requirements.
- TheAceOfHearts 8y agoPretty much. An enterprise customer won't bat an eye at $12K/year, and I imagine it'd pay for itself pretty quickly. I can definitely relate with GP, though. It feels frustrating to learn about an interesting product only to find that it's priced way outside of your budget. At least Gremlin has a public sticker price. Sometimes enterprise services just skip that completely and require you to setup a call with someone in their sales department, which usually means the service is outside of your budget.
- TheSpiceIsLife 8y ago> which usually means the service is outside of your budget. I wonder if this is true though. Maybe it puts off people who otherwise might be able to have a product tailored to their budget???
- dkersten 8y agoIts possible. I know a guy who isn't put off buy "call us" pricing notices and he seems to get good deals, so its certainly possible. Many of us will never find out though, because "call us" means "close tab". Even if I really want their product and am willing to pay a lot for it, my time is too precious to me to waste on calls.
- gfs 8y agoI see that you are on the Rust production user page [0]. Can you talk a little bit about what Rust is used for and how the experience has been? [0]: https://www.rust-lang.org/production/users https://www.rust-lang.org/production/users
- philgebhardt 8y agoHey, I'm an engineer at Gremlin! When you install Gremlin onto your linux hosts for infrastructure experiments, you're using binaries that were completely written in Rust. I would be lying if I said there wasn't a bit of a learning curve (coming from mostly working with Java). Most of that can be attributed to the memory management concepts built into Rust. At first you fight the compiler a bit (asking things like, why am I not allowed to reference this variable?!), but you soon learn to love and rely on the compiler as it builds more confidence in the runtime behavior of the product. One game changer for Rust is the treatment of Errors as first class citizens. It's literally built into the native types that Rust wants you to work with. That's huge for our product, given it runs in an inherently error-prone environment.
- gfs 8y agoThanks for the reply. I always anticipated Rust being a good fit for a daemon like tool. Not having to install a separate runtime and have things statically linked is a nice benefit. I know it's not the only language that is capable of this but being able to leverage the other bits of Rust helps with productivity as well.
- ksmail99 8y agoHi, we are startup using a lot of lambda, fargate, rds and dynamodb. Will gremlin work for this? I didn't see any mention of support of fargate or lambda on your website.
- lklig 8y agoWe've got you covered! Gremlin supports severless products with application layer fault injection. Take a look at our docs for more: https://www.gremlin.com/docs/application-layer/overview/ https://www.gremlin.com/docs/application-layer/overview/
- ksmail99 8y agoThanks that will work nicely with our lambda functions which are in Java. How about python? We are running python django in fargate. so it is possible to bring up a new container or add gremlin in the existing container. Is this possible?
- lklig 8y agoGlad to hear it! Additional language support is top of mind. Node is up next and python is high on the list. Regarding your app running in Fargate, you can do either. Hop over to our #support slack channel and we can help out more! https://www.gremlin.com/slack https://www.gremlin.com/slack
- keyle 8y agoI laughed out loud at "Failure as a service". Thanks for that.
- ingrid 8y agoHa, after working on building Uber’s chaos monkey (which was hard and took a while to build) and working with Netflix’s chaos monkey — it’s super nice to see Gremlin release this service so anyone can see the benefits of chaos engineering. I hope they add a “random chaos” feature to keep engineers on their feet. ;-)
- espeed 8y agoNB: This company "Gremlin, Inc", its product "Gremlin Free", and its use of the Gremlin name is in no way affiliated with or related to Apache TinkerPop™ Gremlin, its ASF marks, name, the open-source Gremlin graph programming language, ASF TinkerPop Gremlin Graph Traversal Machine (GSM), associated libraries, or the Gremlin Graph developers group formed in 2009. http://www.apache.org/foundation/marks/faq/ http://www.apache.org/foundation/marks/faq/
- yodon 8y agoIt's also probably not related to the 1984 movie "Gremlins" or the 1970's car of the same name (listed as one of the ugliest cars of all time[0]) [0] https://www.cbsnews.com/pictures/worlds-15-ugliest-cars/7/ https://www.cbsnews.com/pictures/worlds-15-ugliest-cars/7/
- espeed 8y agoASF marks were filed at formation. Identifying and distinguishing the use of similar and potentially confusing marks (esp in software products) is one of the required duties. The use of ASF marks in software products is prohibited to prevent confusion.
- deleted 8y ago[deleted]
- yodon 8y agoTinkerPop is a graph database. This is an infrastructure testing system. The likelihood of confusion between those two use cases is comparable to the likelihood of confusion with a movie or a car. You have a trademark, that's awesome. Even McDonalds trademark, which is one of the broadest, is scoped.
- espeed 8y agoI was confused when I first saw the project, cartoon logo and "Gremlin Free" name, and I'm intimately familiar with both the Apache TinkerPop Gremlin open-source project, its third-party libraries, and the extensive ASF legal process we went through registering and identifying the use of marks. Read up on your trademark, IP and copyright law. The use of similar marks is not permitted when there is potential for confusion, such as two different software projects, esp with overlapping audiences. NB: TinkerPop is NOT a "graph database", it is a collection of software libraries for connecting to, using, and managing graph databases and distributed processing platforms. Gremlin is a primary part of that stack -- Gremlin proper is the programming language, and the Gremlin GTM is the runtime.
- djb_hackernews 8y agoIs anyone aware of a chaos tool that isn't a SaaS (free or not) and doesn't require using Spinnaker like the current Netflix chaos tool does?
- lklig 8y agoYes, we compiled a list of all the OSS alternatives to Chaos Monkey here! https://www.gremlin.com/chaos-monkey/chaos-monkey-alternatives/ https://www.gremlin.com/chaos-monkey/chaos-monkey-alternativ...
- jedberg 8y agoThe old chaos monkey didn’t require spinnaker. You can find it here: https://github.com/Netflix/SimianArmy https://github.com/Netflix/SimianArmy
- isuckatcoding 8y agoHow do you prevent abuse of this tool?
- lklig 8y agoSecurity is extremely important to us. Clients authenticate to our control plane either with a secret string or a certificate. Clients can be revoked at any point from our webapp and as well if the client loses communication to our control plane, any ongoing attack is halted. Check out our security page for more: https://gremlin.com/security https://gremlin.com/security
- Negitivefrags 8y agoSo here is what I don't get about this stuff. What happens to the in-flight requests? Don't a few users run into random errors whenever a host is killed unexpectedly? You could have your loadbalancer retry everything that fails, but then wouldn't every single request in your app have to be idempotent?
- LoSboccacc 8y agowell for example in our systems all api calls only moves from a know state to another known state and any call failure redirects the client/user to the dashboard trough an error handler so they have to reload the last good state saved on the database. not perfect, but having a server crash is not much different than having a connection reset by a wifi status change or an upload timing out due the mobile network going away or the user navigating away or closing the browser.
- Negitivefrags 8y agoIt sounds like you are saying "The in-flight requests fail" to me. I really don't like the idea of saying that it's simply okay to give random users a bad user experience like that when you are actually killing servers yourself all the time.
- dbaggerman 8y agoIt's a different approach to managing risk -- minimizing impact of failure rather than minimizing the likelihood of failure. It's nice to know that you can kill a process and the only impact is that in-flight requests fail, rather than having a more significant outage if a process crashes and the failover doesn't work, or the process doesn't automatically restart, etc. If you accept that requests will fail you can build retries into the system. It's a lot harder to make a system more resilient if you avoid testing the failure scenarios.
- lklig 8y agoExactly! Chaos engineering is all about thoughtfully planned out experiments, to observe what the user experience will be when something fails. Doing this on your own terms allows you to improve the experience so that your customers aren't affected. You can decide what happens when an in-flight request is dropped, whether you hold onto the state somehow and retry or the client could fail gracefully with a relevant error message.
- big_whoop 8y agoIt's reached a point where I actually want outages, because I just don't fucking care anymore. Burn everything to the ground. None of the perpetual availability is actually productive. Most of these websites are bullshit. If an outage occurs, five nines out of ten, the lost money is hypothetical, not yours, and founded on a dubious hypothosis of how advertising hypotizes and subliminally influences all the mindless simpletons we presume our users to be.
- goldenkey 8y agoFailure as a service doesn't make all that much sense considering that a many failure scenarios would make the target host inaccessible to Gremlin. How does Gremlin handle this?
- lklig 8y agoGood question! All of the network attacks have a whitelisting capability, to keep the host accessible. This isn't an issue with state attacks, as the client will come back online once the host reboots. And with resource attacks the client typically remains active, if your application is handling starved resources well.
- debaserab2 8y agoWhat infrastructure size does one need to have where this technique is beneficial? Genuinely curious where the threshold is.
- farazbabar 8y agoMultiple criteria: 1. When you go from one machine running the code to more than one 2. Any system that may experience failures and detection of such failures and recovery is desirable 3. Most distributed systems due to the failure scenarios inherent in such systems.