5 ms·
Failsafe – failure handling with retries, circuit breakers and fallbacks
- ckugblenu 10y agoQuite interesting. It shows potential to be used in numerous use cases. Anyone know of similar projects in other languages like Python and Javascript?
- sync 10y agoWe use re for Javascript, it works well: https://www.npmjs.com/package/re https://www.npmjs.com/package/re
- rekwah 10y agoAlthough, not feature parity with this project, Pybreaker[0] for the circuit breaker patterns in Python. [0] - https://github.com/danielfm/pybreaker https://github.com/danielfm/pybreaker
- rdli 10y ago(Full disclosure: co-founder of Datawire) We released a microservices development kit (MDK) last week that includes similar semantics (e.g., circuit breakers, failover) that implements these semantics in Python, JavaScript, Java, and Ruby. The implementation is actually written in a DSL which we transpile into language native impls. We do this to insure interop between different languages. We're working on updating our compiler to support Go and C#, adding richer semantics, and making the service discovery piece pluggable (currently there's a dependency on our own service discovery). https://github.com/datawire/mdk https://github.com/datawire/mdk
- Rauchg 10y agoWe use `async-retry` which implements `node-retry` in a way that's friendly to usage with `Promise` and `async/await`. https://github.com/zeit/async-retry https://github.com/zeit/async-retry
- deleted 10y ago[deleted]
- Rapzid 10y ago.Net has Polly https://github.com/App-vNext/Polly https://github.com/App-vNext/Polly
- garthk 10y agoSee also: Twitter's Finagle [1] for the JVM, and Bouyant [2] providing Finagle-as-a-microservice on localhost for language independence. 1: https://twitter.github.io/finagle/ https://twitter.github.io/finagle/ 2: https://buoyant.io https://buoyant.io
- nitrogen 10y agoVery cool. Consistent and clear retry, backoff, and failure behaviors are an important part of designing robust systems, so it's disappointing how uncommon they are. If I were starting a new Java project today I would almost certainly want to use this library instead of the various threads and timers I had to hack together years ago.
- heisenbit 10y agoIndeed this is conceptually hard stuff. The reason for that I believe is that the problems one is solving are system level problems and not local ones. Another way to look at this: It is the other guys problem. A lot of naive retry strategies sort of work until one has a larger number of clients to deal with. I still remember the time trying to get through to a base-station designer who refused to acknowledge the need to do exponential back-off and other mitigation steps. We ran into interesting times shortly later in the field on the management system side. Personally I would also put in a bit of randomness to spread out requests when all clients were initially impacted at the same time and were thus synchronized.
- jodah 10y agoGood example of where random retry delays would be valuable. I filed this as a feature to add for the next release: https://github.com/jhalterman/failsafe/issues/39 https://github.com/jhalterman/failsafe/issues/39
- fdsaaf 10y agoBeware of runaway retries: https://blogs.msdn.microsoft.com/oldnewthing/20051107-20/?p=33433 https://blogs.msdn.microsoft.com/oldnewthing/20051107-20/?p=... Personally, I'd rather systems fail quickly, with retries only at the highest (application) and lowest (TCP) levels.
- dredmorbius 10y agoA note on the name: "fail-safe" in engineering doesn't mean that a system cannot fail, but rather, that when it does, it does so in the safest manner possible. The term originated with (or is strongly associated with) the Westinghouse railroad brake system. These are the pressurised air brakes on trains, in which air pressure holds the brake shoes open against spring pressure. Should integrity of the brakeline be lost, the brakes will fail in the activated position, slowing and stopping the train (or keeping a stopped train stopped). https://en.m.wikipedia.org/wiki/Railway_air_brake https://en.m.wikipedia.org/wiki/Railway_air_brake Fail-safe designs and practices can lead to some counterintuitive concepts. Aircraft landing on carrier decks, in which they are arrested by cables, apply full engine power and afterburner on landing. The idea is that should the arresting cable or hook fail, the aircraft can safely take off again. https://en.m.wikipedia.org/wiki/Fail-safe https://en.m.wikipedia.org/wiki/Fail-safe Upshot: "fail safe" doesn't mean "test all your failure conditions exhaustively". It may well mean to abort on any failure mode (see djb's software for examples). The most important criterion is that whatever the failure mode be, it be as safe as possible, and almost always, based on a very simple and robust design, mechanism, logic, or system. From the description of this project, it strikes me that it may well be failing (unsafely?) to implement these concepts. Charles Perrow, scholar of accidents and risks, notes that it's often safety and monitoring systems themselves which play a key role in accidents and failures.
- pm90 10y agoThis is a great comment. Why do you think the project fails to implement the concepts that you mention?
- daenney 10y agoI think that what they're trying to get at is that having libraries that (for example) wrap failures in retry modes isn't necessarily failing safely. It can very well obscure problems in your implementation or other parts of the systems you're talking to. Having it fail safely can just as well be "abort execution" and visibly log it so as to raise the problems with those that might be able to solve the root cause. There's certainly something to be said for retry strategies in places that involve a lot of network chatter but please don't also forget to add some kind of back off to it so you don't end up retry-overloading a system that's trying to recover.
- ap22213 10y agoIt seems like a well-thought, fluent interface to what lots of Java developers (especially Java 8 ones) inevitably have to write themselves.
- SwellJoe 10y agoThis title would be 100% better with "for Java" on the end.
- _Codemonkeyism 10y ago... for JVM languages.
- cpitman 10y agoHow is this distinct from Hystrix (https://github.com/Netflix/Hystrix https://github.com/Netflix/Hystrix)? Why should I use one over the other?
- jodah 10y agoGood question. Someone asked that recently on Github - here's a quick comparison: https://github.com/jhalterman/failsafe/wiki/Comparisons#failsafe-vs-hystrix https://github.com/jhalterman/failsafe/wiki/Comparisons#fail...
- vikiomega9 10y agoIs there a more detailed comparison? For example, >Executable logic can be passed through Failsafe as simple lambda expressions or method references. In Hystrix, your executable logic needs to be placed in a HystrixCommand implementation It's not apparent to me what the advantage of either interface is. In both situations I have to define a "lambda" and hold state somewhere(either as an object field or passed into the lambda). Unless I'm something here, either seems acceptable.
- jodah 10y ago> Is there a more detailed comparison? There's nothing more detailed that I know of. Is there a particular feature area/comparison you're curious about? I can add a bit more detail. > It's not apparent to me what the advantage of either interface is. In both situations I have to define a "lambda" What I meant by this bit is that the user experience is different. Failsafe can be used with method references or lambda expressions [1], which are a nice, concise way of wrapping executable logic with some failure handling strategy. You cannot do this with Hystrix since all logic must be wrapped in a HystrixCommand impl, which cannot be implemented as a lambda. > either seems acceptable. Like anything, it just depends on what you want. If retries and general purpose failure handling, consider Failsafe. If request collapsing, thread pool management and monitoring, consider Hystrix. [1]: https://github.com/jhalterman/failsafe#synchronous-retries https://github.com/jhalterman/failsafe#synchronous-retries
- mandeepj 10y agoPlease find some of these patterns for .net\azure\c# stack here - https://msdn.microsoft.com/en-us/library/dn568099.aspx https://msdn.microsoft.com/en-us/library/dn568099.aspx