4 ms·
I think the biggest knock against serverless is the same as microservices: your points of failure grow significantly.
by everdev 8y ago
I think the biggest knock against serverless is the same as microservices: your points of failure grow significantly.
- moduspol 8y agoThe points of failure you are responsible for maintaining decrease significantly, though. And the ones you aren't responsible for are ones that the rest of the world uses every day at 1000x the scale you do.
- staticassertion 8y agoI don't see how. Your points of failure were always there, but at most they're more explicit since "service failure" conflates with "network failure" - something I think is beneficial, since you always had to handle "service failures" but they were implicit. That is - a piece of code in a monolith may fail, and a piece of code in a microservice may fail, but I've already written the code to handle network errors for the microservice case, which means I've also implicitly handled the "bug in code" case.
- SahAssar 8y agoThat assumes that a local function call has the same failure rate as a remote function call to a microservice, which in my experience is very much not true. If I have a local function in the same language I can pretty much assume that a call to that function will actually call that function. With a remote call over HTTP or whatever I can't, so that is an additional failure I need to handle.
- staticassertion 8y agoI'm not sure I understand. Why would an RPC call a different function than what you expect? I'll grant you that there is more complexity in this approach, but I believe that fault tolerance is something you improve with microservices, not something you regress on.
- squeaky-clean 8y agoIt's not that it would call a different function, but that sometimes the RPC will fail to call the function. You can't get a network error calling a function which is in memory on the same process.
- staticassertion 8y agoThis is what I was saying though. Yes, you have to handle the failure case of "network call". Microservices add this failure case. But you already had to handle the case of "code blew up because of a bug". By forcing you to handle comm errors like network failure, you also force people, implicitly, to handle "code blew up because of a bug" errors. Even though it adds a second error case, you pretty much handle them the same way in the same place. There was always an error case - the fact that code may blow up.
- fauigerzigerk 8y ago>But you already had to handle the case of "code blew up because of a bug". I'm handling bugs very differently than network failures though, because network failures are usually temporary while bugs are usually (or even by definition) permanent. Dealing with temporary outages in lots of places is extremely difficult. You may need retry logic. You may need temporary storage or queuing. You may need compensating transactions. You may need very carefully managed restarts to avoid DDOSing the service once it comes back online. There may be indirect, non-deterministic knock-on effects you can't even test for properly. Microservices cause huge complexity that is very hard to justify in my opinion.
- staticassertion 8y ago> I'm handling bugs very differently than network failures though, because network failures are usually temporary while bugs are usually (or even by definition) permanent. Depends on the bug - there are transient bugs that are not networking related. But let's assume it's a "hard error" ie: a consistently failing bug. I would say where that bug is makes a huge difference. If it's a critical feature, that bug should probably get propagated. If it's a non-critical feature, maybe you can recover. By isolating your state across a network boundary, recovery failure is made much simpler (because you do not need to unwind to a 'safe point' - the safe point is your network boundary). But it often depends how you do it. I personally prefer to write microservices that use queues for the vast majority of interactions. This makes fault isolation particularly trivial (as you move retry logic to the queue) and it scales very well. If you build lots of microservices with synchronous communications I think you'll run into a lot more complexity. Still, I maintain that faults were already something to be handled, and that a network bound encourages better fault handling by effectively forcing it upon you.