5 ms·
> Network boundaries cause far more problems than they solve, They cause exactly 0 extra problems. A call from function A to function B can fail due to B havin
by staticassertion 4y ago
> Network boundaries cause far more problems than they solve,
They cause exactly 0 extra problems. A call from function A to function B can fail due to B having a bug. A call from service A to service B can fail due to B having a bug or a network failure. Either way, failure is possible and has to be handled - the network only makes that more obvious.
Further, a call between functions can cause mutated shared state - not the case across a boundary, they physically do not share mutable state.
> and you've just shifted the complexity to now securing the network
Not really. Fundamentally you have split your service capabilites up - now you can apply least privilege as you desire.
- blowski 4y agoYou drop in “…or a network failure” like it’s a rare occurrence that’s easy to handle.
- staticassertion 4y agoFrequency is irrelevant to the complexity. As I said, you have to handle the idea of cross-boundary failures, such as with modules that have bugs. If you aren't taking steps to do so, you're not writing robust code. Anyway, yes, persistent network failures are rare for many people.
- blowski 4y ago> Frequency is irrelevant to the complexity So if something happens 50% of the time you should treat it in the same as if it happens one in 100 billion?
- staticassertion 4y agoMaybe? I can't compare 50% to 100 billion. If your computer crashed every 100 billion instructions that would be a problem. If it crashed every other instruction, that would be very slightly more (or the same amount) of a problem. The point is that if you have a function call, you have the opportunity for a bug/ failure. Networks don't change that - you have the opportunity for a bug/ failure. The major difference is that services have stronger failure isolation.
- blowski 4y agoIt sounds like you’re designing a very hypothetical bit of software.
- staticassertion 4y agoAs opposed to designing software that is already implemented?
- travisd 4y agoAt some point, you can't gracefully handle bugs in other peoples code. If a function you call causes a SEGFAULT, in the vast majority of software, you're not expected to handle that. That's an invariant error, and you probably want some way to detect that it happened so you can fix it, but it's not reasonable to ask every caller of every function to handle that (in the same way we don't consider "the earth blew up" to be a reasonable thing to protect against, even though it its technically possible). There's simply not enough time and money to protect against every possible edge case in most software (NASA projects aside). The argument here is that network issues are exceedingly common in microservice environments and so aren't actually an edge failure case, so you actually have to worry about them way more than you would worry about a function in a different module causing a SEGFAULT.
- staticassertion 4y agoThe point is not to handle individual bugs, it is to handle all failures. This is the difference between a "defensive programming" approach and the "let it crash"/ "zen of erlang" approach. Actors are designed such that they have failure isolation, which means they can react to errors in other actors without worrying about their own state. They then have two options based on one of two bug classes - transient and persistent. Persistent errors are propagated to the supervisor. Transient errors are either retried or propagated. It doesn't matter if it's a network error, a disk error, a timeout, a crash, a cosmic radiation bit flip - your approach is always one of those two. So adding more failure cases doesn't "matter" in terms of your error handling, although you may want to adopt helpful patterns in the nuances of "retry". The frequency of errors will obviously increase with a network error (arguably very very little), but the pattern is fundamental to resiliency. If your network is truly so unreliable that you can not pay that cost, don't do it. I don't think most people are developing on networks that fail for long periods of time frequently.
- kaba0 4y agoBut now you are talking squarely about Erlang actors, not microservices in general. The runtime gives you all the needed guarantees here.
- ori_b 4y agoNetwork calls can appear to fail, but actually succeed. Local function calls don't have to deal with byzantine failures[1]. [1] https://en.wikipedia.org/wiki/Byzantine_fault https://en.wikipedia.org/wiki/Byzantine_fault
- staticassertion 4y agoFunction calls can, of course, fail with side effects. Idempotency is always a desirable property.
- nicoburns 4y agoWith a monolith you can put everything inside a database transaction and have a an entire request's worth of logic succeed or fail together. That's a lot easier to manage that having parts of the logic spread over multiple systems succeed and other parts fail.
- staticassertion 4y agoSo use transactions? I don't understand what part of microservices prevents that. In fact, transactions are pretty fundamental to reliable systems. https://www.hpl.hp.com/techreports/tandem/TR-85.7.pdf https://www.hpl.hp.com/techreports/tandem/TR-85.7.pdf From the abstract: >It is pointed out that faults in production software are often soft (transient) and that a transaction mechanism combined with persistent process-pairs provides fault-tolerant execution -- the key to software fault-tolerance.
- nicoburns 4y agoWell you can't use database transactions across multiple connections, so presumably this would involve you implementing your own transaction and rollback system. That's a lot more complexity than using a system that just works out of the box.
- staticassertion 4y agoI don't really understand. If I have a database, and a service is talking to it, it can open a transaction. If I then want to talk to other services, and rollback that transaction based on what happens with those, I can do that. Microservices changes nothing about this. If you want to remove transactions by splitting up your logic such that it operates in terms of sequences or something, you can do that, but that's just a choice like any other.
- nicoburns 4y agoIf the other services you talk to mutate state then rolling back those changes is non-trivial.
- manigandham 4y agoA failure is obvious all by itself. Network boundaries just turn it into a much bigger failure. And network failures are far more common and harder to test, handle and recover from. > "Further, a call between functions can cause mutated shared state - not the case across a boundary, they physically do not share mutable state." This is false as state is not tied to your process nor does it require a network leap to add isolation. > "now you can apply least privilege as you desire." How exactly? It's not magic, you still have to apply them, and now it requires more strategies and effort to accomplish.
- staticassertion 4y agoI'm ignoring the first two points since I'm tired of explaining these things to people - you can read the papers/ watch the talks I've linked. > How exactly? It's not magic, you still have to apply them, and now it requires more strategies and effort to accomplish. Yes, we have tons of tooling for process isolation. Splitting a service into two services means you can isolate two processes instead of one, which means you break up the capabilities unique to each. I used the word "apply" so I don't know why you're saying "you still have to apply them"... it's literally what I just said.
- manigandham 4y agoSecuring the network is more work than just securing the code, because now there's a network in the way. For all the repetition you have on this thread, can you summarize it with the actual benefit that you have gained in a serious production use?
- staticassertion 4y ago> Securing the network is more work than just securing the code, because now there's a network in the way. I very much disagree. > For all the repetition you have on this thread, can you summarize it with the actual benefit that you have gained in a serious production use? I already have summarized it. All I've done since is correct people being incorrect with regards to my summary. If you want a specific example, here's a blog post I wrote a long time ago (the dates are incorrect since we moved websites): https://www.graplsecurity.com/post/architecting-for-performance-and-security-at-the-same-time https://www.graplsecurity.com/post/architecting-for-performa...