4 ms·
> programming for distributed environments as if you had one giant single-thread computer, and function calls could arbitrarily happen over IO or not Ex-Amazon
by ex_amazon_sde 6y ago
> programming for distributed environments as if you had one giant single-thread computer, and function calls could arbitrarily happen over IO or not
Ex-Amazon SDE here. Message-passing libs like 0mq tried to push this idea.
They never became very popular internally because passing data between green threads vs OS threads vs processes vs hosts vs datacenters is never the same thing.
Latencies, bandwidths and probability of loss are incredibly different. Also the variance of such dimensions. (Not to mention security and legal requirements)
Furthermore, you cannot have LARGE applications autonomously move across networks without risks of cascading failures due to bottlenecks taking down whole datacenters.
Often you want to control how your application behaves on a network. This is also a reason why OTP (in Erlang or implemented in other languages) is not popular internally.
- hcarvalhoalves 6y agoI don't dispute one might want to have this control. At the same time, I don't see a theoretical reason one can't have the same programming model, and choosing what computation runs local vs. distributed is just a configuration. A function call or a remote call won't change the domain logic, it's an implementation detail - but today, this leaks into all kinds of decisions, starting from microservices and working backwards to glue everything together again.
- 1penny42cents 6y agoThe programming model is fundamentally different between local vs distributed. Imagine if every local function call in your monolith significantly increased in error rate and latency. Your optimal algorithms would change - you can't just loop over a network call like you can with a local function call. That's what happens when you take a local invocation and put it across a network boundary. The probability of things going wrong just goes up.
- hcarvalhoalves 6y agoI'm not proposing to code like a monolith - but I don't see a reason there can't be a programming model that unifies that. APIs like Spark, for instance, make it largely transparent - processes are scheduled by the driver in nodes w/ hot data caches, and processes that fail are transparently retried; and yet this doesn't impact how you write your query. Effects systems are another thing can help us reason around i/o as function calls and handle the fail-ability. I would bet it's a matter of lacking good accepted patterns instead of a theoretical impossibility.
- 1penny42cents 6y agoIt's not a theoretical impossibility: you can wrap every function call in a monad today if you want. It's just a practical nightmare.
- ex_amazon_sde 6y agoIt's way more than "just a configuration". > A function call or a remote call won't change the domain logic Understanding performance and minimizing failure modes and their impact on larger systems makes for a whole career as "SRE". Making remote calls all over a codebase creates behaviors that are practically impossible to debug or optimize. But the blocker is the network impact of large applications and the emergence of cascading failures.
- 1penny42cents 6y agoThis point is what so many miss. The conceptual decomposition seems nice, but the physical implications of that decomposition are horrible. The further you split a function caller from its caller, physically, the more chaos you're asking for. It's quite simple to see from a physical perspective but from an abstract perspective it's all lines and boxes.