4 ms·
Original title: A Note on Distributed Computing > A better approach is to accept that there are irreconcilable differences between local and distributed comput
by vaughan 4y ago
Original title: A Note on Distributed Computing
> A better approach is to accept that there are irreconcilable differences between local and distributed computing, and to be conscious of those differences at all stages of the
design and implementation of distributed applications. Rather than trying to merge local and remote objects, engineers need to be constantly reminded of the differences between the two, and know when it is appropriate to use each kind of object.
Aka. Transparent remoting.
- deleted 4y ago[deleted]
- dboreham 4y ago> Transparent remoting Presumably you mean "Transparent remoting isn't a thing"? Paper authors were at Sun, so had experience with NFS (the design of which assumed that transparent remoting was a thing).
- ninkendo 4y agoNFS is my favorite example of how to do distributed computing incorrectly. Take a filesystem API that assumes fast, local access, and that individual I/O operations are cheap and immediate (like stat()’ing a file or resolving a symlink), and suddenly make it so all these operations may have to block for hundreds/thousands of milliseconds while accessing something over a remote server potentially on the other side of the world. It causes so many issues: syscalls that can’t be interrupted (and thus you can’t ctrl+c a process), stuck mounts that can’t be unmounted, hanging for ages on `ls`’ing a folder with a lot of symlinks… NFS locking up is the most common answer to the typical interview question “what could cause a load average in the hundreds on a machine with zero CPU usage?”: because there’s a stuck NFS server and the pending IO operations put the processes in the run queue. I’m glad there’s a paper that articulates my thoughts about this subject better than I can.
- zozbot234 4y ago> It causes so many issues: syscalls that can’t be interrupted (and thus you can’t ctrl+c a process), stuck mounts that can’t be unmounted, hanging for ages on `ls`’ing a folder with a lot of symlinks… These things have nothing to do with distributed systems, they can happen on any machine. That's what the "Not ready reading drive A. Abort, Retry, Fail?_" prompt is for.
- dekhn 4y agoumount -l -f will unmount (logically) a stuck mount and let you mount something else there.
- DonHopkins 4y agoAnd before that there was /etc/fastboot or the L1-A two finger salute!
- pmontra 4y agoI hit that a couple of weeks ago after some 30+ years. I forgot how it was. And the per client per directory file entry cache. It doesn't play nice with systems with producers that create files and consumers that look for them. By the way, Amazon's EFS is based on NFS.
- _a_a_a_ 4y agoOk, so be constructive: what should an interface for remote file access look like?
- dekhn 4y agoOpen(path) -> handle, Read(handle, offset, size), Write(handle, size), SnapshotForAppend(handle, newpath), Close And write is write-once, from the start of the file.
- ninkendo 4y agoThe most important thing is that it should be asynchronous, and only asynchronous. A syscall like read() blocks by default (suspending the whole thread) until an NFS server comes back with a response. A more appropriate interface should not allow callers to operate in a blocking way. Blocking incurs a lot of problems, because the kernel ends up waiting a ridiculously long default timeout for the server to respond, and the calling thread is blocked the whole time. Making a remote-access API asynchronous (callback-based), or at least having O_NONBLOCK being the only allowed option, would help a lot here. But really, the other problem is that filesystems have too much baggage around assumptions of speed/locality/reliability that just blindly treating a remote server as if it’s a local filesystem is just a bad idea in the first place. Too much client software is written with the assumption that filesystems are fast (ie. you don’t need to cache it, stat()’s are cheap, etc) that it’s just better to not conflate the two.
- abecedarius 4y agoRemote pretending to be local is asking for failure. But local can trivially conform to abstractions designed for the remote case. Languages like Erlang and E showed that programming that way doesn't have to suck. When I read this paper (it's been a long time) it seemed to overlook this angle on the problem.
- mistrial9 4y agohmm maybe? the repeated references to Object-Oriented structures means partly, that 'code + data + an interface' exists systematically, as the basis for interaction. What you do with that to expose to the human mind is quite flexible (with costs).
- zozbot234 4y ago"Local conforming to the remote case" will be slow and have a lot of overhead, unless the "remote" abstractions are carefully designed to scale down efficiently to a single node. You can view the SSI (single system image) approach as trying to do exactly that.
- o_nate 4y agoThe paper specifically discusses that option (making local conform to remote abstractions) but finds it inadequate for various reasons. Stated succinctly: "Rather than encouraging the production of distributed applications, such a model will discourage its own adoption by making all object-based computing more difficult."
- abecedarius 4y agoThanks, you're right and my memory was fuzzy. But the strong claims like that in the paper are about trying to completely erase the local/remote difference, which neither E nor Erlang did. Instead they reduced the pain in a way that this paper leaves an impression is not worth pursuing: > Rather than using those resources in attempts to paper over the differences between the two kinds of computing, resources can be directed at improving the performance and reliability of each. > ... it is a mistake to attempt to construct a system that is “objects all the way down” if one understands the goal as a distributed system constructed of the same kind of objects all the way down. (E.g. in E the objects are the same kind; it's the references to objects that have two kinds.)
- dang 4y agoWe changed the title back to the original, in keeping with the site guidelines (see https://news.ycombinator.com/newsguidelines.html https://news.ycombinator.com/newsguidelines.html). Among other things, it's much better when searching for previous threads on the same article. (I found several previous submissions, but they didn't have comments). We also changed the URL from https://scholar.harvard.edu/waldo/publications/note-distributed-computing https://scholar.harvard.edu/waldo/publications/note-distribu... - might as well go straight to the paper. It's a good submission—thanks!