3 ms·
This is clever, and the incremental approach is a good engineering direction. That said, readers should be aware that the closures approach is a dead end for a
by timdierks 12y ago
This is clever, and the incremental approach is a good engineering direction. That said, readers should be aware that the closures approach is a dead end for a large & complex service: unless you can figure out how to serialize the full closure state to data that can be exchanged between machines, you're always going to be RAM-constrained and subject to machine failure.
The loss of state isn't a big deal for a news site, but if you lost the user's state in the middle of a significant process (e.g. dropping the cart mid-checkout), it would be a big problem, even if it was infrequent.
What would be great if a language provided a convenient way to repose closures into data blobs that could be exchanged between machines or sent via the browser (an untrusted channel that would have to be secured with cryptography). That said, it's not obvious how to capture the semantics of the full closure of the system state without significant work (look at how fragile object serialization / pickling is). Definitely a good area for language research; I'm sure there's some advances I'm not familiar with.
- lysium 12y agoOne difficulty with storing closures is storing and transferring functions, in particular between machines. There is progress in this direction, but it is not easy.
- drostie 12y agoOne thing that was exciting about Datomic to me (and still is, I just don't have any projects big enough to use it for) is that you get an immutable database state. You can emulate this in any normal database, but it gets harder to enforce the discipline among other members of your team. The basic idea is that you can serialize those same closures by just pointing to "the state of the database at revision #5968", and then, though the database moves on, you can always use that ID to compute the view of that database at that point. It does the same "heavy lifting" that storing these closures is doing, but you can easily share that ID across a distributed service with no "expired links" problems. It's worth mentioning that you can't send a closure in a non-functional context. That is, if Alice sends a closure to Bob, it cannot any longer be the case that Alice's other operations can mutate Bob's state. So you must be serializing an "orphaned" environment tree with a bunch of closures which point at different nodes of that tree. You could definitely do this even better by stealing some ideas from Smalltalk: encapsulate all of the states in some computational node (the original notion of "object" in OOP) which interacts with all of the other parts of the system communicate with by message-passing, and nothing else. To change the code on-the-fly, you just swap out the "code part" of some node for a new code part, and perhaps transform the state, queuing up the messages while you do so; then you can start replaying those messages to the new code. The benefit is that now at any time the nodes can move around servers arbitrarily, as long as you've got a good name-resolution service to tell you where the object is now. In other words: (1) interpret all the things so that code and data are the same; (2) shared-state is your enemy; serialize orphaned states only; (3) you have to explicitly handle the case where someone makes a request while you are sending their closure to another server.