4 ms·
A similar approach is so-called "object prevalence" (http://en.wikipedia.org/wiki/Object_Prevalence). http://en.wikipedia.org/wiki/Object_Prevalence). Basicall
by dk 19y ago
A similar approach is so-called "object prevalence" (http://en.wikipedia.org/wiki/Object_Prevalence). http://en.wikipedia.org/wiki/Object_Prevalence). Basically, keep all the data in RAM and write a journal of changes. On startup, the journal is played back, incrementally restoring the state. Snapshots are taken periodically to keep the journal down to a manageable size.
It's a form of the Command design pattern and has some nice properties. You can get transparent thread synchronization by executing queries in parallel but serializing commands that modify state. For web apps, the HTTP request offers a natural representation for commands (but of course you'd want to strip them down to their essence). You can get fault-tolerance and scalability by feeding the command stream to replica servers. (State-changing commands must be executed by the master server but queries can be load-balanced across the replicas.) And if you keep the journals around, you have a complete history of the application's state.
EDIT: See also http://www.advogato.org/article/398.html http://www.advogato.org/article/398.html
- dk 19y agoI probably should have pointed out the implications for disk I/O. In many cases, the serialization of the command in the journal has a smaller footprint on disk than the data that's modified. Consider an extreme case where a small HTTP POST touches dozens or hundreds of records. And appending data to a log is essentially an ideal disk access pattern. Of course this can't be said for the snapshots, but you can offload that task to a replica server.