4 ms·
Yes and no. The most important part of replication is to have an algorithm that's capable of merging two document histories that are not identical. This core b
by davisp 15y ago
Yes and no.
The most important part of replication is to have an algorithm that's capable of merging two document histories that are not identical. This core bit of CouchDB is quite important and often overlooked in terms of its replication scheme. Randal Leeds is currently hacking this into PouchDB and I'm quite excited to see the algorithm in a non-functional language so that more people can see its simplicity without worrying about learning a new programming paradigm.
On the other hand the atomic update of the two indexes is quite important for efficient replication. Without the atomic update of a by-update-sequence index it would require a full table scan for each replication. Its definitely a necessary optimization, but not a sufficient optimization. The per-doc revision history merging is the special sauce that makes things work.
In unrelated news:
http://twitter.com/vmische/status/100956289387077633 http://twitter.com/vmische/status/100956289387077633
- mark242 15y agoI would go so far as to "bypass" the spec of CouchDB -- at least for Pouch -- by doing a sort of auto-compaction, clearing out the database of the previous revision at storage time, and only storing the latest revision. This doesn't stop you from doing an effective merge, it simply stops the original notion of Couch having multiple branches of the same document. Again, in the browser, to sacrifice indexedDB space, I think that's an okay step to take. The real win will be when a WebWorker is able to run through a view and automatically add the results of that to a separate dbspace. It looks like viewQuery is the beginnings of that.
- tilgovi 15y agoI'm the one working on this right now and this is exactly what I'm doing. However, old document bodies are not the only way to store history. Documents have a revision history, possibly 'stemmed' to a maximum height, which is an array that traces a path from a certain depth back toward the root. All the leafs are conflicts. This is the algorithm davisp is talking about. While I'm not going to leave the document bodies of superseded revisions accessible (by overwriting them in IndexedDB), I still need to store these history paths so that merging a new revision (either via interactive edit or replication) can properly generate a conflict when applicable.
- tilgovi 15y agoAlso, the version on master does what you suggest and makes no attempt to accommodate replication-based merging, thereby dealing implicitly with merging revision histories stemmed to a height of 2. "Real" CouchDB replication is what we aim for, though, and so this algorithm must be written.