4 ms·
> Question: How does Mercurial deal with garbage collection? After all, the desire for garbage collection is by far not unique to Git -- any version control sys
by rbehrends 9y ago
> Question: How does Mercurial deal with garbage collection? After all, the desire for garbage collection is by far not unique to Git -- any version control system that has the equivalent of `commit --amend` and rebase should provide it.
Garbage collecion is an issue that is 100% unique to Git. No other VCS even thinks about throwing user data in the repository away without the user explicitly telling it to. Once you have a user telling you to throw the data away, it can do that. There is no need for GC; this is purely an artifact of Git's implementation. I'm honestly not sure why you think you'd even need a GC for `hg commit --amend` or `hg rebase` (or similar operations in other VCSes).
> As for not needing the extension, it seems to me that having "dangling" commits would be very difficult to use without some decent visualization of the dangling commits, such as what the blog post shows with `hg show`.
I'm not sure where you get the idea. This feature is, after all, not unique to Mercurial. It's Git that has the oddball semantics that no other VCS on earth has. Mercurial has been able to graphically show the graph for ages and the ability to just list open heads, too (`hg heads`). If you look at the code, the implementation of the show command is largely just a templated graphlog of a particular revset. For example, `hg wip` [1] (for "work in progress") has been doing something similar just using revsets and templates from core Mercurial.
> The point of the purely functional data structures isn't to achieve atomicity (although potentially being a bit more robust to power loss etc. is certainly a nice side effect), it's a way of thinking about version control. I've never heard functional programmers use atomicity as the main argument for immutable data structures, either...
This is because functional languages do not have to worry about their state being destroyed by the user hitting Control-C or a power outage. This will simply terminate the program, whereas for Git it will interrupt a transaction in progress.
> The whole point of Git's design is that it chose a robust and crystal clear way of thinking about distributed versioning as its underlying model of what version control is, and then simply provided tools for manipulating that DAG.
This is what other version control systems do, too, without relying on purely functional data structures. The fact that the data structures are purely functional is, after all, not a property that is visible to the user other than through the side effects of garbage collection.
[1] http://jordi.inversethought.com/blog/customising-mercurial-like-a-pro/ http://jordi.inversethought.com/blog/customising-mercurial-l...
- nhaehnle 9y ago> I'm honestly not sure why you think you'd even need a GC for `hg commit --amend` or `hg rebase` (or similar operations in other VCSes). Maybe you don't call it GC, but doing a rebase in Git leaves the old, pre-rebase version around. That is a feature: over the years, it has happened to me more than once that I'd missed something when resolving complex conflicts during the rebase. Being able to refer back to the state from before the rebase was very helpful in these cases. If Mercurial throws the old version away unconditionally, that would suck very much indeed. If Mercurial keeps the old version, then perhaps at some point in the future I'd really rather have all that old data removed as a simple matter of saving disk space. Surely nobody wants to do that manually? Hence: you either have a system that makes it much easier than Git to lose data, or you need garbage collection. I don't know what Mercurial does, but somehow, the fact that this dilemma isn't obvious to you -- somebody who clearly seems to know a lot about Mercurial -- doesn't instill a lot of confidence in it.
- rbehrends 9y ago> If Mercurial throws the old version away unconditionally, that would suck very much indeed. Which is why it isn't done. In core Mercurial, the old revisions are stored in a backup bundle in a separate backup directory. Note that bundles can transparently be used as read-only repositories, so you can view their logs as though they were still part of the parent repo, diff against them, pull from them, etc. With the evolve extension, those revisions will simply be marked as obsolete, with obsolescence markers showing which revisions were replaced by which. The commits will be hidden, but are still part of the repository. If you ever want to get rid of the old revisions, you'd have to use (say) `hg strip -r 'exctinct()'`, which would store them as bundles as described above, or clone the repository and delete the old repository. Plus, there are public, draft, and secret changesets. Public changesets are immutable and cannot be changed without user override. Bazaar rebase will simply hide the old revisions; you can recover them with `bzr heads --all`. To permanently delete the revisions, you have to clone the repository and delete the old version (and all backups). And, of course, there's rarely a reason to use rebase in Bazaar. > Hence: you either have a system that makes it much easier than Git to lose data, or you need garbage collection. As I described above, neither. In every case, you need to go through several steps each requiring the user to affirmatively express their desire to delete data. And as disk space is really cheap these days, hardly anyone ever actually deletes the data in practice, as there's no point to it. > I don't know what Mercurial does, but somehow, the fact that this dilemma isn't obvious to you -- somebody who clearly seems to know a lot about Mercurial -- doesn't instill a lot of confidence in it. I think your dilemma is largely an imaginary one, fretting over a resource (disk space) that is too plentiful to require micromanagement. Keep in mind that most of the data in your repository will come from other people; there's only so much source code or text that a single person can write in a day. If you're generating massively large binary assets, a DVCS is probably the wrong tool, anyway, because of scaling concerns. This inherently limits the amount of "wasted" data that you can have in a repository to a percentage of the repository size.