8 ms·
That was very informative, but I wondered if there is an easy-to-read reference somewhere about WHY git works the way it does? I am old enough to have used SCC
by deepspace 3y ago
That was very informative, but I wondered if there is an easy-to-read reference somewhere about WHY git works the way it does?
I am old enough to have used SCCS, RCS and CVS extensively. Each had their faults, but Git is the only VCS I have used where dealing with merge conflicts is unintuitive enough that I sometimes end up with the repository in an unusable state. I am sure I am doing something wrong, but I would like to understand why.
The VCS that maps closes to the way my brain works is ClearCase. You essentially have a versioned file system, and you can set up a view to present any previous state of that file system. Of course, administration is a nightmare, it is not distributed, it is expensive, yada, yada. But when using it I always felt I knew exactly what was going on under the covers, which is not the case with Git at all.
- zadokshi 3y agoI used git for years now. I’m comfortable with all of the basic functions. It is my experience that I have had several times where merging/branching has caused a repo to “break” It doesn’t happen often, but when it does it’s super frustrating and this I avoid the merging workflows where possible. I’m sure if I educate myself a bit more about git and pay more careful attention to the details, I wouldn’t occasionally have this problem, but Git really shouldn’t be like this.
- jayshua 3y agoWhat does a broken repo look like? What does broken mean here?
- appplication 3y agoGit could be a lot better in a lot of ways, particularly from a developer experience perspective. I’m a little surprised we haven’t seen a meaningful successor. One example is how git will deceive you and tell you you’re up to date with your remote (e.g. origin/main). What it means is that your local branch is up to date with its local concept of remote, and makes no statement guarantees about the actual state of the remote. Which is really a nonsense concept that does not need to exist. Similarly the whole concept of needing to specify “origin” at all is a bit bonkers and does no favors. Why is it that I can pull from a remote branch, commit some changes, run ‘git push’ and git has no idea what branch I want to push to. Another example: if main is a protected branch, don’t let me accidentally commit to it locally. I could keep going with the examples but I won’t. And yeah you can forgive all this by quibbling that git was written in a time when internet access was not ubiquitous, and of course all these decisions make sense because x, y, or z advanced edge case for advanced users only, and I’m a shitty engineer because all of this complexity secretly makes my life better and I’m just too simpleminded to appreciate it. Really though, if you rewrote git from a principles first approach (with developer experience being one of those principles), it certainly would not look like how it looks today. There is too much complexity, too many ways to do things, and too many bad decisions around defaults. Treat it like a proper distributed system, perhaps even backed by a real database. It’s not special because the data is code. The fact that it’s treated as such is the reason it feels so weird.
- TylerE 3y agoBlame the Linus Torvalds personality cult. Only way that terrible UX got traction in the first place. If anyone OTHER than Linus had written it, it would have been mocked, and deservedly so.
- appplication 3y agoYeah, I originally had some Linus snark in there but deleted it. My above comment is being rapidly downvoted though, which is in some sense validating.
- Yodel0914 3y agoThat's a pretty big claim. Git solved real problems at the time in a novel and extremely useful way. Coming from SVN (or god forbid VSS), the power and flexibility of git more than made up for its difficulty to grasp. Why it "won" over alternatives with cleaner UX (eg Hg) is a different question, and I think has a lot to do with github.
- spaniard89277 3y agoGit may solve real problems but honestly, it's a PITA to use. Just for basic tasks needs quite a lot time sink to understand what's going on and what do you need to do. It gets in the way of the workflow IMO.
- TylerE 3y agoMercurial came out essentially simultaneously, solved the same problems, with a much more coherent UX.
- Yodel0914 3y agoEarly on, I was definitely a proponent for hg over git. I assumed it would at the very least remain a peer, given how much more sane the UX is. I was wrong though, and I don't think it's just because of github. Named branches being baked in turned out to actually be a real pain in mercurial. Sure, you can use bookmarks (and we did) but starts to quickly feel like you're swimming against the current.
- Yodel0914 3y agoI've been using git for I guess 15 years, on a variety of server platforms (github, atlassian, MS devops, plain ssh) and branching models, with teams of various sizes, and I've never had a repo "break". Branching and merging is what git is amazing at. The worse thing I've had happen is that someone was lazy when merging and just did a "take mine" on 100s of files, thinking that they'd do a "proper" merge later. Of course git doesn't work like that and their lazy merge had to unwound very carefully.
- nonethewiser 3y ago> It is my experience that I have had several times where merging/branching has caused a repo to “break” What do you mean? Like you resolve the conflict correctly and git stops working? Stops working how? That sounds very strange, and anything else just seems like PEBKAC.
- eru 3y agoI wouldn't go so far as the term PEBKAC. It's true that git does not 'break' from a merge; but merge conflicts (and rebase conflicts) can still be frustrating to resolve for ordinary users. And after things get too frustrating, they often do really random stuff that then might accidentally break the repo for real, or at least get it into a state where they don't have enough knowledge to recover. Git's underlying data model is fine, but the user interface can be quite lacking. For example, 'git checkout' is a mess of almost unrelated functionality thrown together under one command. I think the idea behind 'git switch' is a good one, and git could benefit from a complete overhaul of its user interface. Well, at least if you ignore the switching costs. Old farts like you and me have gotten used to the quirks of the bad old interface, and learned all the barely coherent options to 'git reset'.
- aliher1911 3y agoFor me git reflog was always a lifesaver when a complicated rebase or merge fails. You can always checkout previous SHA and then just return your HEAD to pre screw-up state. In case of remote repo that already ended up "broken" you can still force push to those branches, but others will have to discard theirs. Otherwise you never force push to anything shared.
- mabbo 3y agoIt's definitely not a super intuitive model, but it is a powerful one. As I see it, git is a distributed database of snapshots of a directory of files. Every commit is a new snapshot. It's designed the way it is to achieve that goal while minimizing the space used, yet keeping an entire history of the project. It also has tools like git-diff to better compare those snapshots. But distributed matters most because it was literally designed for managing the Linux kernel development- a large open source project with thousands of contributors working concurrently. Merge conflicts are hard in this model because you're trying to put two snapshots together, not just two diffs.
- eru 3y ago> It's designed the way it is to achieve that goal while minimizing the space used, yet keeping an entire history of the project Most of git's design was done before they thought of 'minimizing the space used'. Originally, they didn't even do deltas and just stored complete snapshots for every object. > Merge conflicts are hard in this model because you're trying to put two snapshots together, not just two diffs. I don't think that's the case at all. Thanks to three-way-merges git has just as much access to the diffs as any other version control system when merging. It's just that in git diffs are a derived data structure, not a source of truth, but that doesn't make a difference.
- hooper 3y agoOne of the ideas behind Pijul is that implicit vs explicit diffs does make an important difference sometimes: https://pijul.org/manual/why_pijul.html https://pijul.org/manual/why_pijul.html
- ta8645 3y ago> ClearCase [...] it is not distributed [...] That's the key reason Git can't work the same way, and what makes it so powerful, and sometimes hard to grok. Git is based on a Directed Acyclic Graph (DAG) of committed changes. The DAG is shared by everyone. Each commit is immutable, and only extends the DAG with a new node. You do not alter any of the DAG that other people already possess, you are creating a NEW chain of nodes connected to the DAG. And that's it, that's how it allows anyone to make changes, at any time, on any number of systems. Because all operations are strictly additive to the DAG. Every time you commit changes, you add another immutable node onto the DAG. And at the very same moment, on another unrelated server, someone else is adding unrelated immutable nodes as well. Later, you may obtain some of those remote changes, they can be fetched into your local repository. You can fetch them without merging them. You will just have a copy of how someone else extended the DAG. But there will be no connection between the changes you made locally, and the changes the other person made in their repository. When you merge, you're simply updating the DAG with a new node, it will point at two, previously disconnected chains of commits. Both sides of the merge represent immutable changes that extended the DAG starting from a single commit, the branching point. And that's what you see in merge conflicts. The first block shows what a section of code looks like in the local branch. And the second block (after the equal signs) shows what that same section of code looks like in the remote branch you're merging. <<<<<<< local a X c ======== a Z c >>>>>>>> remote That's it.
- seesaw 3y ago+1 for clearcase. I enjoyed the time I used clearcase, and never found another one that was as pleasant to use as an end user.
- Borg3 3y agoUgh. Ive been both using ClearCase and also were Admin of it. While as user it was pretty fine to use, as Admin I saw how complicated and fragile the whole thing was. They distributed model is absolutly terrible. You need powerfull box as a VOB server and probably be equal or even more powerfull box if you want to use dynamic views. We ended up using snapshot views. I would choose GIT every single time over ClearCase.
- deleted 3y ago[deleted]
- jamie_ca 3y agoA lot of its design decisions are based around the data storage model, and tooling built to operate on those data structures. I recall a good write-up from a decade gone now, but no dice googling for it. The short version is probably: - Everything is a blob, a text file named after the SHA1 of its content. - Files are just themselves. - Directories list <entry sha1> <name> for their entries (file or subdirectory). - Commits list a Directory (the project root), and some metadata about the commit like the author, commit message, and the parent(s) of the commit. - A branch then is just an end-user-named reference to a commit's hash. Everything flows from that - SHAs are reused, if you're doing a diff and two directory entries have the same sha referencing a file, there's no change. Switching a branch is modifying the special HEAD branch content, and recursively walking it to rehydrate the filesystem (and comparing to the previous checkout to optimize, skipping whole directories that don't change).
- yencabulator 3y agoMy attempt at explaining that is at https://eagain.net/articles/git-for-computer-scientists/ https://eagain.net/articles/git-for-computer-scientists/
- wayfinder 3y agoI think Git is conceptually simple... each commit is just conceptually a diff from a previous commit -- possibly from two commits. Branches and tags are just pointers to a commit. A merge is when you join two commits and thus may result in one merge conflict, while a rebase is resetting your branch to some other target commit and then /re-applying/ _each_ of your commits on top, which may generate more than 1 merge conflict (due to each _one_ of your commits effectively being re-committed all over again) but hides the fact that you applying old changes to a newer base. But the tooling around Git is not great. I think showing merge conflicts properly so you don't mess them up is also tooling issue more than anything. I'm going to plug this app called SmartGit which I have no relationship with that I don't think a lot of people have tried but it's awesome. Doesn't obscure anything about Git and shows merge conflicts well IMO. Costs $ though.
- eru 3y ago> I think Git is conceptually simple... each commit is just conceptually a diff from a previous commit -- possibly from two commits. Branches and tags are just pointers to a commit. That's backwards. Git commits are complete snapshots of the state of your repository at the time. You can compute diffs between arbitrary commits, be that between child and parent commits, or completely unrelated commits. Have a look at https://stackoverflow.com/questions/4129049/why-is-a-3-way-merge-advantageous-over-a-2-way-merge https://stackoverflow.com/questions/4129049/why-is-a-3-way-m...
- wayfinder 3y agoOh I know. I do that all that time. (And in SmartFGit, you can select any two commits.) But I said conceptually and that’s how I see them and it works for me.
- OskarS 3y agoConceptually, though, that's the wrong model. Pijul works like that, but git does not: commits are not diffs, they are full snapshots of the repository, including history. If you cherry-pick a commit from one branch to another, that's an entire new commit, it's not just "this patch moved from one branch to another". If your idea of a commit is "it's a diff to the previous commit", you'll get into trouble.