8 ms·
Linus Torvalds: “GitHub creates useless garbage merges”
- darig 5y agoProgrammers create useless garbage merges
- warpech 5y agoMerges done from a GitHub Pull Request (PR) contain a reference to that PR. The reference is not even a full URL, but just a number (e.g. "#42"). You need to use the GitHub web UI to turn this number into a URL (or deduct it if you know the repo URL). If you have that, then you can go to the Pull Request web page when you can read the whole story about the merge. What Linus rightfully wants is the whole story right in the commit message of the merge. I think GitHub purposely makes condensed commit messages for merges as a way to hook users to their web UI.
- anonymoushn 5y agoThese issue numbers are non-portable even within github! Two repositories may have the same commits in them but different sets of issues.
- est31 5y agoMy favourite is when someone collaborates on a PR and they create PRs for the PR's branch in the forked repo. That repo has its own numbers so you'll end up with merge commits referencing PR #4 in the main repo.
- tadfisher 5y agoBecause forks are GUI sugar for refs in the main repo. This is also why Actions have convoluted security implications.
- est31 5y agoNo I didn't mean the "you can replace the fork's URL with the main repo one and make the commit seem part of the repo". That's a weird github oddity to keep in mind, but not the issue I meant. I meant when that PR including its merge commits gets merged. Then those commmits become proper members of the main repo.
- jareklupinski 5y agothis came very close to exposing private repo details to the wrong users for me recently currently trying to find an exploit using this 'feature'
- throwaway984393 5y agoI don't work for GitHub, but I would eat my shoes if GitHub made any conscious choice at all in their design regarding keeping people on the web UI. It sounds almost exactly like a decision made to keep some sort of backwards compatibility without having to do any extra engineering, stray far from vanilla Git, or make the user experience worse for regular users.
- est31 5y agoThe issue with full URLs is renames of user or project names. I don't trust the inbuilt redirection support of Github this much. Plus, tools that don't shorten full URLs would get extremely long lines in their commit messages. I think the Github issue numbers translate 1:1 to Gitlab if you use the Github->Gitlab exporter. If you used full URLs, it would either have to keep the URLs, or change commit hashes, neither a really nice option. Also, not that chromium did something different. They reference tons of internal bug IDs, change IDs, and this commit msg also has an URL with the non-standard domain go (no TLD): https://github.com/chromium/chromium/commit/8b4b1831e333a43d2da2d642d1b05256b1a921ed https://github.com/chromium/chromium/commit/8b4b1831e333a43d...
- anonydsfsfs 5y ago> Also, not that chromium did something different. They reference tons of internal bug IDs, change IDs, and this commit msg also has an URL with the non-standard domain go (no TLD): The kernel does that too. Those key-value pairs are called "trailers", and are semi-standardized at https://git.wiki.kernel.org/index.php/CommitMessageConventions https://git.wiki.kernel.org/index.php/CommitMessageConventio... The "git interpret-trailers" command exists to work with this kind of data.
- alerighi 5y agoThese has two problems, first it will impose you to use the official or a compatible GitHub interface, not the command line or whatever git GUI interface you want. Even with that, it is impossible to use offline, while your git history is local and can be consulted whenever you want without a connection, and you can make whatever changes you want and push them when you have internet access. The second problem is that it's not portable, it's something internal of GitHub, that is fine till you use GitHub, when you need to migrate to another server, or host your own git repository, for reasons (GitHub changes plans, GitHub closes down, GitHub no longer wants to host your project, etc) what do you do?
- dathinab 5y agoHaving a number short-form is a de-facto standard across most "version control+issue tracker" services. For nearly all cases it's good enough, projects like the Linux kernel are in the end an exception wrt. their requirements and workflow. Most project are just hosted in one place so the shorthand is always clear, and has the benefit of not being "broken" if a project is renamed/moved. Furthermore GitHub always allows the user to specify the commit message, it just prefills some hint. But it can't really know your style of using git so it can't do much more then just prefilling the hints. But you can always write them the way you want to filling in all the details needed. Through IMHO the problem are non-ff git merges in general, not only do they motivate bad commit messages but also does "not having merge conflicts" in no way mean there is no problem. I had more then one time where a "no conflict" git merge introduced bugs. Projects like Linux are a bit special due to the size, number of contributers and sizes of contribution. But for most projects I can just strongly recommend to only allow FF merges and locally rebase branches on master if it doesn't work. Preferable with users using interactive rebases to create nice commit histories, alternatively squashing works to but git is just _bad_ at squashing and so are tools based ontop of it (e.g. GitHub).
- VenTatsu 5y agoIf the data is not stored in the git repo itself there is no workable solution. Any offered standard is just a bandaid on the problem. If a project is renamed or relocated from one name to another it might (and should) preserve issue and merge related information and discussion, but any full URLs will have changed, making full URLs fail in specific cases. Using any form of shortened sequential identifiers will fail when referenced in commits merged in from forks of the project that have their own conflicting identifiers. To solve this in any real way there needs to be a way that when a project is forked it carries information about issues and merges with it, and when commits from that fork are merged back into the main project, or merges from a fork of a fork are merged all the way back to the original project, all metadata about those merges are carried to the original project. Linus' argument is that the correct place to store that is in the merge commit's message. That it is the responsibility of the person doing the merge to provide that information, and that GitHub does a bad job of providing tools to do that.
- 5y ago
- 908B64B197 5y agoClickbait tittle. It's not the content of what's merged that he's complaining about, it's the default formatting from the GitHub UI (no commit message for instance).
- almost_usual 5y agoHow is it clickbait, it’s what he said. > That's another of those things that I really don't want to see - github creates absolutely useless garbage merges, and you should never ever use the github interfaces to merge anything.
- 908B64B197 5y agoThe tittle sort of implies that "people contributing from GitHub try to merge garbage into the kernel tree".
- deleted 5y ago[deleted]
- Jare 5y agoHe said many things. Picking the one that, out of context, sounds like something irritating and controversial, is the very definition of clickbait.
- chrisseaton 5y agoIs there a way to sign your merge commits on GitHub? How do they get trusted?
- Felk 5y agoYes. Here's an excerpt from their documentation on <https://docs.github.com/en/github/authenticating-to-github/managing-commit-signature-verification/about-commit-signature-verification https://docs.github.com/en/github/authenticating-to-github/m...>: > GitHub will automatically use GPG to sign commits you make using the GitHub web interface
- chrisseaton 5y agoSays it signs the commit with its own key. I guess you have to trust GitHub.
- Felk 5y agoWell, yes. The question was whether you can sign _on GitHub_, so your private key has to be available to GitHub. You can always sign locally if you don't trust GitHub.
- drexlspivey 5y agoWhat else would they be signing with? They don’t have your key obviously
- chrisseaton 5y agoWell that was my point - I wonder why we haven't set up a system that lets me sign the merge commit. Otherwise it's a commit purported to be authored by me but when you look it's actually signed by someone else.
- ruuda 5y agoIt’s even worse, if somebody rebase-merges a pull request that you authored (thereby creating a new commit that you did not author), GitHub will show you as the author (without a separate committer, like it normally does when author and committer differ), and put “verified” next to it, which usually means that they verified that it was signed by your GPG key, but in this case, it means that the commit was created by GitHub. https://twitter.com/vmulps/status/1386717970458677250 https://twitter.com/vmulps/status/1386717970458677250
- bborud 5y agoSounds like it would be easy for Github to fix? No? Yes?
- sp332 5y agoMaybe, but they've never cared about supporting the Linux kernel team's workflow.
- deleted 5y ago[deleted]
- bborud 5y agoI can't say I know much about the Linux kernel team's workflow, but sensible merge commits would be useful.
- joelbluminator 5y agoDoes this bother a lot of people? Has anyone ever opened an issue on this? If not I say why bother...Linus is hard to please
- nickelpro 5y agoGithub already offers several way to incorporate PRs, including squashing and rebasing. It also offers repo admins the ability to _forbid_ maintainers from using the default Github merge. I think that's plenty of work dedicated to fixing this issue.
- bifrost 5y agoOne of the few times I agree with Linus.
- asveikau 5y agoI think Linus generally has well reasoned technical opinions. I find it a little unusual that you decided to communicate that it's rare that you agree with him. I think it's much more common to say his communication style has problems. I think he has really good intuition on technical stuff. This is not to say Linux or Linus are perfect, but you can usually pick out certain of his opinions and find them to be pretty grounded.
- bifrost 5y agoLinus is why I generally don't use Linux unless I absolutely have to.
- asveikau 5y agoI guess looking at your user info you like BSD? I like *BSD better than Linux. But it's not due to personal feelings about Linus, and I respect his talents and good eye for a lot of technical things.
- bifrost 5y agoYep, been using FreeBSD on the Desktop since 1996. I've used Linux for plenty of things, but I tend not to unless I need something thats written specifically for Linux. There's very little I need Linux itself for. My issues with Linus were a few design decisions, how stuff gets into the kernel, how the distributions work, how things are poorly tested/etc. Sadly now that FreeBSD is using ZoL I'm encountering a bunch of problems related to that :(
- Sharlin 5y agoI mean, he’s usually right. Even if historically the way he has expressed himself has not always been the most agreeable.
- deleted 5y ago[deleted]
- systemvoltage 5y agoBehind the rage, there is good wisdom in Linus’s rants.
- bigcorp-slave 5y agoHot take: why do we still care? Linus says something negative about something, big shocker. This is just the tech crowd equivalent of a tabloid at the supermarket checkout isle, Tom Cruise is secretly a ferret-lover, read all about Bob’s torrid affair with Mallory!
- lopatin 5y agoBecause it's fun
- chrisseaton 5y agoThat's the biggest thing to learn from bullies like Linus - he hates everything and thinks everyone's an idiot, so where does he go when he actually needs to express something strongly? He can't go anywhere. It's not an effective way to communicate.
- temp8964 5y agoExactly. Linux became successful only after Linus was re-educated.
- chrisseaton 5y agoYou're being snarky, but that's not even what I said anyway. You can still be successful even if your communication is limited.
- nixpulvis 5y agoPlease fuck right off with this.
- Dylan16807 5y ago> That's the biggest thing to learn from bullies like Linus - he hates everything and thinks everyone's an idiot, so where does he go when he actually needs to express something strongly? When he has a list of complaints, that is his strong reaction. And why would he need to go stronger?
- deleted 5y ago[deleted]
- ameixaseca 5y agoFor anyone that haven't tried this yet, please have a look at 3-way merging. A number of tools support it, including simple ones like `meld`. I've been using it for a while and it has been a huge help on complicated merges.
- kklisura 5y agoI had to google what Torvalds currently do: ..."What is your job?" Torvalds replied, "I read and write a lot of email. My job really is, in the end, is to say 'no.' Somebody has to say 'no' to [this patch or that pull request]. And because developers know that if they do something that I'll say 'no' to, they do a better job of writing the code." [1] Do we have any idea what will happen to Linux once Torvalds is not there? Is there any example in history of OSS development as to what happens when BDFL [2] is not there anymore? [1] https://www.zdnet.com/article/linus-torvalds-im-not-a-programmer-anymore/ https://www.zdnet.com/article/linus-torvalds-im-not-a-progra... [2] https://en.wikipedia.org/wiki/Benevolent_dictator_for_life https://en.wikipedia.org/wiki/Benevolent_dictator_for_life
- Valmar 5y agoAt best, we can just hope that Greg has the same resilience as Linus ~ and that Greg's chosen successor also has the same grit
- thih9 5y agoFor those unfamiliar, Greg Kroah-Hartman is the maintainer of the Linux stable branch since 2013: https://en.wikipedia.org/wiki/Greg_Kroah-Hartman https://en.wikipedia.org/wiki/Greg_Kroah-Hartman
- garciasn 5y agoThere is no publicly acknowledged succession plan but Linus has previously commented it may take a few months for the community to work out how to move forward but he seems confident they will. It’s just that no one in their right mind would want to do Linus’ job. It’s miserable and stressful; he only does it because he’s been doing it for 30 years and it’s his baby. The outcome will be interesting to be certain.
- moron4hire 5y agoThe problem is that Git is used for two things: the Linux kernel and anything else. The Linux kernel has this really weird system of software development where nobody knows anyone else and they want to share all their changes over email. Torvalds, for reasons I can't figure out, only cares about that one, niche use case.
- kristjansson 5y agoBecause that ‘niche’ use case is what he wrote git for, and it just happened to be useful to the rest of us?
- moron4hire 5y agowhoosh Edit: but it really is a niche use case. Count the number of developers working, the number of lines written, for projects that are not the special snowflake of the Linux kernel. Line of business apps. Consultoware. Fuggin' code bootcamp students selfishly using other people's open source projects for their grad advancement. All this code that nobody, maybe not even the people writing it, cares about. Git is the wrong tool for those jobs, but because Torvalds wrote it, it must be good.
- kristjansson 5y agoWhoops, missed the snark the first time though.
- lifeisstillgood 5y agoCurious that this comes on front page along side the whole Atlassian / Jira debate. One of the major reasons Github does not / cannot do commit messages the way Linus likes is that the reasons why odd is changed are in and supposed to be in the ticket tracker (where non-coders / committers can view / comment / approve.) This is not a technical fix that github can make - this is a fundamental split about where the power lies - most companies do not have the final say, the source of truth, in their codebase and consequently their repos. Linus is on the side of history here. But for those of us coders living in ticket trackers for our daily job, it's a serious culture shock. For non- coders it's an existential threat. On a separate note, I think there is a space for using hashes to compress tickets and all the ancillary evidence and sign offs, into one artifact, and then reference that in the repo commit. But I just can't see it in my head
- dathinab 5y ago> Linus is on the side of history here. Through not seldom the source of truth is also hidden somewhere in the Linux mailing list. Similar a (insane) high amount of projects on GitHub and in companies will hardly ever use the history besides a last few commits. Sure git blame is a thing, and so is bisect but how often is it used on projects besides large open-source projects?
- superbatfish 5y agoFrom the headline, I thought this would be about github’s squash merges, which represent an extreme anti-pattern to any sane developer.
- tkzed49 5y agocan you explain further? In my experience other people's half-baked commit messages are just noise while the pull request is more carefully written, so the squash merge is more intelligible. Additionally, the PR is typically used as the "atomic" code change rather than the individual commits.
- totony 5y agoEven if they have shitty commits, it barely matter if its hidden behind a merge commit. On the off chance that squashed commits wouldve been useful, I much rather have it than not. In fact I once did a PR moving around some files which is visible with multiple commits (move file commit, edit file commit) which was replaced by a sqiash commit that deletes files and adds one with modified content.
- superbatfish 5y agoI agree: Many people have half-baked commit messages. And many erroneously assume that a PR is always a single atomic code change, rather than a sequence of pseudo-atomic changes. As an industry, we should be nudging people away from those bad habits. Instead, the "squash merge" feature nudges them in the opposite direction. It disincentivizes careful commit commit messages and encourages general disregard for the commit log. Collaboration requires good communication. For programmers, that means: - writing clear code - commenting your code - documenting your code - writing a clear summary in the PR description - organizing your changes into cohesive, (mostly) self-contained, reviewable units -- expressed as independent commits in the version history On that last point: The commit log is something worth curating, tending to it with the same care you would use for the other points. Yes, PRs are often small and self-contained enough that it shouldn't be subdivided into multiple commits. But just as often, it's clearer to use multiple commits which can be reviewed in sequence, rather than in one big chunk. In the latter case, it's incumbent on the PR author to present a logical, clean commit history. If you wrote four commits but two of them are inseparable, then squash those two before you submit the PR. Commit messages like "fix typo from last change" are completely unacceptable in a PR. That sort of thing is fine in your own repos, outside of any review process. But when sharing code with others? No way -- that's inconsiderate with others' time, and pollutes a valuable data stream (the commit log). But hey, if there's a good chance that all your commits will be squashed anyway, then why bother curating your commit log at all? Instead of letting bad commit messages slide and simply squashing them, we should gently insist that PR authors fix their commits on their end. Their next PR will be better for it. In the past, I've been occasionally annoyed when my carefully crafted commit log was simply squashed by a PR. I had intended the commits to be reviewed one-by-one. And if a bug is discovered in one commit later on, it would be clear which lines ought to be reverted, and which should stay unchanged. All of that care gets thrown out by a squash. Grr...
- Shadonototro 5y agoHe is not wrong
- Thristle 5y agoThat's just a preference issue. If you just pick rebase instead of merge from the UI the commits will be merged without that "meta commit". No need to get all emotional
- nanis 5y agoWhen I train developers on `git` in combination with GitHub, I emphasize that `git` the tool was written for a specific goal. Linus wrote it to replace BitKeeper for the express purpose of collaborating on the maintenance & development of the Linux kernel. The biggest problem is when developers who do not really understand version control at all (there are a lot of them -- they tend to think of it as a "save" mechanism) are allowed to use all the features of GitHub and can only do their work with buttons. For example, it is common for some developers to think that there are only two options when dealing with merge conflicts: "Accept theirs" or "accept ours" when, in reality, most real merge conflicts require a synthesis. Now, add to this the fact that such developers barely write commit message (seen too many projects with only "merge X and Y", "update", "fix bug", "implement feature" etc). The situation gets worse. Soon, you end up with a project where the only tool for locating any issue becomes `git bisect`.
- dvdgsng 5y agoYes, some devs just can't merge. Even worse, IntelliJ & co have this "magic merge" button, that resolves conflicts automatically, but the results are usually a mess, esp when merging complex features. But because lazy devs are lazy, they just smash that button, commit, push without testing locally and then move on to the next topic. Naturally, we've had countless CICD failures which were attributed to "merging in git is sooo broken" ...
- Mic92 5y agoI still prefer github's UI and PR style over those ugly kernel mailinglists, where there is no merge status and no good CI integration and no good search unless you force yourself into using some complex mail setup.
- im_dario 5y agoNow, I wish there were a guide to creating proper pull requests. From what I read, I got some things, but not a straightforward or actionable path. The only clear thing is the unmistakable feeling of not doing PRs properly.
- seba_dos1 5y agoThere's https://www.kernel.org/doc/html/latest/maintainer/pull-requests.html https://www.kernel.org/doc/html/latest/maintainer/pull-reque... Note that you generally don't send pull requests when contributing to Linux unless your changeset is unusually big, so pull requests are usually only used by maintainers to forward their stuff to other maintainers. Regular contributors submit patches instead: https://www.kernel.org/doc/html/latest/process/submitting-patches.html https://www.kernel.org/doc/html/latest/process/submitting-pa...