6 ms·
GitHub’s Large File Storage is no panacea for Open Source
- paulddraper 11y ago> Case in point: if a very popular Github repository (such as the one for the Linux kernel) decided to start using LFS for some of their files, they would instantly alienate all of their users. They would no longer be able to properly fork the project, or even clone it to get its binary files stored via LFS. Nobody would be able to send a pull request to Linus as a result without considerable effort. Odd example. Linux doesn't use GitHub pull requests.
- Deinumite 11y agoYeah I am assuming OP doesn't realize that patches are sent to the Linux kernel through emails.
- megastep 11y agoI do, actually. I just noticed that there was a large Kernel git repo hosted on Github, and figured it would make a good example of a large, popular repository. I'm not surprised Linus is not actually using Github, so my bad for not stressing that this was more hypothetical than meant as a statement of fact.
- paulddraper 11y agoThere are literally thousands of uber-popular projects that are deeply integrated with GitHub. I recommend choosing one of those for an example. Line NodeJS or something.
- Jasper_ 11y agoThey still use pull requests, though, just in the form of email-based ones. https://git-scm.com/docs/git-request-pull https://git-scm.com/docs/git-request-pull
- jumpwah 11y agoExcept that's not a "pull request".
- Vendan 11y agoExcept it is. Just cause github does pull requests differently doesn't mean you can't do pull requests in pure git. Remember, git came before github.
- jumpwah 11y ago> Remember, git came before github. That's exactly my point. Using a "request-pull" instead of a "pull-request" will probably mean you can do this git-lfs thing with it. My understanding was that it was not working with "forked" repositories, which you need to have to make a "pull request". To make a "request pull" (or to just send a patch file), you don't need to "fork". ;)
- paulddraper 11y agoSure. With git-fu and enough emails, you can replicate anything done in GitHub. I'm just saying it's a very odd choice for an example of GitHub screwing over workflows.
- eridius 11y agoI would imagine that there's a number of people out there who still fork the GitHub mirror of Linux and use that to build their pull requests (which are then submitted via email rather than over GitHub, but which would presumably have the exact same issues). In fact, there are currently 10,532 forks according to GitHub.
- alkonaut 11y agoThe most interesting takeaway for me was that Microsoft seems to privide the only(?) free git hosting that includes LFS? Does anyone know if their repos supports forking in combination with LFS too?
- bsimpson 11y agoThe post seems hyperbolic. I'd love to hear GitHub's rebuttal.
- megastep 11y agoI am the author and yes this was very much unapologetically hyperbolic. At least it got the conversation started.
- ansiton 11y agoI don't think you needed paragraphs like "My guess is that some high-level greedy marketing dickwad, completely unaware of the asinine implications of his brilliant idea, signed off on this dumb-as-a-bag-of-rocks pricing model. He then directed the grunts to somehow implement his grand vision on GitHub’s servers. That’s when shit started to hit the fan." ... to get the conversation started. Your other points were sensible and lucid. This was a distraction and had the paradoxical effect of making me more sympathetic to github. The same looks to be true of other posters in this thread. If you were consciously choosing to take a hyperbolic tone, can I ask if you might reconsider that decision in future posts? Or at least concretely test your idea that calling people "dickwads" and "grunts" gets you more traction. I appreciated you raising the bandwidth question, and comparing it with other services. You made a good argument. Thank you!
- nacs 11y agoIt's his personal blog not some corporate blog or newspaper. I'm not sure it's within your rights to ask him to change the way he writes within his own bubble because you don't like his word choice.
- bsimpson 11y agoIt's just as much within his rights to call someone out for it as it is for someone to write as he likes. Clearly the post was written for an audience. Being needlessly inflammatory could certainly turn the audience off, and/or undercut the author's credibility. Ansiton's advice was both helpful and valid.
- dantiberian 11y agoThere's a lot of assumptions here about GitHub being greedy. I've got no idea how much money it costs GitHub to support Open Source projects, but it must easily be in the millions. I think that by this point GitHub deserves the benefit of the doubt before launching into vicious accusations.
- neandrake 11y agoThere's still the issue where turning on LFS makes forks unusable - regardless of price/profit this still seems like a major issue. Based off all the author's descriptions it sounds more like GitHub are still in the process of figuring out how best to get LFS management in their largely-collaborative environment. I wouldn't be surprised if GitHub has upcoming changes to resolve some of these issues.
- SwellJoe 11y agoWhile I don't think github is deserving of "vicious accusations", I do believe it is foolish to assume that the github we know today will be the github of tomorrow. SourceForge.net was once an excellent and trustworthy steward of Open Source software projects. It was predicted by some folks in the free software community that it would not always be the case, and alternatives like Savannah were maintained in order to act as a hedge against that concern. I believe it is more than reasonable to assume that github will change, and it would be downright dangerous to assume that we can rely on a profit-motivated corporation (even one as cool as github currently is) to remain a trustworthy repository forever. So, sure, say nice things about github; I also think github is a good product, and I appreciate their free hosting for OSS projects. And, sure, you should use github if it provides value for you and you're willing to accept the price. But, don't ask me to trust they'll never change, because history indicates they will. It's probably also unfair to suggest that someone criticizing some valid concerns about github's current behavior, based on their own experience with Open Source projects hosted at github, are making "vicious accusations".
- pbiggar 11y agoReally? These aren't vicious? "My guess is that some high-level greedy marketing dickwad, completely unaware of the asinine implications of his brilliant idea, signed off on this dumb-as-a-bag-of-rocks pricing model." "All the marketing material pimping GitHub’s LFS support [...]. I do not believe this is unintentional." "This is completely batshit. The side effect of this pernicious, greedy pricing model is to [...]" "I honestly couldn’t believe that GitHub would be willing to do something that shortsighted, visibly motivated by greed from the cash they thought they could extract from some of their users"
- i386 11y agoShock and horror: commercial company has a paid value add. GitHub is not a charity.
- facetube 11y agoI wouldn't necessarily describe a paid feature that breaks all forking for an entire repository a "value add".
- megastep 11y agoYes, my criticism was not that they're trying to make money from this, but rather that they are crippling their core product in an effort to monetize it some more.
- i386 11y agoNo, they are trying to solve a problem for their customers - storing large files. They charge money for this feature. It's disingenuous to think that they are breaking their product on purpose to extract money from you.
- megastep 11y agoIt is not disingenuous if it is exactly what they are doing. The entire cause of this problem, and the reason it breaks so many things, is because they are trying to monetize bandwidth usage. This model is unsustainable because free users only get a paltry 1GB/month and their attempt to enforce that is what breaks forks. If they would just stop trying to do that, then we would have nothing to talk about. It is not at all uncalled for to criticize the way they are trying to do business, especially when it affects you as an existing customer.
- i386 11y agoIf you value large file support over forking then its a value add.
- cwyers 11y ago> I honestly couldn’t believe that GitHub would be willing to do something that shortsighted, visibly motivated by greed from the cash they thought they could extract from some of their users. They're a business. C'mon here.
- megastep 11y agoMy argument was that this was actually bad for their business. This is not a customer-friendly move.
- i386 11y agoPeople who make products make these sort of tradeoffs all the time and the tradeoff is rarely something thats permanent. For the fast majority of people who need this kind of functionality (LFS), breaking forking is a inconvenience compared to the value being added.
- megastep 11y agoI would completely agree with you if that was the way GitHub presented that feature, being open about the consequences of adopting it. They haven't done that, and as a result their customers are not properly informed on the trade-offs they are making.
- cwyers 11y agoCustomers are people who pay you.
- eshamow 11y agoI'm not sure I understand why artifacts can't be stored in a different service - even an S3 bucket, if not a real repository service - and fetched dynamically via a build process. Is there a reason why binary blobs need to be stored directly next to code in order to be versioned?
- hyperpape 11y agoAside from a second point of failure, how does this integrate with anything? When you push, what piece of software pushes what where? And who pays?
- eshamow 11y agoYou can put this sort of build framework together with whatever tool you're using (gradle, maven, rake, grunt etc). The idea isn't to shove everything into a storage bucket, but to assemble a toolchain using components that are fit for purpose. Git is fundamentally not fit for purpose as an artifact repository. There are tools that are. -Eric
- megastep 11y agoThe way I understand it, it should be possible, though it's more complicated since they seem to infer the LFS URL from the repo URL by default. So if you wanted to say keep your repo on Github, and store your LFS files on S3, you'd need to explicitly tell git where to write the files. There are configuration values for that. Also you'd need the necessary LFS server piece on Amazon's side.
- eshamow 11y agoI'm thinking that rather than using git for versioning the binary artifacts as well, you tag and version your git repo, then tag and name/label your artifacts in another storage service. You then allow a build tool to assemble from both locations.
- nkurz 11y agoThis seems like an odd problem, but I'm not as familiar with Git as I should be. Is there not a reasonable way to download only the most recent version of these large binary files on the initial request, and then download the historical versions only in the (likely very rare) case that the user actually wants to use them? This would seem more useful in this case than hoping that binary diffs the repository small enough.
- adrusi 11y agoAlso not very well versed in git, but my understanding is that there is a way to clone a repo to only include latest revisions, but that this limits usage of git. I believe that fixing this was an area of active development a few months ago, its possible it already landed.
- gh02t 11y agoShallow clones are the term. It used to be that you couldn't pull remote changes or push local changes to/from a shallow clone, but that was fixed with v1.9 (early 2014). I'm not sure how LFS interacts with shallow clones though, as it's really a separate system that works in tandem with git more than a part of git itself.
- icebraining 11y agoIf you're talking about binary files merged into Git itself (not Git LFS, which is a separate mechanism), you can use "git clone --depth <n>" to get only the latest <n> revisions of the tree, and then use "git pull --unshallow" if you need to fetch the rest of the history.
- alkonaut 11y agoCan I do that automatically so that only binaries are fetched shallow, and text is fetched deep? Otherwise it's not very useful.
- icebraining 11y ago
- lemevi 11y agoEdit: I was wrong, however I learned from the conversation so I am leaving it here! Thanks to those who corrected me. > On the other hand, source files being mostly text, they are more intelligently handled and typically only differences between revisions are stored in the commits. This is completely incorrect, git stores whole blobs from one commit to the other. svn stored patches, but git does not. Every version of a file is stored in its entirety in your git tree since the beginning of the repository's existence. This is one of the reasons why git is so fast. You can go through your objects in your .git directory and verify this for yourself[0]. $ find .git/objects -type f .git/objects/ff/a5d733354ae6f8bdc67764d58d87c9a3161f66 .git/objects/ff/deb08f4856bd6eb5b31d7f800b3e480ae3e2e0 $ git cat-file -p ffa5d733354ae6f8bdc67764d58d87c9a3161f66 ...file contents appear... [0] https://git-scm.com/book/en/v2/Git-Internals-Git-Objects https://git-scm.com/book/en/v2/Git-Internals-Git-Objects
- saurik 11y agoThis is only true for recent commits: as you accumulate commits, garbage collections are performed of the loose blobs and the remaining generation is stored into a pack file, which has been carefully ordered by similarity and stored using a delta-encoding. For more information, this chapter from one of the popular online books about git might suffice. https://git-scm.com/book/en/v2/Git-Internals-Packfiles https://git-scm.com/book/en/v2/Git-Internals-Packfiles (edit: After I started responding to your comment, you edited your comment to link to the same book! I recommend you continue reading the later chapters: "you'll never believe how it works" ;P.)
- Vendan 11y agoFalse, git can do both. Run a git gc and check those files again. Chances are, many of your loose object files are missing, but everything still works
- general_failure 11y agoWhile not stored as 'text patches', when the objects are packed (as in pack files), they are stored as binary diffs.
- pbiggar 11y agoHow uncharitable can a single blog post be! The entire post is discredited by the author repeatedly projecting his unfounded opinions onto GitHub, such as "My guess is that some high-level greedy marketing dickwad, completely unaware of the asinine implications of his brilliant idea, signed off on this dumb-as-a-bag-of-rocks pricing model." "All the marketing material pimping GitHub’s LFS support [...]. I do not believe this is unintentional." "This is completely batshit. The side effect of this pernicious, greedy pricing model is to [...]" "I honestly couldn’t believe that GitHub would be willing to do something that shortsighted, visibly motivated by greed from the cash they thought they could extract from some of their users". Charitable explanation for forks not working: they haven't yet written the code to make this work with forks, and it's better to ship something working early, than to make it work in all cases. Charitable explanation for charging for bandwith: bandwidth costs money. (I believe this is a real problem for Dropbox, which doesn't charge for bandwidth but must still pay for it). Also, all CDNs, and also AWS charge for bandwidth. Overall, while GitHub may be able to support it's OSS folks better by changing the pricing on some parts of its product, this post is incredibly uncharitable. I hope the OP will consider removing the unfounded narrative that he's projecting onto GitHub (esp the "marketing dickwad" thing - wtf) and focus on the facts. [Disclaimer: my company partners with GitHub on lots of stuff]
- forrestthewoods 11y agoI wonder if Perforce Cloud will be able to fill this role at all. Probably not. Open Source isn't their target audience. But it could be a consideration. Has anyone tried the new Perforce/Git stuff? Is it any good? We're still on an older pre-Helix version.
- jedbrown 11y agoIn the interest of not propagating this common misconception: "The main problem with Git is that binary files are stored “as is” in the history of the project, so that every single revision of a new binary file (even if just a single byte has changed) is stored in full. [...] On the other hand, source files being mostly text, they are more intelligently handled and typically only differences between revisions are stored in the commits." This is false. Git stores the full version of each file in "loose" format and uses compressed incremental diffs (originally based on xdiff) in packfiles (after "git gc") without distinguishing text vs binary in either case. The issue is that binary files are often compressed themselves (so a one-byte semantic change has nonlocal effect) or have positional references (like jump targets in an executable, causing small changes to cascade). These factors explain the inefficient handling of binary files, but improving efficiency requires changing the semantics. LFS follows in the path of a few other tools (based on smudge/clean filters) that try to hide the semantic difference from the casual user, though that difference seems to bite people more frequently than we'd like.
- paulddraper 11y agoThis. Unlike many other systems, compression in git has nothing to do with commit order or file types or really anything VCS related. The way delta chains work in git are ingenious and transparent. The problem is that "binaries" are large amounts of data with high entropy.
- bhuga 11y agoI suspect setting up the free LFS reference/test server[1] that GitHub provides would have taken less time than writing this post complaining that GitHub isn't free enough. 1: https://github.com/github/lfs-test-server https://github.com/github/lfs-test-server
- nmc 11y agoThis being a test server implementation probably indicates that it is not meant to be run in production environment.
- sytse 11y agoAt GitLab we're working to support LFS. Initial support might or might not work with forks. As now with our Git Annex support storage will be free with a soft limit of 10GB of disk space per project (includes Git, Git Annex and Git LFS data) and there is no bandwidth limit. It will work with public and private projects (both are free).