13 ms·
Announcing Git Large File Storage
- m0th87 11y agoOur solution is likely a lot more duct tape-y, but we developed a straight-forward tool in Go for managing large assets in git: https://github.com/dailymuse/git-fit https://github.com/dailymuse/git-fit There's a number of other solutions open source out there, some of which are documented in our readme.
- justinsb 11y agoThis solves a real problem, but I can't help but feel it is a band-aid hack. The main fundamental advantage (vs implementation quirks of git) I can see is that these files are only fetched on a git checkout. But (of course) this breaks offline support, and it requires additional user action. Wouldn't it have been fairly easy to build exactly the same functionality into git itself? "Big" blobs aren't fetched until they are checked-out? This also has the advantage the definition of "big" could depend on your connectivity / disk space / whatever, rather than being set per-repo.
- ocdtrekkie 11y agoI have a feeling the decision for how to arrange, as a separate thing, is likely to feed the monetization component. Particularly if it's limited to using GitHub's storage.
- acveilleux 11y agoSpecs are open, there's an open server implementation. It might be easiest to set it up with github, but if it catches on, I expect implementations will be readily available from all github-like platforms and as stand-alone.
- ocdtrekkie 11y agoAwesome. I did notice they called it "Git LFS" instead of "GitHub LFS", which should be a clue there, though from other comments I figured it might be GitHub specific.
- jxf 11y agoThis looks really interesting. You basically trade the ability to have diffs (nearly meaningless on binary files anyway) for representing large files as their SHA-256 equivalent values on a remote server. What will be interesting is to see whether GitHub's implementation of LFS allows a "bring your own server" option. Right now the answer seems to be no -- the server knows about all the SHAs, and GitHub's server only supports their own storage endpoint. So you couldn't use, say, S3 to host your Git LFS files.
- icebraining 11y agoYou basically trade the ability to have diffs (nearly meaningless on binary files anyway) for representing large files as their SHA-256 equivalent values on a remote server. That's exactly what git-annex does. Except it can host on your own servers, or S3, or Tahoe-LAFS, or rsync.net, etc. And it's free software. And it supports multiple servers for the same repo, so you have redundancy. Adding an S3 remote is just setting the AWS keys and running a single command: http://git-annex.branchable.com/tips/using_Amazon_S3/ http://git-annex.branchable.com/tips/using_Amazon_S3/
- azernik 11y agoThis is also free software, and you can also use your own server.
- efuquen 11y agoI think the question is why did they role their own solution when there was one already an open and freely one available. If it wasn't suitable in some way I would really like to know why.
- s73v3r 11y agoIf git-annex really worked well, and was easy to use, I imagine there'd be much more uptake of it.
- geoffreyirving 11y agoAfter a quick scan, I'm a bit worried that this is too tied to a server in practice. For example, if I've downloaded everything locally, can I easily clone the whole download (including all lfs files) into a separate repo? If I can, can changes to each be swapped back and forth?
- ot 11y ago> Every user and organization on GitHub.com with Git LFS enabled will begin with 1 GB of free file storage and a monthly bandwidth quota of 1 GB. Does this mean that with the free tier I can upload a 1GB file which can be downloaded at most once a month? Even a small 10MB file, which fits comfortably in a git repo, could be downloaded only 100 times a month. Maybe they meant 1TB bandwidth?
- mdlowman 11y agoI would suspect the point is rather that you have a bunch of megabyte range files, and you rarely update them and don't have to sync. But for most workflows this feature seems targeted at, the free tier seems insufficient.
- jandrese 11y agoI'm having trouble seeing where a 1GB/month quota in any way meshes with "large file" support. The free tier is basically "test out the API, don't even think about using it for real".
- cortesoft 11y agoYes, that is the free tier. If you want to use it seriously, it will cost some money. OR you can use it and host your own file server, for free. I don't think these facts are a problem. They create an open source tool, provide a location to try it out, and a service to pay to use it if you like it and don't want to host yourself. Seems like a fair offer.
- sytse 11y agoGitLab.com offers 5GB per repo support for git-annex (unlimited repos).
- jewel 11y agoThe "filter-by-filetype" approach used here is going to work a lot better for mixed-content repositories than git-annex, which doesn't have that capability built-in (to my knowledge). git-annex has been great for my photo collection (which is strictly binary files). It lets me keep a partial checkout of photos on my laptop and desktop, while replicating the backup to multiple hosts around the internet. At work we have a bunch of video themes that are partially XML and INI files and partially JPG and MP4. LFS would work great for us, except we don't use github (we don't have a need for it.) It looks like this is going to be very simple for that kind of workflow. Just yesterday HN user dangero was looking for this exact sort of thing, large file support in git that didn't add too much complexity to the workflow: https://news.ycombinator.com/item?id=9330125 https://news.ycombinator.com/item?id=9330125
- icebraining 11y agoThe filter-by-filetype can be replace by small script that augments git: http://git-annex.branchable.com/forum/help_running_git-annex_on_top_of_existing_repo/#comment-41049e5e5de08f6cc0b76f9298b8e4c0 http://git-annex.branchable.com/forum/help_running_git-annex...
- sytse 11y agoWould be nice to replace that with git hooks so you can just use regular git commands. Any idea's if that is feasible?
- icebraining 11y agoI'm not sure, but I doubt a hook would do. An alternative would be to have a frontend script to git that would shadow the git command (using shell aliases) and call git-annex when appropriate.
- sytse 11y agoYeah, I don't like shell aliases but it would work. I wonder how the LFS client works.
- vvanders 11y agoThis looks like it misses the mark a bit. As anyone who's worked on project with large binary files(the docs assume PSDs) you need to be able to lock unmergeable binary assets. Otherwise you get two people touching the same file and someone has to destroy their changes. That never makes anyone happy. It's also unseen how good the disk performance is. These two areas are the reason why Perforce is still my go-to solution for large binary files.
- jordigh 11y agoMercurial has largefiles and locking too: http://mercurial.selenic.com/wiki/LargefilesExtension http://mercurial.selenic.com/wiki/LargefilesExtension http://mercurial.selenic.com/wiki/LockExtension http://mercurial.selenic.com/wiki/LockExtension Like other people have noticed, you can have the good parts of a DVCS and the good parts of a CVCS. It doesn't have to be either-or: https://blogs.janestreet.com/centralizing-distributed-version-control-revisited/ https://blogs.janestreet.com/centralizing-distributed-versio... http://bitquabit.com/post/unorthodocs-abandon-your-dvcs-and-return-to-sanity/ http://bitquabit.com/post/unorthodocs-abandon-your-dvcs-and-...
- neandrake 11y agoWe turned on largefiles extension almost a year ago and have come to regret that decision. The major pain point is with integrations, from Eclipse plugin to most any repository management/hosting solution tend to either have bugs or full-on don't support those repos. With plain mercurial workflow it works in most situations, but still rears its ugly head. The maintainers of mercurial classify it as a "feature of last resort" (Read http://mercurial.selenic.com/wiki/FeaturesOfLastResort http://mercurial.selenic.com/wiki/FeaturesOfLastResort). Switching off largefiles requires rebuilding the repository which rebuilds the entire repository. Orchestrating the migration to a new repository for engineering department is also painful which is why we're stuck for the near future (for example ongoing support for a version that's built from largefiles repo with ongoing feature work in non-largefiles repo). The tooling for mercurial tends to lack behind git's, likely due to git's enormous popularity - so I personally would recommend avoiding largefiles extension.
- Doji 11y agoSo basically it's git-annex, but tied to GitHub. http://git-annex.branchable.com/ http://git-annex.branchable.com/
- scott_karana 11y agoYeah, that was exactly my feeling. "Not invented here" much?
- andrewchambers 11y agothey are trying to make a service. If you are making a product you generally want to be in control of its core parts.
- sytse 11y agoIf you use another open source project that gives you control right? It would be nice if everyone reused git-annex like they reused git.
- acveilleux 11y agoAs far as I can tell, less flexible (far less) then annex but it can be made very seamless to the user (no/few special commands and lfs tracking by filemask.) I would think bridging the gap in annex to track by file pattern would be easy but a lot of people might prefer not to know how to make annex go. So using simplicity as differentiator.
- IgorPartola 11y ago(Not a git-annex user here). I suppose functionally, these two are similar. But the use case is different. git-annex seems to be more for managing files and making sure they don't disappear on you. GitHub's new thing is for keeping track of larger objects inside your git project efficiently. Basically, yeah, you can use git-annex to store the PSD, the audio samples, the promo video, etc. but wouldn't it be nice to have it all tied in with your normal project workflow?
- 11y ago
- Rondom 11y agoHas someone had a closer look and can say how this compares to Git-Annex?
- jefurii 11y agoThis and git-annex (and git-fat and others) use the same basic architecture of storing links in Git and schlepping the binaries around separately. Git-annex renames binaries with their SHA256 hashes, puts them in a .git/annex/ dir, and replaces files in the working dir with symlinks. Git-LFS seems to use small metadata pointer files (SHA256 hash, file size, git-lfs version) instead of symlinks. Not sure whether the files reside in something like the .git/annex/ dir; I'm guessing the do or there wouldn't be those pointer files. You can clone a repo without having to download the files. With git-annex you can sync between non-bare and bare repositories without having a central server. Git-LFS seems to have a separate server for binaries. It looks like it may act like a git-annex special remote rather than git-annex's usage of synced/master branch. Git-annex repos share information about the locations of annex files, how many repos contain a given file, etc. You can trust and un-trust repos. It doesn't look like Git-LFS offers this. Git-LFS has a REST API. I'm using an old version of git-annex so I can't say if it does (I think it does). Git-LFS is written in Go, git-annex is Haskell. Git-LFS is a GitHub project. GitHub will offer object hosting. Update: clarity, speling, added a bullet point.
- icebraining 11y agogit annex uses symlinks in indirect mode, but can use small files in direct mode (useful for filesystems which don't support symlinks, like many Android sdcards).
- joeyh 11y agoThe lack of location tracking looks like the most significant difference to me. While the git-lfs documentation does mention that different git remotes can have different LFS endpoints configured, all git-lfs knows about a file is its SHA256. So how can it tell which remote to download the file from? The best it could do is try different remotes until it finds one that has the file. I hesitate to say this means git-lfs is not distributed at all, but it seems significantly less distributed than git-annex, which can keep track of files that might be in Glacier, or on an offline drive, or a repo cloned on a nearby computer, and so can be used in a more peer-to-peer fashion when storing and retrieving the large files.
- niche 11y agoYes! Bringing us all one step closer to the whiysi (we host it you store it) dev paradigm. Bravo!
- Pirate-of-SV 11y agoCan't wait too see what Linus got to say about this. I suppose he got an arguably better solution to the problem?
- jgrowl 11y agoMy guess is that Linus doesn't care about large binary files.
- mahouse 11y agoAh, cool! At last I will be able to store my database backups in GitHub.
- dmitrypolushkin 11y agoHopefully they will provide such functionality, but I don't think in the nearest future.
- duartetb 11y agoDoes this mean gamedevs might start droping Perforce for this? If its not too expensive maybe?
- eropple 11y agoMost developers I know dropped Perforce for git and svn a long time ago. Configure them with the appropriate ignores and it works fine.
- Impossible 11y agoThis hasn't been my experience working in AAA console games, although I could definitely see it being the case in mobile. What developers do you know?
- bananaboy 11y agoAgreed, a lot of developers still use Perforce in my experience (my own indie studio included). I don't think this will make anything better for game developers generally. The big stumbling blocks are: Git is not artist/non-programmer friendly; and an inability to lock a file when editing (specifically binary files or files that are difficult to merge, e.g. complex level files like Unity's).
- eropple 11y agoMobile and mid-tier console developers. Not AAA, but with budgets firmly in the "you spent what, for that?" range.
- Tiktaalik 11y agoThis is certainly the issue that is preventing game devs from adopting Git. On the other hand game devs at this point are very used to Perforce, and it looks like Perforce is interested in solving this problem from the other side, by adding Git features to Perforce Helix and making it distributed.
- dhruvgupta 11y ago
- Poiesis 11y agoHas anyone seen what happens for a user who doesn't have this installed when cloning? I've tried it out but it seems to not affect local clones.
- lesplat 11y agoSo does this mean the large files are actually versioned?
- gabeio 11y agoFrom how I read it, it sounds like a little of yes and no... it's similar to git's way but I am not sure if they are really going to keep versions of all of the old large files... I guess if they are going to be fully reverse compatible like being able to go backwards in git you have to...
- jbramble 11y agoDoes this mean github could become useful for music production?
- joeyh 11y agoIt's interesting that this uses smudge/clean filters. When I considered using those for git-annex, I noticed that the smudge and clean filters both had to consume the entire content of the file from stdin. Which means that eg, git status will need to feed all the large files in your work tree into git-lfs's smudge filter. I'm interested to see how this scales. My feeling when I looked at it was that it was not sufficiently scalable without improving the smudge/clean filter interface. I mentioned this to the git devs at the time and even tried to develop a patch, but AFAICS, nothing yet. Details: <https://git-annex.branchable.com/todo/smudge> https://git-annex.branchable.com/todo/smudge>
- joeyh 11y agoSeems that git status nowadays does manage to avoid running the smudge filter, unless the file's stat has changed. This overhead does still exist for other operations, like git checkout.
- bburky 11y agoAlso, is there any reason Git LFS can't be used as a special remote for git-annex? It would provide an easy way for people to host their git-annex repos entirely on GitHub.
- joeyh 11y agoYeah, git-annex is very interested in having a special remote for everything and anything. And if someone creates 4 shell commands, I could have a demo working in half an hour. The commands would be: lfs-get SHA256 > file lfs-store SHA256 < file lfs-remove SHA256 (optional) lfs-check SHA256 # exit 0 or 1, or some special code if github is not available Presumably the right way would be to use their http api, but these 4 commands seem generally useful to have anyway.
- jedbrown 11y agoAs the author of git-fat, I have to say the smudge/clean filter approach is a hack for large files and the performance is not good for a lot of use cases. The reality is that it's common to need fine-grained control over what files are really present in the repository, when they are cached locally, and when they are fetched over the network. Git-annex does better than the smudge/clean tools (git-fat, git-media, git-lfs) but at somewhat increased complexity. I think our tools have stepped over the line of "as simple as possible but no simpler" and cut ourselves off from a lot of use cases. Unfortunately, it's hard for people to evaluate whether these tools are a good fit now and in a couple years. As for git-lfs relative to git-fat: (1) the Go implementation is probably sensible because Python startup time is very slow, (2) git-lfs needs server-side support so administration and security is more complicated, (3) git-lfs appears to be quite opinionated about when files are transferred and inflated in the working tree. The last point may severely limit ability to work offline/on slow networks and may cause interactive response time to be unacceptable. Some details of the implementation are different and I'd be curious to see performance comparisons among all of our tools.
- zmmmmm 11y agoIs there any hint on pricing? Slighty annoying to have a section titled "Pricing" which .... doesn't tell you the price. I would much rather use my own external server for hosting large files, it is going to need to be price competitive with other options to be interesting I would think.
- ElectricFeel 11y agomy name is Lars & i do projects for LiveIT! this is exciting
- nodesocket 11y ago> Every user and organization on GitHub.com with Git LFS enabled will begin with 1 GB of free file storage and a monthly bandwidth quota of 1 GB. A GB doesn't get you very far if you are working with raw audio and video. Does it make sense to think about storing virtual machines images (.vmdk) in git on GitHub with LFS?
- theli0nheart 11y agoI'm sure GitHub did their due diligence before starting to work on this, but I can't lie: it bums me out a bit that they didn't find git-bigstore [1] (a project I wrote about 2 years ago) before they started, since it works in almost the exact same way. Three-line pointer files, smudge and clean filters, use of .gitattributes for which files to sync, and remote service integration. Compare "Git Large File Storage"'s file spec: version https://git-lfs.github.com/spec/v1 oid sha256:4d7a214614ab2935c943f9e0ff69d22eadbb8f32b1258daaa5e2ca24d17e2393 size 12345 And bigstore's: bigstore sha256 96e31e44688cee1b0a56922aff173f7fd900440f Bigstore has the added benefit of keeping track of file upload / download history _entirely in Git_, using Git notes (an otherwise not-so-useful feature). Additionally, Bigstore is also _not_ tied to any specific service. There are built-in hooks to Amazon S3, Google Cloud Storage, and Rackspace. Congrats to GitHub, but this leaves a sour taste in my mouth. FWIW, contributions are still welcome! And I hope there is still a future for bigstore. [1]: https://github.com/lionheart/git-bigstore https://github.com/lionheart/git-bigstore
- SomeCallMeTim 11y agoLooks like you and GitHub are also both duplicating the git-media extension [1]. I haven't settled on one for my own use, but I'll compare features of bigstore and git-media before I do. Thanks for making your project available! [1] https://github.com/alebedev/git-media https://github.com/alebedev/git-media
- jedbrown 11y agoLooks like you wrote git-bigstore a few months after I wrote git-fat (also Python and a similar design; partially inspired by git-media). It would be interesting to do some performance comparisons and merge our capabilities, perhaps with support for each other's stub formats if we can do it in a compatible way.
- theli0nheart 11y agoYep, agreed. That would be awesome.
- tomphoolery 11y ago
- sytse 11y agoI like the ease of use of 'git lfs track "*.psd"' and being able to use normal git commands after that. Would it be possible to extend git-annex with a command that lets you set one or more extensions? By using git hooks you can probably ensure that the normal git commands work reliably.
- Animats 11y agoDoes Github really do this using git's "smudge" and "clean" filters? That would mean reprocessing the whole file for each access. That's inefficient. It's useful only if someone else is paying for the disk bandwidth, and necessary only if you don't have control of the storage system. Why would GitHub do that to itself?
- markvitals 11y agoThis is very handy for designers, who want to use Photoshop or Illustrator with Git
- sytse 11y agoTo celebrate the broader support for git with large files we just raised the storage limit of GitLab.com to 10GB https://about.gitlab.com/2015/04/08/gitlab-dot-com-storage-limit-raised-to-10gb-per-repo/ https://about.gitlab.com/2015/04/08/gitlab-dot-com-storage-l... also, we're glad GitHub open sourced it and didn't call it assman
- sytse 11y agoSomeone asked if this was temporarily or permanent, it is permanent, see https://news.ycombinator.com/item?id=9344984 https://news.ycombinator.com/item?id=9344984
- luckydude 11y agoBitKeeper has had a better version of this since around 2007. Better in that we support a cloud of servers so there is no "close to the server" thing, everyone is close to the server. What we don't have is the locking. I agree with the people commenting here that locking is a requirement because you can't merge. We need to do that.
- patcon 11y agoCool! Can't want for future integration of content-addressable systems like ipfs :)
- saljam 11y agoWhat's stopping git from storing large files using Merkle trees + a rolling hash? I'm probably missing something since there this, and git-annex, and git-bigstore, and others...
- spb 11y agoI still don't get why you wouldn't just check large binaries into a submodule and host that everywhere you would an annex/LFS.
- deleted 11y ago[deleted]
- callum85 11y agoCan someone explain to me what problem this solves in layman's terms... How are version control systems are "impractical" for large files? Or to put another way, what problems will I run into if I just commit large media files without using this?
- ndepoel 11y agoWith distributed version control systems such as Git or Mercurial, when you clone a repository you get the entire history of that repository (or of a selected branch). This means that if you place large media files directly in the repository, then every clone will contain each and every revision of that file. In time, this will cause an enormous amount of bloat in your repository and slow work on the repository down to a crawl. Cloning a repository several dozens of gigabytes in size is no fun, I can tell you. Centralized version control systems such as Subversion don't have this problem (or at least, to a lesser extent), because as a user you only download a single revision of each file when you check out the repository. Extensions like git-media, git-fat and now git-lfs solve this issue by only storing references to large media files inside the Git repository, while storing the actual files elsewhere. With this, you will only download the revision of the large file that you actually need, when you need it. It's sort of a hybrid solution in-between centralized and decentralized version control.
- amelius 11y agoWouldn't it be nicer if we had something like this on the level of the filesystem, instead of on the level of a version control system? Advantages would be that git and any other user-space application wouldn't need much extension, and files could be opened as if they were on the local file system.
- dmitrypolushkin 11y agoHopefully bup will implement something like that for the backuping.
- silon3 11y agoCan it link to torrent?