8 ms·
Git partial clone lets you fetch only the large file you need
- piliberto 7y ago> One reason projects with large binary files don't use Git is because, when a Git repository is cloned, Git will download every version of every file in the repository. Wrong? There's a --depth option for the git fetch command which allows the user to specify how many commits they want to fetch from the repository
- sewer_bird 7y agoYes, but 95% of devs, even fairly talented ones, don't really know how to use Git.
- colonwqbang 7y agoAuthor seems to be a manager, not necessarily a dev.
- phaemon 7y agoGit is fundamentally very simple. Any dev who doesn't understand exactly how it works is not even remotely talented.
- shaklee3 7y agodepth just let you control the amount of history. It will not let you exclude files that are at the highest depth that you don't want. So while that statement was not accurate, it's not what this feature is intended for.
- ComputerGuru 7y agodepth is broken; it cannot be used for submodules/recursive submodules dependably because most hosts will refuse to serve unadvertised refs. We learned this the hard way. Or maybe it is submodules that are broken. Or git itself.
- chrismorgan 7y agoDepth, submodules and multiple work trees are all half-baked features that work fine up to a point, then start falling over frantically—most notably if you try to use them together.
- shaklee3 7y agoThis is great. We use get lfs extensively, and one of the biggest complaints we have is users have to clone 7GB of data just to get the source files. There's a work around in that you don't have to enter your username and password from the lfs repo, and let it timeout, but that's a kluge.
- elephantum 7y agoThere’s an option for that: GIT_LFS_SKIP_SMUDGE=1 git clone SERVER-REPOSITORY
- derefr 7y agoHas anyone used Git submodules to isolate large binary assets into their own repos? Seems like the obvious solution to me. You already get fine-grained control over which submodules you initialize. And, unlike Git LFS, it might be something you’re already using for other reasons.
- matheusmoreira 7y agoThe problem with git submodules is they can't be used like a hyperlink to another repository. Updating the submodule requires updating the superproject as well. The new commits are invisible to the superproject until that is done. It'd be great if they worked like Python's editable package installations.
- ComputerGuru 7y agoThey can now, with the new-ish submodule update/init --remote. But the problem with sub modules is that you cannot do a shallow fetch (depth 1) because most hosts won’t serve unadvertised refs.
- sjburt 7y agoThen the state of the superproject would depend on when the checkout occurred. That would be disastrous for consistency, you’d be unable to replicate a checkout later or elsewhere. The state of a repo after a checkout should only depend on the commit that was checked out.
- matheusmoreira 7y agoThis is great for vendoring external dependencies that aren't under the developer's control. When the same developer is working on several related but separate projects at the same time, it's too cumbersome. Would be nice if git submodules could also point to a branch instead of specific commits. That way, the superproject's state would not be modified every time the branch is updated.
- 7y ago
- itroot 7y agoAlso --reference (or --shared) is a good parameter to speed-up cloning (for build, for example), if you have your repository cached in some other place. I was using it a long time ago when I was working on system that required to clone 20-40 repos to build. This approach decreased clone timings by an order of magnitude.
- mikepurvis 7y agoDo you actually need clones in that scenario? I worked on a build system that grabbed source from several hundred repos at the starting point, and it turned out to be way faster to just grab it all as tarballs with aria2c.
- madsbuch 7y agoGrapping the tarbell from where? To my best knowledge, tarbell export is not a part of git, but something git hosts provide. Git is a distributed VCS, and we should support keeping it that way.
- mikepurvis 7y agoAlmost any project you work on will have an authoritative copy of the repo in some kind of web-accessible tool, most of which provide a tarball-download function. And GitHub's scheme is pretty much a de-facto standard at this point—GitLab's implementation is an exact copy of it, for example: https://<host>/<org/project>/archive/<ref/branch/tag>.tar.gz Edit to add: Also, git-archive --remote is actually most of the way there, but it's not an HTTP download, of course. :(
- saagarjha 7y agoGitHub doing something one way and GitLab copying it doesn't make a standard.
- andrewshadura 7y agoCareful, with extra large repositories it actually slows down the cloning while, obviously, significantly reducing the space usage.
- smitty1e 7y agoIn AWS, it's worth considering putting those large files in an S3 bucket.
- nikivi 7y agoIs it possible given a git repo (hosted on say GitHub) to only 'clone' (download) certain files from it? Without `.git`
- fizixer 7y agoI believe you're looking for the 'working tree' only. You could do the following: git archive --remote=<your-URL> | tar -t source: https://stackoverflow.com/questions/3946538 https://stackoverflow.com/questions/3946538
- calvinlh 7y agoIf you only want a subset of the repo's files, you can use Github's Subversion interface: https://stackoverflow.com/a/18194523 https://stackoverflow.com/a/18194523
- bspammer 7y agoShort answer is, not easily: https://stackoverflow.com/a/14610427 https://stackoverflow.com/a/14610427 You can get the most recent tree for a repository (no history, just the current state of the repo) with `git clone --depth=1`. That's often good enough for slow connections.
- scarecrow112 7y agoThis is interesting and could be a savior for Machine Learning(ML) engineering teams. In a typical ML workflow, there are three main entities to be managed: 1. Code 2. Data 3. Models Systems like Data Version Control(DVC) [1], are useful for versioning 2 & 3. DVC improves on usability by residing inside the project's main git repo while maintaining versions of the data/models that reside on a remote. With Git partial clone, it seems like the gap could still be reduced between 1 & 2/3. [1] - https://dvc.org/ https://dvc.org/
- vvanders 7y agoAlso known as workspace views in P4. It's interesting to see the wheel reinvented. We used to run a 500gb art sync/200gb code sync with ~2tb back end repo back when I was in gamedev. P4 also has proper locking, it is really the is right tool if you've got large assets that need to be coordinated and versioned. Only downside of course is that it isn't free.
- jasondclinton 7y agoThis kind of comment isn't helpful. Of course, there have been ways to copy large files around since there were networks. What's new in this protocol enhancement is that this works within the context of a Merkle tree-based technology (upon which all DVCS's are based). To use your analogy, yes this is a wheel but it's built with rubber instead of wood and iron.
- vvanders 7y agoI guess I should have expanded more. DVCS is in direct opposition of workflows that include binary files(yes I'm aware that git lfs has locking, it's also centrally orchestrated) because you can't merge almost every binary format. We were using P4 ~15 years ago for these workflows and rather than understanding what made them work people are just rediscovering the same problems that have already been solved. My guess is that we'll next see a solution that dynamically caches most downloaded files in a geographic friendly way, heck we may even call it "P4Proxy". I've seen so much FUD around how git is the "one true workflow" because other solutions "don't scale" when they don't understand the constraints that certain workflows impose. Git/DVCS is great for a lot of things but sometimes you should use the right tool for the job rather than hack something together. [Edit] These reasons are exactly why you see Unreal supporting P4/SVN out of the box[1] and no mention of git. [1] https://docs.unrealengine.com/en-US/Engine/UI/SourceControl/index.html https://docs.unrealengine.com/en-US/Engine/UI/SourceControl/...
- dagmx 7y agoI think this is a very p4 centric view of the world. Locking helps with preventing collisions, but honestly the issue is still always communication. Why are people even touching files they shouldn't be touching? Meanwhile perforce is a pain for code heavy projects and requiring a central perforce server. Git works great there. The issue is neither is a silver bullet for the others workflow and needs, and they both suck horribly for mixed code and binary asset workflows. That's not even considering cost. Meanwhile film studios generally prefer keeping the considerations separate and using symlinks or URI to their data store and that works really well. But that doesn't work great for remote workflows. So again, I think you're applying a very p4, game centric view to this. There are lots of different use cases and team structures that none of these version control systems are able to address in their entirety.
- krupan 7y agoI started a project recently and for the first time ever I've wanted to keep large files in my repo. I looked into git LFS and was disappointed to learn that it requires either third party hosting or setting up a git LFS server myself. I looked into git annex and it seems decent. This, once it is ready for prime time, will hopefully be even better
- danbolt 7y agoIn the AAA games industry git has been a bit slower on the uptake (although that’s changing quickly) as large warehouses of data are often required (eg: version history of video files, 3D audio, music, etc.). It’s nice to see git have more options for this sort of thing.
- jfkebwjsbx 7y agoGit LFS has been a thing for years, though.
- danbolt 7y agoYou’re absolutely right, but larger developers and publishers have been slower to adopt. P4’s GUI/model is also intuitive for non-programming roles to learn and use historically compared to git, so a team with wide skills can ramp up quickly with a unified toolset. A less-technical manager gets a GUI that has versioning across changes from a multidisciplinary team. You can probably guess what inertia that has in a space with higher turnover compared to other industries. As mentioned, things are changing though. git and GitHub have become a mainstay and are what new programmers likely learn in schools. This has a trickle effect on new projects with smaller teams and results in more investment into git setups. I use git in a AAA context at work, and it’s not uncommon to find sentiments from more seasoned game programmers on git that are similar to HN comments about the latest fad in web frameworks.
- jfkebwjsbx 7y agoThe problem is that you blamed Git, rather than your legacy workflows and corporate culture. And again, you keep blaming Git here, now around a lack of intuitiveness and a lack of GUI. Again, Git has multiple GUIs around to choose from and multiple integrations with almost any editor and IDE you can think of, some meant for beginners and trivial usage. And no, things are not "changing" and Git is not to be compared with a "web framework fad". Git became the version control system more than 5 years ago.
- 7y ago
- vicosity 7y agoI'm still unconvinced. Will this provide a user friendly approach to managing design assets.
- madsbuch 7y agoMy impression is that it will use the normal git experience managing design assets. Ie. with this there should be no need for additional tooling. If it works, that would be so great!
- jniedrauer 7y agoThis could actually be a really good solution to the maximum supported size of a Go module. If you place a go.mod in the root of your repo, then every file in the repo becomes part of the module. There's also a hardcoded maximum size for a module: 500M. Problem is, I've got 1G+ of vendored assets in one of my repos. I had to trick Go into thinking that the vendored assets were a different Go module[0]. Go would have to add support for this, but it would be a pretty elegant solution to the problem. [0]: https://github.com/golang/go/issues/37724 https://github.com/golang/go/issues/37724
- lima 7y agoThat does sound like a "you're holding it wrong" issue. As one of the Go team members pointed out, defining a separate module is not a hack, but the intended way of doing it. How would a partial checkout help?
- jniedrauer 7y agoGo modules are built around git, unlike many other languages package systems. That means you don't get to pick and choose what goes into them. Imagine if you had to put an empty package.json in every (non-node) directory of your git repo to exclude it from an NPM package, or an install.py in every (non-python) directory to exclude it from a PyPI package. Multi-language repos would get ridiculous pretty quickly.
- lima 7y agoThis is only an issue if you put a go.mod file at the top level of your repo. We have monorepos with hundreds of modules.
- yencabulator 7y ago> Go modules are built around git Not really. Modules are specced based on zip files and metadata in text files. There's just support for extracting that data from git repos transparently. Here's a slightly out of date write-up: https://research.swtch.com/vgo-module https://research.swtch.com/vgo-module
- beagle3 7y agoThere is one note piece to the puzzle to make git perfect for every use case I can think of: store large files as a list of blobs broken down by some rolling hash a-la rsync/borg/bup. That would e.g. make it reasonable to check in virtual machine images or iso images into a repository. Extra storage (and by extension, network bandwidth) would be proportional to change size. git has delta compression for text as an optimization but it’s not used on big binary files and is not even online (only on making a pack). This would provide it online for large files. Junio posted a patch that did that ages ago, but it was pushed back until after the sha1->sha256 extension.
- pas 7y agoDo ISOs and other large blob types support only partial (block) modification? Wouldn't all subsequent blocks change too?
- beagle3 7y agoSometimes they do - e.g. if you replace a file in the ISO that is the same size up to block alignment, which is common when e.g. editing a text file or recompiling an executable with a minor change. They almost always do when it's a VM image representing a disk - only some blocks change every write. However, with self synchronizing hashes of the kind used by rsync bup and borg, it doesn't matter - you could have a 1TB file, delete a single byte at position 100 - and you only need to store or transfer one new block (with average size 8KB for rsync, configurable for borg) if you already have a copy of the version before the change. It's somewhat comparable with diff/patch but not exactly; it's worse in that change granularity is only specified on average; It's better in that it works well on binary files, does not require a specific reference diff (can reference all previous history), and efficiently supports reordering as well small changes - if you divide a 4000 line text file to four 1000-line sections and reorder them 1,2,3,4 -> 3,1,4,2 you will find the diff/patch to be as long as a new copy, whereas a self synchronizing hash decomposition will hardly take any space for the reordered file given the original.
- pas 7y ago
- microtherion 7y agoThat seems quite useful, though Git LFS mostly does the job. One of my biggest remaining pain points is resumable clone/fetch. I find it near impossible to clone large repos (or fetch if there were lots of new commits) over a slow, unstable link, so almost always I end up cloning a copy to a machine closer to the repo, and rsyncing it over to my machine.
- hinkley 7y agoWhat’s your take on this line? > Partial Clone is a new feature of Git that replaces Git LFS and makes working with very large repositories better by teaching Git how to work without downloading every file.
- microtherion 7y agoI believe partial clone makes the situation a little better, but it's not nearly as good as resumable cloning, because you have to partition your repo in advance.