5 ms·
I sympathise a lot with this post! Git cloning can be shockingly slow. As a personal anecdote, clones of the Rust repository in CI used to be pretty slow, and
by aidanhs 6y ago
I sympathise a lot with this post! Git cloning can be shockingly slow.
As a personal anecdote, clones of the Rust repository in CI used to be pretty slow, and on investigating we found out that one key problem was cloning the LLVM submodule (which Rust has a fork of).
In the end we put in place a hack to download the tar.gz of our LLVM repo from github and just copy it in place of the submodule, rather than cloning it. [0]
Also, as a counterpoint to some other comments in this thread - it's really easy to just shrug off CI getting slower. A few minutes here and there adds up. It was only because our CI would hard-fail after 3 hours that the infra team really started digging in (on this and other things) - had we left it, I suspect we might be at around 5 hours by now! Contributors want to do their work, not investigate "what does a git clone really do".
p.s. our first take on this was to have the submodules cloned and stored in the CI cache, then use the rather neat `--reference` flag [1] to grab objects from this local cache when initialising the submodule - incrementally updating the CI cache was way cheaper than recloning each time. Sadly the CI provider wasn't great at handling multi-GB caches, so we went with the approach outlined above.
[0] https://github.com/rust-lang/rust/blob/1.47.0/src/ci/init_repo.sh#L50-L68 https://github.com/rust-lang/rust/blob/1.47.0/src/ci/init_re...
[1] https://github.com/rust-lang/rust/commit/0347ff58230af512c9521bdda7877b8bef9e9d34#diff-a14d83f2e928fc5906d026a42cb16f021b452709b88bc3fd85c63e741cbd9a42R70 https://github.com/rust-lang/rust/commit/0347ff58230af512c95...
- bertr4nd 6y ago> Contributors want to do their work, not investigate "what does a git clone really do". Exactly this. Especially if the repo and CI pipeline are complicated, it is incredibly easy to just assume “it’s slow” is a fact of life. And from the point of view of the dev-productivity team, well, they have tons of possible issues to deal with at any given time. Not just CI but the repos themselves, the build system, maybe IDEs, debuggers, ... Sure the fix ends up being easy but you have to know to go looking for it.
- IggleSniggle 6y agoWhen you’ve got a billion other tasks to do, you might even know that it could be orders of magnitude faster and still not fix it, simply because of higher priority work. Frankly, I’d rather spend extra time trying to address problems/bugs/potential security holes in the actual shipped code than in fixing a poorly working CI pipeline...and I’m the kind of dev who gets really irritated by these problems. But you have to prioritize. Basically, barring “external” forces like cost overflow, customer unhappiness, or similar...stuff like that gets fixed at an equilibrium point between how much the problem hurts the dev, how adjacent to the codebase the devs current work is, and how interesting/irritating the dev finds the problem.
- auscompgeek 6y agoOut of curiosity, why not use the submodule.<name>.shallow option in .gitmodules?
- aidanhs 6y agoPrimarily because, until you mentioned it now, I wasn't even aware it was an option! That said, I generally shy away from shallow clones and probably wouldn't use it here: - it's a trap for people who ever want to work in that repo normally (we use the trick for more than just LLVM) - I believe shallow clones, over time (e.g. for contributors), are less nefficient than deep clones - I would expect shallow cloning to reuse fewer objects and benefit less from git's design. [0] describes a historic issue on this topic [0] https://github.com/CocoaPods/CocoaPods/issues/4989#issuecomment-193838934 https://github.com/CocoaPods/CocoaPods/issues/4989#issuecomm...