12 ms·
We shrunk our Javascript monorepo git size
- dangsux 2y ago[dead]
- fragmede 2y ago> Large blobs happens when someone accidentally checks in some binary, so, not much you can do > Retroactively, once the file is there though, it's semi stuck in history. Arguably, the fix for that is to run filter-branch, remove the offending binary, teach and get everyone setup to use git-lfs for binaries, force push, and help everyone get their workstation to a good place. Far from ideal, but better than having a large not-even-used file in git.
- larusso 2y agoThe main issue is not a binary file that never changes. It’s the small binary file that changes often.
- abound 2y agoThere's also BFG (https://rtyley.github.io/bfg-repo-cleaner/ https://rtyley.github.io/bfg-repo-cleaner/) for people like me who are scared of filter-branch. As someone else noted, this is about small, frequently changing files, so you could remove old versions from the history to save space, and use LFS going forward.
- lastdong 2y agoIt’s easier to blame Linus.
- cocok 2y agofilter-repo is the recommended way these days: https://github.com/newren/git-filter-repo https://github.com/newren/git-filter-repo
- deleted 2y ago[deleted]
- killingtime74 2y agoShrank
- Sparkyte 2y agoShrinky dinky
- bubblesnort 2y agoHoney, I shrunk the git!
- dougthesnails 2y agoI think I prefer shrunked in this context.
- peutetre 2y agoI was in the pool!
- tankenmate 2y agoWould be correct if it is "We shrank", but from my poor memory of the terminology that is the transitive form, shrunken is the intransitive form. But once again from my poor memory.
- darraghenright 2y agoI've spoken English as my native language for almost five decades and I've never seen/heard the word "shranked" before. This surely cannot be correct. Even the title of the linked article doesn't use "shranked". What?
- forgotpwd16 2y agoCommonly (since ca. 19th century), shrank is used as the past tense of shrink, shrunk as the past particle, and shrunken as an adjective. The title of the linked article uses "shrunk" as past tense and the submitted title was changed to "shrunked" for some reason. "Shranked" was not mentioned anywhere. (But "shrinked" has had some use in the past.)
- eviks 2y agoupd: silly mistake - file name does not include its full path The explanation probably got lost among all the gifs, but the last 16 chars here are different: > was actually only checking the last 16 characters of a filename > For example, if you changed repo/packages/foo/CHANGELOG.md, when git was getting ready to do the push, it was generating a diff against repo/packages/bar/CHANGELOG.md!
- p4bl0 2y agoI was also bugged by that. I imagine that the meta variables foo and bar are at fault here, and that probably the actual package names had a common suffix like firstPkg and secondPkg. A common suffix of length three is enough in this case to get 16 chars in common as "/CHANGELOG.md" is already 13 chars long.
- derriz 2y agoI wish they had provided an actual explanation of what exactly was happening and skipped all the “color” in the story. By filename do they mean path? Or is it that git will just pick any file with a matching name to generate a diff? Is there any pattern to the choice of other file to use?
- snthpy 2y ago+1
- daenney 2y agoFile name doesn’t necessarily include the whole path. The last 16 characters of CHANGELOG.md is the full file name. If we interpret it that way, that also explains why the filepathwalk solution solves the problem. But if it’s really based on the last 16 characters of just the file name, not the whole path, then it feels like this problem should be a lot more common. At least in monorepos.
- floam 2y agoIt did shrink Chromium’s repo quite a bit!
- deleted 2y ago[deleted]
- triyambakam 2y ago> we have folks in Europe that can't even clone the repo due to it's size. What is it about Europe that makes it more difficult? That internet in Europe isn't as good? Actually, I have heard that some primary schools in Europe lack internet. My grandson's elementary school in rural California (population <10k) had internet as far back as 1998.
- RadiozRadioz 2y agoAt least here in Western Europe, in general the internet is great. Though coverage in rural areas varies by country.
- gnrlst 2y agoIn most EU countries we have multi-gigabit internet (for cheap too). Current offers are around ~5 GBIT speeds for 20 bucks a month.
- n_ary 2y agoWell good for you. On my side of europe, I pay €50/- for a cheap 50Mbps(1 month cancellation notice period). I could get a slightly cheaper 100Mbps from a predator for €20/- for first 6 month but then it goes up to €50/- and they pull bs about not being able to cancel if you even move because your new location is also in their coverage area(over garbage copper) and suffers at least 20 outages per month while there are other providers with much cheaper rates and better service. Some EU is still suffering from Telekom copper barons.
- jillesvangurp 2y agoSadly, I'm in Germany. Which is a third world country when it comes to decent connectivity. They are rolling out some fiber now in Berlin. Finally. But very slowly and not to my building any time soon. Most of the country is limited to DSL speeds. Mobile coverage is getting better but still non existent outside of cities. Germany has borders with nine countries. Each of those have better connectivity than Germany. I'm from the Netherlands where over 90% of households now have fiber connections, for example. Here in Berlin it's very hard to get that. They are starting to roll it out in some areas but it's taking very long and each building has to then get connected, which is up to the building owners.
- yunusabd 2y ago> For many reasons, that's just too big, we have folks in Europe that can't even clone the repo due to it's size. What's up with folks in Europe that they can't clone a big repo, but others can? Also it sounds like they still won't be able to clone, until the change is implemented on the server side? > This meant we were in many occasions just pushing the entire file again and again, which could be 10s of MBs per file in some cases, and you can imagine in a repo The sentence seems to be cut off. Also, the gifs are incredibly distracting while trying to read the article, and they are there even in reader mode.
- anon-3988 2y ago> For many reasons, that's just too big, we have folks in Europe that can't even clone the repo due to it's size. I read that as an anecdote, a more complete sentence would be "We had a story where someone from Europe couldn't clone the whole repo on his laptop for him to use on a journey across Europe because his disk is full at the time. He has since cleared up the disk and able to clone the repo". I don't think it points to a larger issue with Europe not being able to handle 180GB files...I surely hope so.
- peebeebee 2y agoThe European Union doesn't like when a file get too big and powerful. It needs to be broken apart in order to give smaller files a chance of success.
- thrance 2y ago
- bubblesnort 2y ago> We work in a very large Javascript monorepo at Microsoft we colloquially call 1JS. I used to call it office.com.. Teams is the worst offender there. Even a website with a cryptominer on it runs faster than that junk.
- wodenokoto 2y agoWe were all impressed with google docs, but office.com is way more impressive. Collaborative editing between a web app, two mobile anpps and a desktop app with 30 years of backwards compatibility and it pretty much just works. No wonder that took a lot of JavaScript!
- tinco 2y agoTo be fair, we were impressed with Google Docs 15 years ago. Not saying office.com isn't impressive, but Google Docs certainly isn't impressive today. My company still uses GSuite, as I don't like being in Microsoft's ecosystem and we don't need any advanced features of our office suite but Google Docs and the rest of the GSuite seem to be intentionally held back to technology of the early 2010's.
- alexanderchr 2y agoGoogle docs certainly haven't changed much the last 5-10 years. I wonder if that's an intentional choice, or if it is because those that built it and understand how it works are long gone to work on other things.
- jakub_g 2y agoActually I did see a few long awaited improvements landing in gdocs lately (e.g. better markdown support, pageless mode). I think they didn't deliver much new features in early 2020s because they were busy with a big refactoring from DOM to canvas rendering [0]. [0] https://news.ycombinator.com/item?id=27129858 https://news.ycombinator.com/item?id=27129858
- 2y ago
- issung 2y agoHaving someone in arms reach to help out that knows the inner workings of Git so much must be a lovely perk of working on such projects at companies of this scale.
- jonathanlydall 2y agoCertainly being in an org which has close ties to entities like GitHub helps, but any team in any org with that number of developers can justify the cost of bringing in a highly specialized consultant to solve an almost niche problem like this.
- jimjimjim 2y agoDid anybody else shudder at "Shrunked"?
- 0points 2y agoEnglish is my third language, also yes.
- amsterdorn 2y agoHoney, I done shrunked them kids
- tankenmate 2y agoShrunken, shrunked ain't no language I ever heard of.
- deleted 2y ago[deleted]
- snthpy 2y agoThanks for this post. Really interesting and a great win for OSS! I've been watching all the recent GitMerge talks put up by GitButler and following the monorepo / scaling developments - lots of great things being put out there by Microsoft, Github, and Gitlab. I'd like to understand this last 16 char vs full path check issue better. How does this fit in with delta compression, pack indexes, multi-pack indexes etc ... ?
- _joel 2y ago> Really interesting and a great win for OSS! Are they going to be opening a merge request to get their custom git command back in git proper then?
- acdha 2y agoIt appears so: https://lore.kernel.org/git/pull.1785.git.1725890210.gitgitgadget@gmail.com/ https://lore.kernel.org/git/pull.1785.git.1725890210.gitgitg...
- rettichschnidi 2y agoI'm surprised they are actually using Azure DevOps internally. Creating your own hell I guess.
- jonathanlydall 2y agoI find the “Boards” part of DevOps doesn’t work well for us a small org wanting a less structured backlog, but for components like Pipelines and the Git repositories it’s neither here nor there for us. What aspects of Azure DevOps are hell to you?
- rettichschnidi 2y agoSome examples, in no particular order. Hampering the productivity: - Review messages get sent out before review is actually finished. It should be sent out only once the reviewer has finished the work. - Code reviews are implemented in a terrible way compared to GitHub or GitLab. - Re-requesting a review once you did implemented proposed changes? Takes a single click on GitHub, but can not be done in Azure DevOps. I need to e.g. send a Slack message to the reviewer or remove and re-add them as reviewer. - Knowing to what line of code a reviewer was giving feedback to? Not possible after the PR got updated, because the feedback of the reviewer sticks to the original line number, which might now contain something entirely different. - Reviewing the commit messages in a PR takes way too many clicks. This causes people to not review the commit messages, letting bad commit messages pass and thus making it harder for future developers trying to figure out why something got implemented the way it did. Examples: - Too many clicks to review a commit message: PR -> Commits -> Commit -> Details - Comments on a specific commit does not shown in the commits PR - Unreliable servers. E.g. "remote: TF401035: The object '<snip>' does not exist.\nfatal: the remote end hung up unexpectedly" happens too often on git fetch. Usually works on a 2nd try. - Interprets IPv6 addresses in commit messages as emoji. E.g. fc00::6:100:0:0 becomes fc00::60:0. - Can not cancel a stage before it actually has started (Wasting time, cycles) - Terrible diffs (can not give a public example) - Network issues. E.g. checkouts that should take a few seconds take 15+ minutes (can not give a public example) - Step "checkout": Changes working folder for following steps (shitty docs, shitty behaviour) - The documentation reads as if their creators get paid by the number of words, but not for actually being useful. Whereas GitHub for example has actually useful documentation. - PR are always "Show everything", instead of "Active comments" (what I want). Resets itself on every reload. - Tabs are hardcoded (?) to be displayed as 4 chars - but we want 8 (Zephyr) - Re-running a pipeline run (manually) does not retain the resources selected in the last run Security: - DevOps does not support modern SSH keys, one has to use RSA keys (https://developercommunity.visualstudio.com/t/support-non-rsa-keys-for-ssh-authentication/365980#T-N10079417 https://developercommunity.visualstudio.com/t/support-non-rs...). It took them multiple years to allow RSA keys which are not deprecated by OpenSSH due to security concerns (https://devblogs.microsoft.com/devops/ssh-rsa-deprecation/ https://devblogs.microsoft.com/devops/ssh-rsa-deprecation/), yet no support for modern algos. This also rules out the usage of hardware tokens, e.g. YubiKeys. Azure DevOps is dying. Thus, things will not get better: - New, useful features get implemented by Microsoft for GitHub, but not for DevOps. E.g. https://devblogs.microsoft.com/devops/static-web-app-pr-workflow-for-azure-app-service-using-azure-devops/$ https://devblogs.microsoft.com/devops/static-web-app-pr-work... - "Nearly everyone who works on AzDevOps today became a GitHub employee last year or was hired directly by GitHub since then." (Reddit, https://www.reddit.com/r/azuredevops/comments/nvyuvp/comment/h1abcf0/ https://www.reddit.com/r/azuredevops/comments/nvyuvp/comment...) - Looking at Azure DevOps Released Features (https://learn.microsoft.com/en-us/azure/devops/release-notes/features-timeline-released https://learn.microsoft.com/en-us/azure/devops/release-notes...) it is quite obvious how much things have slowed down since e.g. 2019. Lastly - their support is ridiculously bad.
- tux3 2y agoFor those wondering where this new git-survey command is, it's actually not in git.git yet! The author is using microsoft's git fork, they've added this new command just this summer: https://github.com/microsoft/git/pull/667 https://github.com/microsoft/git/pull/667
- masklinn 2y agoI assume full-name-hash and path-walk are also only in the fork as well (or in git HEAD)? Can't see them in the man pages, or in the 2.47 changelog.
- tux3 2y agoYep. Path-walk is currently pending review here: https://lore.kernel.org/all/pull.1813.git.1728396723.gitgitgadget@gmail.com/T/ https://lore.kernel.org/all/pull.1813.git.1728396723.gitgitg... It more or less replaces the --full-name-hash option (again a very good cover letter that explains the differences and pros/cons of each very well!)
- clktmr 2y ago[flagged]
- throwuxiytayq 2y agoCan you elaborate how exactly git is at risk here? These posts never do.
- deleted 2y ago[deleted]
- AbuAssar 2y agothe gif memes were very distracting...
- jbverschoor 2y ago> those branches that only change CHANGELOG.md and CHANGELOG.json, we were fetching 125GB of extra git data?! HOW THO?? Unrecognized 100x programmer somewhere lol
- blumomo 2y ago[flagged]
- mark_and_sweep 2y agoAs a German, I assumed he's talking about poor connection speeds.
- mirekrusin 2y agoSize doesn't matter, it's how you use it (no invalid diffs on paths sharing trailing part).
- tom_ 2y agoThey're not actually smaller. It just looks like it because they're further away.
- nkmnz 2y ago> we have folks in Europe that can't even clone the repo due to it's size Officer, I'd like to report a murder committed in a side note!
- wodenokoto 2y agoNice to see that Microsoft is dog-fooding Azure DevOps. It seems that more and more Azure services only have native connectors to GitHub so I actually thought it was moving towards abandonware.
- jakub_g 2y agoParaphrasing meat of the article: - When you have multiple files in the repo which have the same trailing 16 characters in the repo path, git may wrongly calculate deltas, mixing up between those files. In here they had multiple CHANGELOG.md files mixed up. - So if those files are big and change often, you end up with massive deltas and inflated repo size. - There's a new git option (in Microsoft git fork for now) and config to use full file path to calculate those deltas, which fixes the issue when pushing, and locally repacking the repo. ``` git repack -adf --path-walk git config --global pack.usePathWalk true ``` - According to a screenshot, Chromium repacked in this way shrinks from 100GB to 22GB. - However AFAIU until GitHub enables it by default, GitHub clones from such repos will still be inflated.
- kreetx 2y agoI don't think GitHub, or any other git host, will have objections to using it once it's part of mainline git? Also, thank you for the TLDR!
- masklinn 2y ago> I don't think GitHub, or any other git host, will have objections to using it once it's part of mainline git? Fixing an existing repository requires a full repack, and for a repository as big as Chromium it still takes more than half a day (56000 seconds is 15h30), even if that's an improvement over the previous 3 days it's a lot of compute. From my experience of previous attempts, trying to get Github to run a full repack with harsh settings is extremely difficult (possibly because their infrastructure relies on more loosely packed repositories), I tried to get that for $dayjob's primary repository whose initial checkout had gotten pretty large and got nowhere. As of right now, said repository is ~9.5GB on disk on initial clone (full, not partial, excluding working copy). Locally running `repack -adf --window 250` brings it down to ~1.5GB, at the cost of a few hours of CPU. The repository does have some of the attributes described in TFA, so I'm definitely looking forward to trying these changes out.
- leksak 2y agoWouldn't a potential workaround be to create a new barebones repository and push the repacked one there? Sure, people will have to change their remote origin but if it solves the problem that might be worth the hassle?
- jakub_g 2y agoThe article mentions Derick Stolee who dig the digging and shipped the necessary changes. If you're interested in git internals, shrinking git clone sizes locally and in CI etc, Derrick wrote some amazing blogs on GitHub blog: https://github.blog/author/dstolee/ https://github.blog/author/dstolee/ See also his website: https://stolee.dev/ https://stolee.dev/ Kudos to Derrick, I learnt so much from those!
- develatio 2y agoHacking Git sounds fun, but isn't there a way to just not have 2.500 packages in a monorepo?
- Cthulhu_ 2y agoYeah, have 2500 separate Git repos with all the associated overhead.
- develatio 2y agoCan’t we split the packages into logical groups and maybe have 20 or 30 monorepos of 70-100 packages? I doubt that all the devs involved in that monorepo have to deal with all the 2500 packages. And I doubt that there is a circular dependency that requires all of these packages to be managed in a single monorepo.
- smashedtoatoms 2y agoPeople act like managing lots of git repos is hard, then run into monorepo problems requiring them to fix esoteric bugs in C that have been in git for a decade, all while still arguing monorepos are easy and great and managing multiple repos is complicated and hard. It's like hammering a nail through your hand, and then buying a different hammer with a softer handle to make it hurt less.
- crazygringo 2y ago> all while still arguing monorepos are easy and great I don't know anyone who says monorepos are easy. To the contrary, the tooling is precisely the hard part. But the point is that the difficulty of the tooling is a lot less than the difficulty of managing compatibility conflicts between tons of separate repos. Each esoteric bug in C only needs to be fixed once. Whereas your version compatibility conflict this week is going to be followed by another one next week.
- HdS84 2y ago
- vtodekl 2y ago[dead]
- mattlondon 2y agoI recently had a similar moment of WTF for git in a JavaScript repo. Much much smaller of course though. A raspberry pi had died and I was trying to recover some projects that had not been pushed to GitHub for a while. Holy crap. A few small JavaScript projects with perhaps 20 or 30 code files, a few thousand lines of code for a couple of 10s of KBs of actual code at most had 10s of gigabytes of data in the .git/ folder. Insane. In the end I killed the recovery of the entire home dir and had to manually select folders to avoid accidentally trying to recover a .git/ dir as it was taking forever on a poorly SD card that was already in a bad way and I did not want to finally kill it for good by trying to salvage countless gigabytes of trash for git.
- EDEdDNEdDYFaN 2y agobetter question - does the changelog need to be checked in the first place?
- DeathMetal3000 2y agoThey fixed a bug on a tool that is widely used. In what world is questioning why an organization is checking in a file that you have no context on a “better question”.
- tazjin 2y agoI just tried this on nixpkgs (~5GB when cloned straight from Github). The first option mentioned in the post (--window 250) reduced the size to 1.7GB. The new --path-walk option from the Microsoft git fork was less effective, resulting in 1.9GB total size. Both of these are less than half of the initial size. Would be great if there was a way to get Github to run these, and even greater if people started hosting stuff in a way that gives them control over this ...
- dizhn 2y agoThey call him Linux Torvalds over there?
- nixosbestos 2y agoOh hey I know that name, Stolee. Fellow JSR grad here.
- Vilian 2y agoPeople who use git in monorepos don't understand git
- nsonha 2y agoI think the title misses the "Honey, " part
- deleted 2y ago[deleted]