4 ms·
We learned they were 70% autogenerated so probably shouldn't have been in git at all, but our build process relied on that, and didnt want to fix it, so we bodg
by time4tea 3y ago
We learned they were 70% autogenerated so probably shouldn't have been in git at all, but our build process relied on that, and didnt want to fix it, so we bodged it.
- maccard 3y ago> probably shouldn't have been in git at all Something being autogenerated, or binary, doesn't mean it shouldn't be in version control. If step one of your instructions to build something from version control involve downloading a specific version of something else, then your VCS isn't doing it's job, and you're likely skirting around it to avoid limitations in the tool itself. People still use tools like P4 because they want versioned binary content that belongs in version control, or because they want to handle half a million files, and git chokes. In my last org, we vendored our entire toolchain, including SDKs. The project setup instructions were: - Install p4 - Sync, get coffee - Run build, get more coffee. A disruptive thing like a compiler upgrade just works out of the box in this scenario. It's a shame that the mantra of "do one thing well" devolves into "only support a few hundred text files on linux" with git.
- PMunch 3y agoWouldn't Git LFS be the tool for this job? Have the automated tool build a .zip file for example of the translations (possibly with compression level set to 0), then have your build toolchain unzip the archive before it runs. Then check that big .zip file into GitLFS, et voila you now have this large file versioned in Git.
- maccard 3y agoGit LFS isn't the same as git, though. It's better than putting everything in a separate store, but for one it disables offline work, and breaks the concept of D in the DVCS of git. > then have your build toolchain unzip the archive before it runs My build toolchain shouldn't have to work around the shortcomings of my environment, IMO. > et voila you now have this large file versioned in Git. No, it's on a separate http server that is fetched via git lfs. Subtle, but important difference.
- aidenn0 3y ago> it disables offline work, This is a non-issue for images and autogenerated files, since you shouldn't ever be doing a merge on them. > breaks the concept of D in the DVCS of git. git-annex is distributed and works well for files that will never be merged (such as images, or autogenerated files)
- folmar 3y agoIt's good enough for the small usecases, but way behind tools that have first class support for binary files (binary deltas, common compression, ...). Even SVN shines here.
- haizzz 3y agoSeparately we also found that git lfs is not very optimised for large repositories, notably its locking feature which list every file tracked by git for every checkout and commit command.
- avidiax 3y ago> Something being autogenerated, or binary, doesn't mean it shouldn't be in version control. I think the SHA should be in version control. The file should be reproducibly built [1], then cached on a central server. This means that a build target like a system image could be satisfied by downloading the complete image and no intermediate files. And a change to one file in one binary will result in only a small number of intermediate files being downloaded or reproducibly built to chain up to the new system image. This is something that's really lacking in, for example, Git. [1] https://en.wikipedia.org/wiki/Reproducible_builds https://en.wikipedia.org/wiki/Reproducible_builds
- maccard 3y ago> I think the SHA should be in version control. The file should be reproducibly built [1], then cached on a central server. Requiring reproducible builds to handle translations or images is a bit much. Also, if it's cached on a central server, that now means you need to be connected to that central server. If you require a connection to said central server, why not just have your source code on said server in the first place, a la p4? I do agree that NixOS is a great idea, but personally 99% of my problems would be solved if git scaled properly.
- avidiax 3y agoYou can always build from source in this scenario. The cache server lets you skip two things. First, you can prune the leaves of the tree of intermediate files you might need. Second, where you do need to compile/build/link/package, etc., you can do only those steps that are altered by your changes. So you save CPU time and storage space. > why not just have your source code on said server in the first place, a la p4? That would be great. A version of git where cloning is almost a no-op, and building is downloading the package assuming you haven't changed anything. I'm not aware of how p4 allowing this. My recollection of perforce is that I still had most source files locally.
- jtsiskin 3y agohttps://git-lfs.com/ https://git-lfs.com/
- thechao 3y agoThis is precisely why every ASIC (HW) company I'm familiar with uses P4. ASIC design flows rely critically on 3rd party tooling, that must be version/release specific. You can't rely on those objects being available whenever. They get squirreled away and kept, forever.
- Karellen 3y ago> In my last org, we vendored our entire toolchain, You vendored all your compilers/language runtimes in the source control repo of each project? Including, like, gcc or clang? WTF? > It's a shame that the mantra of "do one thing well" devolves into "only support a few hundred text files on linux" with git. Because the Linux kernel source tree and its history can accurately be described as "a few hundred text files". Yeah, right.
- 0xcoffee 3y agoIt's not that unusual, we vendor entire VM images which contain the development environment. (Codebase existed since before docker). And it works well, need to fix something in a project that was last update 20 years ago? Just boot up the VM and you are ready.
- xorcist 3y agoI don't think that was the question but rather why commit to git? Having local commits intermingled with an upstream code base can make for really hairy upgrades, but I guess every situation is slightly different.
- maccard 3y ago> but rather why commit to git? Well we don't put them in git, we put them in perforce because git keels over if you try and stuff 10GB of binaries into it once every few months. I think the real question is the other way around though, why _not_ use git for versioning when that's what it's supposed to be for? Why do I have to verison some things with git, and others with npm/go build/pip/vcpkg/cargo/whatever?
- tom_ 3y agoI've worked on a couple of game projects that did this. Build on Windows PC, build for Windows/Switch/Xbox One/Xbox Serieses/PS4/PS5/Linux. I was never responsible for setting this up, and that side of things did sounda bit annoying, but it seemed to work well enough once up and running. No need to worry about which precise version of Visual Studio 2019 you have, or whether you've got the exact same minor revision of the SDK as everybody else. You always build with exactly the right toolchain and SDK for each target platform.
- tomjakubowski 3y agoDoes perforce have features which make vendoring easier? Just curious why I see P4 called out here and in the replies too.
- tom_ 3y agoIt just does a pretty good job of dealing with binary files in general. The check in/check out model is perfect for unmergeable files; you can purge old revisions; all the metadata is server side, so you only pay for the files you get; partial gets are well supported. And so, if you're going to maintain a set of tools that everybody is going to use to build your project, the Perforce depot is the obvious place to put them. Your project's source code is already there! (There are various good reasons why you might not! But "because binary files shouldn't go in version control" is not one of them)
- IshKebab 3y agoIt's not an unbreakable rule that generated or binary files should not be in Git. It's a rough guideline. Partly because Git is bad at dealing with binary files. There are plenty of cases when including generated files is appropriate. It has many advantages over not doing that - probably the biggest are * Code review is much easier because you can see the effect on the output. * It's easier to find the generated files because they're next to the rest of your code. IDEs like it much more too. In fact the upsides are so great and the downsides so minimal I would say it should be the default option as long as: * The generated files are not huge. * The generated files are always the same. Even when they are huge it might still be a good idea, but you can put the files in a submodule or LFS. I do that for a project that has a really difficult to install generator so users don't need to install it.
- deleted 3y ago[deleted]
- Cthulhu_ 3y agoI'm on the fence with this one. My previous project was Go & Typescript with a range of generated files; I committed the generated files, so that they would flag up in code reviews if they were changed, avoiding hidden or magic changes. I also didn't automatically regenerate, avoiding churn. That said, if the autogenerated output is stable, it's fine. After all, in a sense, compiling your code is also a kind of autogenerating and few people will advocate for keeping compiled code in git.
- haizzz 3y agoTo clarify here, "generated" and "autogenerated" were bad choices of words. They're translations created by humans and is dependent on the strings in code. See also: https://news.ycombinator.com/item?id=37318052 https://news.ycombinator.com/item?id=37318052