4 ms·
Generating code is a powerful tool that I think everyone should have in their toolbox — use it wisely of course, but don’t avoid it on principle. This describe
by codeflo 5y ago
Generating code is a powerful tool that I think everyone should have in their toolbox — use it wisely of course, but don’t avoid it on principle.
This describes an interesting approach that I hadn’t seen. I personally wouldn’t have considered putting the result in the same file because I usually like to gitignore any and all generated code, mostly to reduce phantom merge conflicts. However, some colleagues have the opinion that those conflicts are trivially resolved by rebuilding and thus don’t matter, and having diffs of the generated code show up in pull requests adds value.
What do others think, check in generated code or put it in gitignore?
- maccard 5y ago> However, some colleagues have the opinion that those conflicts are trivially resolved by rebuilding and thus don’t matter This is how mistakes happen, and when dealing with things like serialisation is an absolute _minefield_ > nd having diffs of the generated code show up in pull requests adds value. They do have a point here, but is there a reason that the generated code diff couldn't be done separately to being commited? > What do others think, check in generated code or put it in gitignore? Always ignored, no questions asked. Generated code is a build artifact, the same way an object file or a jar is. It just happens to be human readable. Your build system should be 100% bulletproof and pick up changes to your input files and correctly regenerate the output files for you, anything less should be considered a p0 bugfix. If you need to cache these files for whatever reason, do so at a different level (sccache/container/whatever).
- twic 5y agoIf you're going to check it in, at least have a CI job which rebuilds it from scratch, then fails if there are any changes. I'm not persuaded of the value of having diffs in source control, but i haven't looked at concrete examples. This would be an interesting post for someone to write. For performance, i'd rather follow a distributed caching approach. Hash the true source, look for /some/nfs/mount/or/s3/path/${hash}.tar.gz, if it exists, unpack it into the generated source directory, if not, generate the source, tarball it, and upload that to the said path. Takes some care to hash the right things (needs to include the version and configuration of the generation tool, etc). Again, you could have a CI job checking that the caches are correct if you're worried.
- dilap 5y agoI like checking it in; on balance, I think it makes life nicer: Simpler to get started running the code, less compute wasted running generators (this is especially nice when doing something like bisect), and it's nice to be able to just see all your code, right there in the repo. I actually think the biggest downside is the diffs -- often there's no easy way (e.g. in GitHub) to hide all generated code, which can be annoying. Merge conflicts don't bother me at all, since they're so easy to resolve; I agree with your colleagues there. (Though that would be trickier if you had files mixing generated and hand-written code! Contra the original article, my instincts/advice would be to always keep generated code well-isolated from hand written code, ideally using some standard convention [e.g., all generated stuff is in /generated/ path].)
- HelloNurse 5y agoKeeping generated code in a segregated source code repository, with the advantage of inspecting diffs without file pollution, should be technically easy. Commands to take snapshots of every build ("git add *", "git commit", "git rev-parse" on the actual source repository, "date" etc.) are easy to script.