16 ms·
I remember a Rich Hickey talk where he described Datomic, his database. He said "the problem with a database is that it's over there." By modeling data with imm
by MathMonkeyMan 1y ago
I remember a Rich Hickey talk where he described Datomic, his database. He said "the problem with a database is that it's over there." By modeling data with immutable "facts" (a la Prolog), much of the database logic can be moved closer to the application. In his case, with Clojure's data structures.
Maybe the the problem with CI is that it's over there. As soon as it stops being something that I could set up and run quickly on my laptop over and over, the frog is already boiled.
The comparison to build systems is apt. I can and occasionally do build the database that I work on locally on my laptop without any remote caching. It takes a very long time, but not too long, and it doesn't fail with the error "people who maintain this system haven't tried this."
The CI system, forget it.
Part of the problem, maybe the whole problem, is that we could get it all working and portable and optimized for non-blessed environments, but it still will only be expected to work over there, and so the frog keeps boiling.
I bet it's not an easy problem to solve. Today's grand unified solution might be tomorrow's legacy tar pit. But that's just software.
- AtlasBarfed 1y agoI want my build system to be totally declarative Oh the DSL doesn't support what I need it to do. Can I just have some templating or a little bit of places to put in custom scripts? Congratulations! You now have a turing complete system. And yes, per the article that means you can cryptocurrency mine. Ansible terraform Maven Gradle. Unfortunate fact is that these IT domains (builds and CI) are at a junction of two famous very slippery slopes. 1) configuration 2) workflows These two slippery slopes are famous for their demos of how clean and simple they are and how easy it is to do. Anything you need it to do. In the demo. And sure it might stay like that for a little bit. But inevitably.... Script soup
- lelanthran 1y agoAlternative take: CI is the successful monetization of Make-as-a-Service.
- lenkite 1y agoNo, you keep your build system declarative, but you support a clean plugin API that permits injection into the build lifecycle and allow configuring/invoking the plugin with your DSL.
- KronisLV 1y ago> Part of the problem, maybe the whole problem, is that we could get it all working and portable and optimized for non-blessed environments, but it still will only be expected to work over there, and so the frog keeps boiling. Build the software inside of containers (or VMs, I guess): a fresh environment for every build, any caches or previous build artefacts explicitly mounted. Then, have something like this, so those builds can also be done locally: https://docs.drone.io/quickstart/cli/ https://docs.drone.io/quickstart/cli/ Then you can stack as many turtles as you need - such as having build scripts that get executed as a part of your container build, having Maven or whatever else you need inside of there. It can be surprisingly sane: your CI server doing the equivalent of "docker build -t my_image ..." and then doing something with it, whereas during build time there's just a build.sh script inside.
- kqr 1y agoThis sounds a lot like "use Nix".
- justinrubek 1y agoUnfortunately, that's the last thing a lot of people want to hear, despite it saving a whole lot of heartache.
- zhengyi13 1y agoI mean, sure (also bazel I think), but I feel like that's because the learning curve for these tools to a first approximation looks a bit like the infamous EvE Online learning curve[0]. [0]: https://imgur.com/gallery/eve-online-learning-curve-jj16ThL https://imgur.com/gallery/eve-online-learning-curve-jj16ThL
- KronisLV 1y agoI mean, if it's easy enough to actually get your average developer to use it, then sure. In my experience, things that are too hard will just not be done, or at least not properly.
- dapperdrake 1y agoTransactions and a single consistent source of truth with stuff like observability and temporal ordering is centralized and therefore "over there" for almost every place you could be in. As long as communications have bounded speed (speed of light or whatever else) there will be event horizons. The point of a database is to track changes and therefore time centrally. Not because we want to, but because everything else has failed miserably. Even conflicting CRDT change merges and git merges can get really hairy really quickly. People reinvent databases about every 10 years. Hardware gets faster. Just enjoy the show.
- MathMonkeyMan 1y agoI haven't used Datomic, but you're right that the part that requires over there is "single consistent source of truth." There's only ever a single node that is sequencing all writes. Perhaps as a result of this, it provides strong [verified ACID guarantees][1]. What I got from Hickey's talk is that he wanted to design a system that resisted the urge to encode everything in a stored procedure and run it on the database server. [1]: https://jepsen.io/analyses/datomic-pro-1.0.7075 https://jepsen.io/analyses/datomic-pro-1.0.7075
- MortyWaves 1y agoIt’s why I’ve started making CI simply a script that I can run locally or on GitHub Actions etc. Then the CI just becomes a bit of yaml that runs my script.
- j4coh 1y agoAre you not worried about parallelisation in your case? Or have you solved that in another way (one big beefy build machine maybe?)
- MortyWaves 1y agoHonestly not really… sure it might not be as fast but the ability to know I can debug it and build it exactly the same way locally is worth the performance hit. It probably helps I don’t write C++, so builds are not a multi day event!
- maccard 1y agoHow does that script handle pushing to ghcr, or pulling an artifact from a previous stage for testing? In my experience these are the bits that fail all the time, and are the most important parts of CI once you go beyond it taking 20/30 seconds to build. A clean build in an ephemeral VM of my project would take about 6 hours on a 16 core machine with 64GB RAM.
- thechao 1y agoSheesh. I've got a multimillion line modern C++ protect that consists of a large number of dylibs and a few hundred delivered apps. A completely cache-free build is an only few minutes. Incremental and clean (cached) builds are seconds, or hundreds of milliseconds. It sounds like you've got hundreds of millions of lines of code! (Maybe a billion!?) How do you manage that?
- maccard 1y agoIt’s a few million lines of c++ combined with content pipelines. Shader compilation is expensive and the tooling is horrible. Our cached builds on CI are 20 minutes from submit to running on steam which is ok. We also build with MSVC so none of the normal ccache stuff works for us, which is super frustrating
- DrBazza 1y agoYour build should be this: build.bash <debug|release> and that's it (and that can even trigger a container build). I've spent far too much time debugging CI builds that work differently to a local build, and it's always because of extra nonsense added to the CI server somehow. I've yet to find a build in my industry that doesn't yield to this 'pattern'. Your environment setup should work equally on a local machine or a CI/CD server, or your devops teams has identically set it up on bare metal using Ansible or something.
- nrclark 1y agoAgreed with this sentiment, but with one minor modification: use a Makefile instead. Recipes are still chunks of shell, and they don’t need to produce or consume any files if you want to keep it all task-based. You get tab-completion, parallelism, a DAG, and the ability to start anywhere on the task graph that you want. It’s possible to do all of this with a pure shell script, but then you’re probably reimplementing some or all of the list above.
- gchamonlive 1y agoJust be aware of the "Makefile effect"[1] which can easily devolve into the Makefile also being "over there", far from the application, just because it's actually a patchwork of copy-paste targets stitched together. [1] https://news.ycombinator.com/item?id=42663231 https://news.ycombinator.com/item?id=42663231
- dgfitz 1y agoYou invoke CMake/qmake/configure/whatever from the bash script. I hate committing makefiles directly if it can be helped. You can still call make in the script after generating the makefile, and even pass the make target as an argument to the bash script if you want. That being said, if you’re passing more than 2-3 arguments to the build.sh you’re probably doing it wrong.
- nrclark 1y agoYes to calling CMake/etc. No to checking in generated Makefiles. But for your top-level “thing that calls CMake”, try writing a Makefile instead of a shell script. You’ll be surprised at how powerful it is. Make is a dark horse.
- reactordev 1y agoThe rule for CI/CD and DevOps in general is boil your entire build process down to one line: ./build.sh If you want to ship containers somewhere, do it in your build script where you check to see if you’re running in “CI”. No fancy pants workflow yamls to vendor lock yourself into whatever CI platform you’re using today, or tomorrow. Just checkout, build w/ params, point your coverage checker at it. This is also the same for onboarding new hires. They should be able to checkout, and build, no issues or caveats, setup for local environment. This ensures they are ready to PR by end of the day. (Fmr Director of DevOps for a Fortune 500)
- pxc 1y agoYou still inevitably need a bunch of CI platform-specific bullshit for determining "is this a pull request? which branch am I running on?", etc. Depending on what you're trying to do and what tools you're working with, you may need such logic both in an accursed YAML DSL and in your build script. And if you want your CI jobs to do things like report cute little statuses, integrate with your source forge's static analysis results viewer, or block PRs, you have to integrate with the forge at a deeper level. There aren't good tools today for translating between the environment variables or other things that various CI platforms expose, managing secrets (if you use CI to deploy things) that are exposed in platform-specific ways, etc. If all you're doing with CI is spitting out some binaries, sure, I guess. But if you actually ask developers what they want out of CI, it's typically more than that.
- michaelmior 1y agoA lot of CI platforms (such as GitHub) spit out a lot of environment variables automatically that can help you with the logic in your build script. If they don't, they should give you a way to set them. One approach is to keep the majority of the logic in your build script and just use the platform-specific stuff to configure the environment for the build script. Of course, as you mention, if you want to do things like comment on PRs or report detailed status information, you have to dig deeper.
- 1y ago
- layer8 1y agoYes, the build system should be independent from the platform that hosts it. Having GitHub or GitLab execute your build is fine, but you should as easily be able to execute it locally on your own infrastructure. The definition of the build or integration should be independent from that, and the software that ingests and executes such definitions shouldn’t be a proprietary SaaS.