11 ms·
Write Gitlab CI Pipelines in Python Code
- nhoughto 5y agoReminds me of the original atlassian bamboo pipeline as code impl that had you write the pipeline as Java or Groovy, can’t remember specifically. They got dunked on for it not being a declarative file. Now we’ve got a pipeline being defined by a declarative file being generated by code.
- delusional 5y agoWe use bamboo at my job, and it's java that generates a yaml file that is then shoved into bamboo. It's a completely shitshow. why would i use a turing complete language to generate a declarative yaml file containing turing complete shell scripts? Why not just write the shell script in the first place?
- lyjackal 5y agoBecause the shell scripts are only run from _within_ pipeline steps. Gitlab CI yaml has no way of conditionally creating the pipeline steps, or the step's contents.
- DrSarez 5y agoThat is interesting. In fact I've googled for it and found a blog post with really nastay Java code defining a pipeline. Of course, in code you can do nasty things and create insane pipelines. Thus I think it is good having both - the declarative base and a generator on top. The declarative base ensures having an easy way to get into pipeline mechanics. However if you plan to write really complex pipelines, the pipeline-as-code approach jumbs in. At this point your'e firm with the basics and know what your'e doing in code. Amazon is doing the same with the Cloud Development Kit (CDK). With CDK you can code your infrastructure in a number of languages (java, csharp, .net, python, typescript), which was synthesized into cloud formation to be finally deployed. For smaller projects and teams not firm with one of those languages, plain CFN may be much better. However after learning CDK you won't create any infrastructure without it.
- takeda 5y agoTo me it reminds me of troposphere[1]. It's similarly using an imperative language (also python) to generate a declarative file (CloudFormation). [1] https://github.com/cloudtools/troposphere https://github.com/cloudtools/troposphere
- mdaniel 5y agoJetBrains took that approach in Space, also, only choosing Kotlin in order to side-step(?) the "groovysayswhat?" ambiguity and anti-discoverability of Jenkinsfile
- octopoc 5y agoWhat I would really like to see is a CI system that lets me write a script in a language of my choice instead of defining a pipeline config file. That way I can run the pipeline locally, put breakpoints in, etc. Nuke [1] gets close but there are still a lot of tasks that don't have C# bindings, such as publishing build artifacts and uploading test results. While I'm dreaming about my perfect CI, I'd also like the ability to download benchmark results from previous commits so that I can generate trend graphs and publish them in the build results. To do this right the CI system would have to have an API using REST, GraphQL, gRPC or some such API format that generates clients in many languages. That way they don't have to maintain bindings in every language. [1] https://nuke.build/ https://nuke.build/
- gravypod 5y agoI think the main issue is that most build systems are not expressive enough to do complex packaging, code gen, etc that most projects need. I've been using Bazel for a few projects/teams and it's worked really well. Once your build system is as expressive as you need you get a lot of freedom. Bazel also let's you execute completely local builds and remote builds triggered locally. If you set this up debugging CI issues is amazingly simple.
- octopoc 5y agoTotally agree. I've never used Bazel; it sounds a lot like Nuke. I had a coworker whose job was to maintain the CI. His commit messages would look like this sometimes: > Fix CI issue with blah blah blah > Hmm that didn't work lets try something from stackoverflow > Build fix > Build fix please work > Please > I hate my life > I am tired and hungry, I want to go home > Stupid yaml The CTO would call him like a week later to make sure he was okay. This happened often before we switched to Nuke. It rarely happened afterwards. It's awesome being able to debug your CI locally.
- gravypod 5y agoYep, and if you get everything working smoothly enough developers can use the same very polished tooling for their own workflow.
- nerdbaggy 5y agoGitlab can load its ci file from a remote web server. What I have done is generate the yaml programmatically and this gets returned from the web server. Less than ideal but it at least allows dynamic creation
- StavrosK 5y agoYou can just have a CI job generate a CI config file and have GitLab CI load that into a child pipeline. Look up dynamic pipelines.
- lyjackal 5y agoYa, as I understand, that's all this python ci lib is doing. I find it a little annoying that your first have to start a worker to create the pipeline file, and then again to run the pipelines steps
- remram 5y agoIn this day and age it's not that weird. GitHub's own staff are spawning CI jobs just to add a label to a each new issue https://github.blog/2021-04-28-use-github-actions-manage-docs/ https://github.blog/2021-04-28-use-github-actions-manage-doc...
- nerdbaggy 5y agoI feel like there was some weird edge cases I was having with child pipelines. Maybe it was something with artifacts, don’t remember.
- VOBOSHI 5y agoArtifacts, passing environment variables from the main pipeline to the child pipeline, expanding variable references in the child pipeline and getting unexpected results... We're migrating our shared pipeline away from child pipelines. Maybe one day we'll do a write-up. :)
- dnsmichi 5y ago
- zie 5y agoThis is cool, but I'd rather my CI jobs are just k8s or nomad or systemd jobs. Then the code we deploy is the same as the code we use to build, and it doesn't matter WHAT you do in build land, go nuts and do whatever you want. CI systems will either grow into general purpose code runners or they will wither and die.
- kmstout 5y agoSeveral weeks ago I poked around Github's CI workflow references. It seems that there's a clear path for crafting a workflow that simply grabs the text of a each issue (as it's submitted) and appends it to the documentation before closing the issue, perhaps marking it as "expected behavior". In short, an evening of tinkering could automatically turn all of a project's bugs into features for the foreseeable future. This would represent a quantum leap in software quality assurance.
- sytse 5y agoAt GitLab we take the 'Release notes' section of each issue and use it to automatically generate the release post. For an example of a release post see https://about.gitlab.com/releases/2021/04/22/gitlab-13-11-released/ https://about.gitlab.com/releases/2021/04/22/gitlab-13-11-re...
- zmmmmm 5y agoAnother step in the endless cycle of configuration vs code. It's not an accident that we are in a deep cycle of constrained configuration languages (yaml/json/etc etc). People chose to go there because before that we had a cycle of using programming languages and people hated it, for all sorts of good reasons. Now I see we are on the way back into adding wrappers around the static config files, to turn them back into programming languages. This is happening all over the place - because guess what, it turns out, people hate static config files too, for all sorts of good reasons. I am not sure if we will ever reach a compromise here or if we are just going to have to put up with endless change and churn because nobody is ever happy or at least willing to just settle for things that are "ok" but not "perfect".
- mixologic 5y agoI think the main issue with using programming languages to run your CI/CD is that the whole point of CI/CD is to test and verify that your code works. Now you have code that may or may not be buggy, running your CI/CD so you'll have to set up CI/CD to test your test running code, ad nauseum. People might hate constrained configuration, but the idea is that you dont have to prove that what you're configuring is actually whats happening.
- ViVr 5y ago> Now you have code that may or may not be buggy Doesn't your solution of replacing the possibly buggy code with configuration leave you with configuration that may or may not be buggy?
- zmmmmm 5y agoIts mainly a question of it being deterministic and reproducible I think. The presumption is that you then go and test a bunch of stuff based on whatever was built, so no, it will not be buggy. But if the code that actually goes to production is actually different to what is tested (because your rebuilt it based on a config with some runtime behavior that executed differently), you are compromising the whole thing and all bets are off. People will try to split the difference and generate a static config with code and test that as a reproducible entity, but then you have another set of tradeoffs (not least, complexity).
- remram 5y agoI think if we're going that way, it would be nice to have a single library that can generate configurations for multiple CI systems (GitLab CI, GitHub Actions, Jenkins, CircleCI, etc)
- LennyWhiteJr 5y agoIt's almost like we've turned yaml into some sort of primitive assembly file that must be compiled by a higher-order language. What was once created for human readability, has now turned into a machine generated and machine parsed format.
- ORioN63 5y agoIt was created for human readability, but always with machine interop in mind, otherwise English/<NatLang>, would prob. do just fine in most cases, especially if sprinkled with some markup.
- aiisjustanif 5y ago> “We use bamboo at my job, and it's java that generates a yaml file that is then shoved into bamboo. It's a completely shitshow.” > “What I have done is generate the yaml programmatically and this gets returned from the web server. Less than ideal but it at least allows dynamic creation” They amount of polarization in this comment section is making me dizzy.
- freetime2 5y ago> We started this project because of Gitlab CI yaml files growing over thousands of lines. Is this typical? All of our Gitlab CI files are well under 100 lines. What sorts of things are these pipelines doing that require so much configuration? Our CI steps are basically: * Build * Run some static analysis * Test * Publish build artifacts With each step taking only a few lines. Most of the “heavy lifting” is managed by other tools like npm, or some scripts we have checked into the project, and our CI process just kicks off those steps.
- globular-toast 5y agoMy largest one is 200 lines. But yeah, most of the heavy lifting should be done by other tools. The most common "script" for a build step should be simply "make". For Python, almost everything is run through tox, so I have "tox" for tests, "tox -e wheel" for the packaging etc.
- orf 5y agoIf you've got thousands of repositories deploying services written in different languages to many different environments and multiple clusters within each environment then your CI files can get quite lengthy. Using Docker helps standardize things, but our shared pipeline repo is still 3.8k lines of YAML.
- kissgyorgy 5y agoIt seems like a lot of people don't understand that having a DECLARATIVE language/configuration/whatever such a huge advantage it's insane. It's so easy to write, you can never make logical mistakes, avoid all kind of bugs, it's just easy to learn and you can just write the desired result. I never understood these project where somebody makes a very easy declarative thing and makes it imperative. Programming in Python is order of magnitude "harder" and error prone than writing a simple YAML file.
- BiteCode_dev 5y ago> Programming in Python is order of magnitude "harder" and error prone than writing a simple YAML file. For a simple config, sure. But as soon as you have something more complicated, you end up with: - a frankenstein monster DSL, with pseudo flow control + funcs backed in through a patched up templating system - weird error messages, no stack trace, and type errors that accumulates - no DRY - no tooling: type checks, linting, de bugger, logger. Forget about it. Case in point: ansible playbooks. So unless your config is going to stay under 20 lines, just use a real programming language. If you don't want to use Python, fine. Use dhall, nix, jsonet or something else.
- erinnh 5y agoYou have linting and logger in Ansible. Type checks, kind of depend. Debugging is definitely something I also would like to see something better coming out of Ansible than what is currently available.
- dvdkon 5y agoSlightly relevant: I made a hack to write Ansible playbooks in Python [1]. It doesn't fix everything, for example runtime control flow still has to be described as attributes of tasks, but it does help with parametrising. It's a question whether it's worth it to add such a hack or if I should just make do with YAML. [0]: https://gitlab.com/-/snippets/2112382 https://gitlab.com/-/snippets/2112382
- BiteCode_dev 5y ago
- CSDude 5y agoI have made something similar for Tekton (Kubernetes) with Jsonnet, https://mustafaakin.dev/posts/2020-04-26-using-jsonnet-to-generate-dynamic-tekton-pipelines-in-kubernetes/ https://mustafaakin.dev/posts/2020-04-26-using-jsonnet-to-ge... but the eternal cycle of configuration vs code is always bothering me in every few years and make me question myself.
- deleted 5y ago[deleted]