4 ms·
I work for GitLab and I was responsible for introducing release changes for this cycle. That is correct, we deploy our RCs to production. I consider RC as only
by marinj 9y ago
I work for GitLab and I was responsible for introducing release changes for this cycle.
That is correct, we deploy our RCs to production. I consider RC as only a point in time snapshot of a release that will be sent to public.
One thing that is incorrect is that we push all our RCs to production. For example, this release RC1 did not get to GitLab.com because during deployment to our other environments we found an issue that could have caused a large problem at scale. However, we executed a lot of QA and FA tests against all our RCs. You can see this here https://gitlab.com/gitlab-org/release/tasks/issues?scope=all&utf8=%E2%9C%93&state=closed&label_name[]=QA%20task https://gitlab.com/gitlab-org/release/tasks/issues?scope=all... .
If we only deployed a final release, we would most likely be overwhelmed with the amount of changes GitLab receives and it would be
hard for us to monitor and limit the impact.
We are working on getting to continuous deployment to GitLab.com, and as part of this push we started right now with the tools we have at our disposal. One of the ideas to get to stable, non impactful deploys was to create many RCs. Thinking was that if we have a smaller delta between the RCs, we can more effectively check the changes and reduce the overall impact. We also wanted to make sure that we over-communicate our status updates so that we can get feedback in case we overlooked something.
Was it problem free? Nope. Did it limit the impact to users? I firmly believe so.
We have a lot of work to do to get to blue-green deployments and continuous deployment on GitLab.com scale while we also continue to ship to on premise customers. I do believe that we will get there the best way we know how to, and that is iterating on changes.
- lucideer 9y ago> One thing that is incorrect is that we push all our RCs to production I hadn't recorded them or looked it up, hence the use of the word "seem", but this is at least good to hear. > Was it problem free? Nope. Did it limit the impact to users? I firmly believe so. You may be right. It might have been worse, and perhaps releasing more frequent low-impact changes helped, but I'm somewhat unconvinced. Orchestration was known not to be stable; given that, I would typically try and minimise deployments until orchestration had been more comprehensively tested and/or staging/canary environments had better parity with prod. Additionally, the deployments were done in the middle of work days (as has been mentioned elsewhere), at highly inconvenient times (just before Christmas!), with little to no warning to users, hinting that bad internal planning must have played some part. When queried, the reply was literally that there is no schedule[0] . This response does not instill confidence. > I do believe that we will get there the best way we know how to, and that is iterating on changes. I really hope this is true—I'm still using Gitlab myself—but while I was an advocate 6 months ago, I've very much put such advocacy on hold for the moment. Also, just to mention, I really appreciate there are GL employees on HN commenting on these threads. I've received a reply before elsewhere (on perf. and Gitaly) and it's been enlightening and informative. I really do like the transparency in this organisation, and have done my best to be as patient as possible with the stability/perf. issues up until now. But it's really become a bit ridiculous at this stage. [0] https://twitter.com/gitlabstatus/status/943574131425140736 https://twitter.com/gitlabstatus/status/943574131425140736
- slrz 9y ago> Additionally, the deployments were done in the middle of work days (as has been mentioned elsewhere), at highly inconvenient times (just before Christmas!) It's always the middle of a work day in some part of the world. It's always going to be inconvenient for someone. Do these updates really cause significant outages or just slower response times due to invalidated caches or whatever? The dogfooding-preleases-on-Gitlab.com idea in general seems perfectly reasonable to me. They're the ones best equipped to deal with any resulting issues so catching them before the majority of on-premise users is going to hit them seems like a very good idea. If they wouldn't do it, do you think the possible issues would magically disappear until GA? How?