8 ms·
CircleCI Down
- oxff 4y ago> Update - We are investigating multiple possible causes, including database changes and code changes. Sounds like they haven't got the first clue about what is causing it.
- ismayilzadan 4y agoIt has been going on for almost 6 hours now. It does feel like they haven't got a clue.
- onion2k 4y agoI don't think you can say that. They might know what happened, but it could still be hard to recover from. Catastrophes happen. If you deploy something that has a destructive migration you can't easily roll back without reverting to a backup and there's a problem that you've not seen in QA then you're in for a bad time. This is compounded if you also discover your backup process hasn't worked properly for a while. If that happens you're facing some serious downtime, and the dilemma of either trying to fix the problem, or trying to rollback to the last working backup. There's a good reason why grumpy old devs like me insist on writing docs, having playbooks, testing everything including non-code stuff, and we still fear major deploys. I have scars from exactly those sorts of disasters. Hopefully the devs at Circle get past this with as little stress as possible, and they learn from what went wrong.
- ismayilzadan 4y agoYou are totally right, catastrophes happen and I also wish that they get past this with as little stress as possible. The whole reason of my assumption was the lack of description in their updates for the incident that is going on for 6 hours. Maybe little more detail would give me a hint that everything is under control, but I didn't feel that when I read their updates.
- a1445c8b 4y agoIt’s probably just because they had the choice of either focusing all their energy on fixing the problem asap or setting aside some of it to write a more detailed description that’s also fit for public consumption. Given the severity, they probably chose the former since whatever descriptive, reassuring description they put out there isn’t going to be actionable anyway.
- z00b 4y agoHi folks. As the CircleCI CTO, I appreciate your patience here and all the feedback. It's true that we are focused on getting customers moving again over sharing more detailed information, but will aim to do better in providing a bit more in our updates. status.circleci.com provides real time updates for both how we're tackling outages and more detailed incident reports. We will post more information there about this incident once we are on the other side and have comprehensive detail.
- plumefar 4y agoWhen facing such large scale issues, communicating properly is very hard: Several teams might be investigating several possible root causes in parallel, and you might change your mind over time as to what is the most probable root cause. So you might end up communicating something ("we think it comes from X, we're fixing it that way"), just to find yourself changing you mind a few minutes later. Changing your message is usually not well perceived, even though that's actually normal during an investigation. I would not like to be in charge of the communication. Finding the balance between saying too much or too little is tricky.
- brightball 4y agoI'm not all that surprised. A friend saw a phishing email that was imitating them because they lacked a DMARC record. Sent them explicit instructions on how to fix it by adding a DMARC policy and all they did was create a p=none record that doesn't prevent direct imitation. That's definitely the first step, but eventually you need to turn it up to p=quarantine for it to do you any good and it's been a while (several weeks). Shouldn't have needed a random user to point it out in the first place. I just don't have a tremendous amount of confidence that they take their infrastructure seriously at this point.
- ShakataGaNai 4y agoSo they did the thing they were recommended but didn't take some further steps, on this one issue. Clearly that means they are totally incompetent? Even though the people dealing with DMARC issues are probably IT & Marketing, not the DevOps & Engineering people who are running the product.
- brightball 4y agoA p=none record is barely different from not having a record at all...and yes at this point a tech company without an enforced record is a major red flag. It's been a decade since the standard went public, it's required at the federal level already and in many EU countries it's being mandated for businesses in general. Most 3rd party senders today already insist that you setup DKIM as part of your setup process and if that happens, you're going to pass a DMARC check. It's hard to setup for older companies with thousands of servers in their own data centers that are each individually sending email. Cloud native companies sending their email through a few 3rd parties like Sendgrid/Postmark or a newsletter tool are EASY to setup. I'm mentioning this on a post about their infrastructure being down for 6 hours because yes, it's related. Email delivery for the primary domain is absolutely an IT, Engineering, Operations and Security problem, not a marketing problem. It goes directly to the application especially when one of the main facets of the application is to send emails about your repos and login credentials. Blame shifting it to the marketing department does not hold up. When multiple people are commenting on this post about just how frequently their outages are happening it shows a problem in the overall infrastructure mindset for it to continue. Maybe they know exactly what the problem is and somebody higher up is keeping them from fixing it in order to prioritize other things. Either way, for company that's supposed to be providing a core devops function to have outages that frequently as well as making it dead simple to spoof email that looks like it's coming straight from them...it's not a good look.
- PrimeDirective 4y agoLately, it's been down almost weekly. Not a fan of these types of services myself, but we do use it at work.
- folkrav 4y ago> Not a fan of these types of services myself What do you mean by "these types of services"?
- capableweb 4y agoCircleCI is a CI/CD solution (ala SaaS) that you don't host yourself. Many (myself included) prefer to host mission-critical services ourselves to avoid untimely downtime.
- folkrav 4y agoI know what CircleCI is. Just wanted to understand what part of it you were referring to, which seems to be SaaS in general. I'm honestly not really convinced about self-hosting really avoiding untimely downtime, but whatever works for you and your team. E.g. I've worked in a business who self-host their Gitlab instance, there was a non-negligible amount of work for backups/upgrades on top of troubleshooting performance issues once in a while, amongst other things in the same vein.
- capableweb 4y agoThe part about "untimely downtime" is about that CircleCI decides themselves when to push updates (which is the most common reason services has downtime), instead of you deciding when to upgrade/push updates. If you have a big migration/change coming up, you'd put pause on upgrading the CI/CD service as you don't want to muck with it while pushing out other organization-wide changes. Granted, self-hosting comes with it's own share of problems too, no solution is a silver-bullet without any issues, but being able to "freeze" things to a stable mode helps to stabilize other processes.
- 4y ago
- proxysna 4y agoMy monthly Circle CI downtime, yay.
- hnlmorg 4y agoI wish it was only monthly. We have a Slack alert whenever there is reported downtime and it goes off at least once a week. Though granted it's seldom as business impacting as this current outage. The problem CircleCI faces is that most hosted VCS now support CI/CD tools, as do most enterprise clouds. These will all have better integration with most peoples systems because they'll already be using the VCS or the public cloud (and if you're not using the public cloud you'd likely favor Jenkins / Concourse / etc over a cloud CI solution). So CircleCI's relevance is constantly being eaten at. The last thing they need it to damage their own reputation with these constant outages. I really hope CircleCI can turn around as it's great to see some competition but at this point in time I'm not feeling to optimistic about their long term future.
- kldx 4y agoIn my experience, GitHub actions is equally flaky
- frays 4y agoThis was the final straw. I'm switching over to Buildkite now.
- keithpitt 4y agoOh hai!
- carimura 4y agoI'm switching to X and will be happy until either, a) X gets large enough to matter and also has scale issues. b) X gets acquired by Y and sends a "what a great journey we're so excited that X will have Y's resources" email and then inevitably becomes just another forgotten tool under Y /s sorta
- x86_64Ubuntu 4y ago>...inevitably becomes just another forgotten tool under Y Not before Y cuts all development to X. It doesn't always happen, but if you see Embarcadero buying a piece of your tech stack, find an alternative, immediately.
- bpicolo 4y agoJenkins is the most flexible system I've worked with, but Buildkite has been great and just gets the job done reliably. Big fan.
- kawsper 4y agoWe've been with Buildkite since 2016, it's been a great experience, I really like the tool, and their team is amazing!
- mr337 4y agoWelcome to greener pastures :) Been on BK for about 6mo now and still loving it.
- Nextgrid 4y agoI wonder how much it's going to take before people realize that maybe a single server somewhere in the office running Jenkins isn't that bad of an idea after all. Unless you're Google, "scale" will inherently not be a problem, and risks of operator error can be reduced by scheduling maintenance at times where an accidental outage won't impact your business.
- xyst 4y agoin some cases that single server is bob's 2016 MBP because management is too cheap to buy a dedicated server
- justin_oaks 4y agoExtra points if management is too cheap to buy a dedicated server, but is willing to pay for 3rd party services.
- xcambar 4y agoTo operate properly at the scale of hundreds of engineers pushing changes, you absolutely need a cluster of machines and a team to operate it. It is significantly cheaper and more efficient to pay a third party.
- justin_oaks 4y agoI agree with the sentiment that people should evaluate whether or not they need an external service to run their builds. That said, there are a number of reasons to not use the Jenkins server in the office: 1) Someone on staff needs to maintain it. 2) A single hardware failure can cause significant downtime. 3) Your office internet service may have limited bandwidth and be a bottleneck for your build or artifact deployment 4) Having your server on-site may be considered a security risk. I'm not saying that a server in the office is a bad idea, I'm just saying that each business needs to consider the advantages and disadvantages. I'm sure there are those who could get by just fine with a server in the office.
- egwor 4y ago
- nerdjon 4y agoI still fail to see the heavily opinionated appeal of CircleCI over running a dockerized Jenkins instance (and agents) in AWS. (Or GitHub Actions or any other managed CI environment) We get all the customization I want and it scales just fine. But I am also still annoyed that when CircleCI announced their templates, they did not offer you the ability to have private templates (or something along those lines, it left a bad taste in my mouth and we moved to Jenkins a couple months later)
- tedmiston 4y agoI used CircleCI a long time ago, just before and after 2.0, and thought it was fine. These days GitHub Actions is awesome and full featured and does everything I want though without getting in the way. I doubt I'd use anything besides GHA unless I wanted to decouple from GitHub as a dependency. Even then, I wouldn't be surprised if someone else hasn't already come up with an open source or local runner for the GHA yaml files as a stopgap.
- watermelon0 4y agoTwo downsides of GHA that I could find: - you are limited to 2 core / 7 GB instance, whereas CircleCI offers up to 16 cores / 64 GB (for example, building software usually scales proportionately to machine size, and this can in theory be up to 8x faster on CircleCI) - no support for ARM instances
- tedmiston 4y agoThe details you listed sound accurate for the built-in GitHub-hosted runners. GHA also supports bringing your own self-hosted runners [1][2] where you install their agent, so you could, e.g., use an ARM instance on AWS with tons of cores. It looks like CircleCI offers this as well. [1]: https://docs.github.com/en/actions/hosting-your-own-runners/about-self-hosted-runners https://docs.github.com/en/actions/hosting-your-own-runners/... [2]: https://github.com/actions/runner https://github.com/actions/runner
- tauntz 4y ago
- shdh 4y agoHands down the worst CI/CD platform out there. GitLab CI #1 in my opinion.
- thinkindie 4y agoto all those saying "why you don't spin a GitLab CI instance" - we are a small team, we want to focus on shipping code that adds value to our customers, not maintaining something that has been largely commoditised.
- aftbit 4y agoCircleCI used to be an absolutely awesome way to set up CI. The default images just worked for the vast majority of cases. Sure, it always took a bit of fiddling with YAML, but IMO was way ahead of Jenkins et al when it came out. Then they reworked their YAML format to enable pipelines (and make everything way more complicated, and break all our existing flows). Then they reworked their images to make them more efficient or something, but now there is almost never an image that has everything I need, so I have to build my own image in docker (with yet another CircleCI job). Now that Github Actions exists, we have been slowly migrating everything off CircleCI. Too bad they lost the simplicity that brought us there in the first place.