10 ms·
Docker was unavailable in Ubuntu/Debian repos
- justinsaccount 10y agoTitle is misleading. The 'apt.dockerproject.org' host had a broken release file. This is not an Ubuntu or Debian maintained repository.
- vox_mollis 10y agoBroken, or compromised?
- cjbprime 10y agoProbably broken, since a competent attacker would have been able to avoid creating a checksum mismatch. My company's actually done the same thing before (same error), by putting Cloudfront in front of our APT repo -- it cached the main packages file inappropriately, causing the checksum mismatch.
- deleted 10y ago[deleted]
- poooogles 10y ago>Does this mean that Docker -- a major infrastructure company -- does not have any on-call engineers available to fix this? It appear to be that way. Reminds me when all of the reddit admins were stuck on a plane on the way back from a wedding [1]. Remember kids, improve your bus factor. http://highscalability.com/blog/2013/8/26/reddit-lessons-learned-from-mistakes-made-scaling-to-1-billi.html http://highscalability.com/blog/2013/8/26/reddit-lessons-lea...
- StavrosK 10y agoAnd don't put everyone on the same bus, I guess.
- sleepychu 10y agoActually since the aim is to avoid being hit by a bus, inside a bus seems like a pretty good place to put everyone.
- deleted 10y ago[deleted]
- mpnordland 10y agoAssume N is the total number of buses. Where N=1 your solution works. For N>1 the engineers can still be hit by a bus. This also ignores all other vehicles such as garbage trucks, semis and stealth bombers, all of which can take out your bus.
- Sanddancer 10y agoYou want to avoid sticking everyone on the same (air)bus too. https://en.wikipedia.org/wiki/Pacific_Southwest_Airlines_Flight_1771#Aftermath https://en.wikipedia.org/wiki/Pacific_Southwest_Airlines_Fli...
- FireBeyond 10y agoAlso called (by me) as "Baldrick's Bullet": https://www.youtube.com/watch?v=pKRxX3s3JlM https://www.youtube.com/watch?v=pKRxX3s3JlM Captain Blackadder: Baldrick, what are you doing out there? Private Baldrick: I'm carving something on a bullet, sir. Captain Blackadder: What are you craving? Private Baldrick: I'm carving "Baldrick", sir. Captain Blackadder: Why? Private Baldrick: It's part of a cunning plan, sir. Captain Blackadder: Of course it is. Private Baldrick: You know how they say that somewhere there's a bullet with your name on it? Captain Blackadder: Yes? Private Baldrick: Well I thought that if I owned the bullet with my name on it, I'll never get hit by it. Cause I'll never shoot myself... Captain Blackadder: Oh, shame! Private Baldrick: And the chances of there being two bullets with my name on it are very small indeed. Captain Blackadder: Yes, it's not the only thing that is "very small indeed". Your brain for example- is brain's so minute, Baldrick, that if a hungry cannibal cracked your head open, there wouldn't be enough to cover a small water biscuit.
- 10y ago
- Bromskloss 10y ago> all of the reddit admins were stuck on a plane on the way back from a wedding Clearly, we need to operate in cells, so that no one knows everybody and will have everybody over for weddings and other parties.
- bitJericho 10y agoHaha. Well its not uncommon to avoid putting an entire team on one plane. Pretty risky for a company to do that!
- madgar 10y agoYeah... the entire point of an on-call rotation is to specify who is available for incident response... if everybody is on the same plane at once then by definition nobody is on-call.
- bitJericho 10y agoAnd in the extremely unlikely event of a plane crash, the company goes belly up.
- oldmanhorton 10y agoThere was a comment a bit below that suggested that those who paid for commercial support got it 24/7, but if that's true, id imagine the fix that commercial support would have given to paying customers would have fixed it for everyone else too...
- shykes 10y agoDisclaimer: I work at Docker. I believe commercial releases are downloaded from a separate infrastructure (to be confirmed). Either way, the availability of Docker packages, free or commercial, is critical infrastructure and we should treat it as such. IMO our primary infrastructure team should have been involved, and someone should be on call for this. We'll do a post-mortem, find the root cause, and take corrective action as needed. Apologies for the inconvenience.
- voltagex_ 10y agoYou may want to lock that GitHub thread soon, it's getting argumentative and not very helpful.
- shykes 10y ago> You may want to lock that GitHub thread soon, it's getting argumentative and not very helpful. In open-source we call that "thursday" :)
- zjaffee 10y agoWhat ever happened to community over code?
- shykes 10y agoPart of a healthy community is accepting that people disagree a lot, have different values, and communicate their ideas in very different ways. Where we draw the line is if people are being intimidated, bullied, insulted, or anything that even remotely resembles harassment. Although I personally feel that some of the comments in that thread are pretty unfair and poorly informed, they don't seem to violate the social contract.
- Rapzid 10y agoI don't think this would surprise anyone that has used Docker Hub in their CD pipeline. So many reliability issues. We moved off it last year. When we went to cancel our subscription the other month downgrades/cancellations were broken on the site as a known issue; had to open a support ticket. Most of the UI issues were still present along with some new ones.
- kfrz 10y agoThere are two tiers to Docker's support, for certain, but as pointed out on the Github issue by a Docker team member here (https://github.com/docker/docker/issues/23203#issuecomment-223326996 https://github.com/docker/docker/issues/23203#issuecomment-2...), there's a definite urgency sensed by the team.
- tsuresh 10y agoFrom the GitHub issue thread, I see a lot of people being angry for their production deployments failing. If you directly point to an external repo in your production environment deployments, you better not be surprised when it goes down. Because shit always happens. If you want your deployments to be independent of the outside world, design them that way!
- agentgt 10y agoAfter reading the bug report I am surprised so many people are using a remote package repositories for their build machines builds... then again I'm not too surprised I guess. I'm not that familiar with Docker but I am of package/dep management (from deb, jars, npm, eggs etc) and you most certainly want to use a mirrored package repository (jfrog, sonatype, or whatever) for this reason and many more other reasons (bandwidth, security, control, etc). So if you did have issues with the outage I would look into one of those mirroring tools. At the minimum it will speed up your builds.
- karterk 10y agoMaybe this is the norm in big enterprises, but I have not actually come across any company which hosts a local package repository for commonly available packages.
- LTheobald 10y agoReally? We certainly do here. Linux packages are mirrored. Programming dependencies are mirrored via tools like Nexus or Artifactory. Like tsuresh said - stuff happens. What if you internet connection went down for a long period of time. You couldn't continue working. It takes very little to setup, gives you fall over but also makes installing dependencies sooo much faster.
- draw_down 10y agoThey do it where I work, because our ops and dev teams like it when deploys don't randomly break.
- 10y ago
- 0x0 10y agoIt's scary how most people in that thread seem to be more concerned about forcing an installation, rather than pause and consider why the hashes might be wrong and why it might not be a good idea to install debs with incorrect hashes. If the apt repo was compromised (but the signing keys were not), this is very likely exactly the symptom that would appear.
- cjbprime 10y ago> If the apt repo was compromised (but the signing keys were not), this is very likely exactly the symptom that would appear. I don't think that's correct. It would pass a checksum test and fail a signature test with a "W: GPG Error". The checksum test is not about cryptographic security, it's just about files referenced by the Packages file having the same hash that the Packages file declares them to have. You don't need any signing keys to make that happen.
- 0x0 10y agoWhat's more suspicious: Bad hashes or bad signatures? What would an attacker choose if their goal was to get as many people as possible to force install?
- cjbprime 10y agoIt's impossible to force install the packages when they have bad hashes (hence the severe breakage here), and it is possible to install the packages when they have bad signatures if you didn't import the gpg key or don't run with signature checking. So I'd guess a rational attacker would choose a bad signature. But attackers can be irrational; it doesn't prove it's not an attack. Just not my intuition.
- 0x0 10y agoThat's interesting, I'm assuming you're talking about apt now. I don't think dpkg checks signatures if you install straight from a .deb. :)
- shykes 10y agoHi, I work at Docker. Here is my reply on the github thread: https://github.com/docker/docker/issues/23203#issuecomment-223326996 https://github.com/docker/docker/issues/23203#issuecomment-2... I am copying it below: <<< Hi everyone. I work at Docker. First, my apologies for the outage. I consider our package infrastructure as critical infrastructure, both for the free and commercial versions of Docker. It's true that we offer better support for the commercial version (it's one if its features), but that should not apply to fundamental things like being able to download your packages. The team is working on the issue and will continue to give updates here. We are taking this seriously. Some of you pointed out that the response time and use of communication channels seem inadequate, for example the @dockerststus bot has not mentioned the issue when it was detected. I share the opinion but I don't know the full story yet; the post-mortem will tell us for sure what went wrong. At the moment the team is focusing on fixing the issue and I don't want to distract them from that. Once the post-mortem identifies what went wrong, we will take appropriate corrective action. I suspect part of it will be better coordination between core engineers and infrastructure engineers (2 distinct groups within Docker). Thanks and sorry again for the inconvenience. >>>
- de0x 10y agoWhy don't you "do devops"?
- NetStrikeForce 10y agoMaybe you are being downvoted because your comment is too short, but it's something that crossed my mind when reading the end of the parent comment. The best thing is, this might end up being the best proof of why you need to embrace devops methodologies and maybe take advantage of tools like Docker while doing so :)
- Sanddancer 10y agoThis isn't doing devops, this is release management, which is something that traditional sysadmins do all the time. Making sure that the image file and repository information match up is pretty basic, and that a release is deployed correctly is pretty basic. A project like this, I'm surprised they don't have tools like nagios constantly checking to make sure downloads are working and that checksums, etc all match up, preferably on the servers before whatever load balancing system you have points at them. Deployment can, and should, be as atomic and possible, regardless of who is pushing the go button.
- ajarmst 10y agoI often wonder why the community's response to issues with an open/free/community package is to give the maintainers a strong argument to discontinue it in favour of a commercial one, or just abandon it altogether.
- ajarmst 10y ago"Why I Haven't Fixed Your Issue" --- http://www.brycematheson.io/post/why-i-havent-fixed-your-issue/ http://www.brycematheson.io/post/why-i-havent-fixed-your-iss...
- therealmarv 10y agoI think this is a combination of chains which are dependent on eachother, especially when you use Travis CI: 1st) apt-get not flexible enough to ignore that error on apt-get update 2nd) Travis CI having so much external stuff installed, it's a big big image which has more failure points 3rd) Docker repo failed.
- nativityscenes 10y agoHey, I work for Docker too. Our quick 3 hour fix on this issue was quite impressive, I hope the community feels the same. We do the cloud native devops fixes faster than anyone, hands down! I personally haven't slept in days in anticipation of being able to resolve such a high-profile issue, and reap the glory and fame that come with such as fix. Hopefully more people purchase our commercial product as a result of this outage!
- dang 10y agoPlease stop.
- deleted 10y ago[deleted]
- mapleoin 10y agoThis is a really bad title. There is nothing wrong with either Ubuntu's or Debian's repositories. The problem is with Docker's repositories of Ubuntu/Debian packages.
- perlgeek 10y agoOutages or mis-configurations can happen to pretty much any source of packages you use, be it debian, pypi, npm, bower or maven repositories, or source control. Anybody remember left-pad? So as soon as you depend heavily on external sources, you should start to think about maintaining your own mirror. Software like pulp and nexus are pretty versatile, and give you a good amount of control over your upstream sources.
- smegel 10y agoSometimes paying for RHEL isn't a bad thing.
- mschuster91 10y ago... which is why the clever sysop mirrors his packages and tests if an update goes OK before updating the mirror. If you're running more than three machines or regularly (re)deploy VMs, it is a sign of civilization to use your mirror instead of putting your load on (often) donated resources. It's the same stupid attitude of "hey let's outsource dependency hosting" that has led to the leftpad NPM desaster and will lead to countless more such desasters in the future. People, mirror your dependencies locally, archive their old versions and always test what happens if the outside Internet breaks down. If your software fails to build when the NOCs uplink goes down, you've screwed up.
- brazzledazzle 10y agoI'm a bit disappointed that people are willing to make public criticisms of Docker when it's their builds that are failing. They made the decision to depend on a resource that could be unavailable for a large number of reasons entirely unrelated to Docker or their infrastructure. Just like the node builds that failed this should cause you to rethink how you mirror or cache remote resources not prompt you to complain about your broken builds on a github issue page. There may be things you'll never be able to fully mirror or cache (or could just be entirely impractical) but an apt repository is definitely not one of them.
- willejs 10y ago+1 !