9 ms·
Docker Systems Status: Full Service Disruption
- danvesma 1y ago...well this explains a lot about how my morning is going...
- atymic 1y agoResult of AWS outage https://news.ycombinator.com/item?id=45640754 https://news.ycombinator.com/item?id=45640754
- reader_1000 1y ago> We have identified the underlying issue with one of our cloud service providers. Isn't it everyone using multiple cloud providers nowadays? Why are they affected by single cloud provider outage?
- postexitus 1y agoNot only they are not using multiple cloud providers, they are not using multiple cloud locations.
- lvncelot 1y agoI think more often than not, companies are using a single cloud provider, and even when multiple are used, it's either different projects with different legacy decisions or a conscious migration. True multi-tenancy is not only very rare, it's an absolute pain to manage as soon as people start using any vendor-specific functionality.
- dijit 1y ago> as soon as people start using any vendor-specific functionality It's also true in circumstances where things have the same name but act differently. You'd be forgiven for believing that AWS IAM and GCP IAM are the same thing for example, but in GCP an IAM Role is simply a list of permissions that you can attach to an identity. In AWS an IAM Role is the identity itself. Other examples; if you're coming from GCP, you'd be forgiven for thinking that Networks are regional in AWS, which will be annoying to fix later when you realise you need to create peering connections. Oh and while default firewall rules are stateful on both, if you dive into more advanced network security, the way rules are applied and processed can have subtle differences. The inherent global nature of the GCP VPC means firewall rules, by default, apply across all regions within that VPC, which requires a different mindset than AWS where rules are scoped more tightly to the region/subnet. There's like, hundreds of these little details.
- DiggyJohnson 1y agoSounds like we’ve walked a similar path on this. Especially with IAM and network policies. > There’s like hundreds of these little issues Exactly. If it is a handful of things that is fine. It’s often as you describe.
- OtherShrezzing 1y agoI think there's some irony in Docker being impacted specifically, as they're one of the main tools to help achieve true multi-tenancy.
- DiggyJohnson 1y agoDepends on if you’re using Docker or Podman Desktop versus straight Docker/Podman and where you’re pulling your images from.
- ikiris 1y agoMulti cloud is just a way to have the outages of both.
- brookst 1y agoAnd even if you think it’s important enough to justify the expense and complexity, it’s times like this when you discover some minor utility service 1) is a critical dependency, and 2) is not multi-cloud. Complex systems are hard.
- rcxdude 1y agoBecause it's hard enough to distribute a service across multiple machines in the same DC, let alone across multiple DCs and multiple providers.
- nobleach 1y agoLooking at the landscape around me, no. Everyone is in crisis cost-cutting, "gotta show that same growth the C-suite saw during Covid" mode. So being multi-provider, and even in some cases, being multi-regional, is now off the table. It's sad because the product really suffers. But hey, "growth".
- madisp 1y agothey are using multiple cloud providers, but judging by the cloudflare r2 outage affecting them earlier this year I guess all of them are on the critical path?
- roywiggins 1y agoYou can be multi-cloud in the sense that you aren't dependent on any single provider, or in the sense that you are dependent on all of them.
- hunter2_ 1y agoA bit like the ambiguity of search facets: if I select one facet, I get results that match, but if I add a second facet, should the results expand (OR'ing my selections) or contract (AND'ing my selections)? Presumably they should be OR'd if they belong to the same category (like selecting multiple colors, if any given result has only one color) but AND'd otherwise (like selecting a color and a size). But then a category could consist of miscellaneous features, and I want results that have every feature I've selected, which goes against the general case.
- oldpersonintx2 1y ago[dead]
- jelder 1y agoNo, that's pretty rare, and generally means you can't count on any features more sophisticated than VMs and object storage. On the other hand, it's pretty embarrassing at this point for something as fundamental as Docker to be in a single region. Most cloud providers make inter-region failover reasonably achievable.
- richardwhiuk 1y agoAlmost all cloud providers help here by having inter-region failures as well. There are multiple AWS services which are "global" in the sense that they are entirely hosted out of AWS East 1
- wredcoll 1y ago> Isn't it everyone using multiple cloud providers nowadays? Why are they affected by single cloud provider outage? No? I very much doubt anyone is doing that.
- DiggyJohnson 1y agoMulti cloud is not nearly as trivial as often implied to implement for real world complex projects. Things get challenging the second your application steps off the happy path
- walkabout 1y ago> Isn't it everyone using multiple cloud providers nowadays? Oh yes. All of them, in fact, especially if you count what key vendors host on. > Why are they affected by single cloud provider outage? Every workload is only on one cloud. Nb this doesn’t mean every workflow is on only one cloud. Important distinction since that would be more stable.
- pmontra 1y agoBecause even if service A is using multiple cloud providers not all the external services they use are doing the same thing, especially the smallest one or the cheapest ones. At least one of them is on AWS East-1, fails and degrades service A or takes it down. Being multi-cloud does not come for free: time, engineers, knowledge and ultimately money.
- KronisLV 1y agoI guess people who are running their own registries like Nexus and build their own container images from a common base image are feeling at least a bit more secure in their choice right now. Wonder how many builds or redeployments this will break. Personally, nothing against Docker or Docker Hub of course, I find them to be useful.
- frenkel 1y agoOnly if they get their base images from somewhere else...
- bravetraveler 1y agoPull-through caches are still useful even when the upstream is down... assuming the image(s) were pulled recently. The HEAD to upstream will obviously fail [when checking currency], but the software is happy to serve what it has already pulled. Depends on the implementation, of course: I'm speaking to 'distribution/distribution', the reference. Harbor or whatever else may behave differently, I have no idea.
- Sphax 1y agoWe run Harbor and mirror every base image using its Proxy Cache feature, it's quite nice. We've had this setup for years now and while it works fine, Harbor has some rough edges.
- thephyber 1y agoI came here to mention that any non-trivial company depending on Docker images should look into a local proxy cache. It’s too much infra for a solo developer / tiny organization, but is a good hedge against DockerHub, GitHub repo, etc downtime and can run faster (less ingress transfer) if located in the same region as the rest of your infra.
- jsmeaton 1y agoGuess where we host nexus..
- jdthedisciple 1y agoSo thus far today outages are reported from - AWS - Vercel - Atlassian - Cloudflare - Docker - Google (see downdetector) - Microsoft (see downdetector) What's going on?
- ta1243 1y agoOr they all rely on AWS, because over the last 15 years we've built an extremely fragile interconnected global system in the pursuit of profit, austerity, and efficiency
- benrutter 1y agoWait, Google and Microsoft rely on AWS? That seems unlikely? (does it? I wouldn't really know to be honest)
- thephyber 1y agoIt’s very likely they’ve bought companies that were built on AWS and haven’t migrated to use their homegrown cloud platforms.
- ta1243 1y agoMore likely the outage reports for google and microsoft are based around systems which also include aws
- ssl-3 1y agoIn terms of user reports: Some users don't know what the hell is going on. This is a constant. For instance: When there's a widespread Verizon cellular outage, sites like downdetector will show a spike in Verizon reports. But such sites will also show a spike in AT&T and T-Mobile reports. Even though those latter networks are completely unaffected by Verizon's back-end issues, the graphs of user reports are consistently shaped the same for all 3 carriers. This is just because some of the users doing the reporting have no clue. So when the observation is "AWS is in outage and people are reporting issues at Google, and Microsoft," then the last two are often just factors of people being people and reporting the wrong thing. (You're hanging out on HN, so there's very good certainty that you know what precisely what cell carrier you're using and also can discern the difference betwixt an Amazon, a Google, and a Microsoft. But lots of other people are not particularly adept at making these distinctions. It's normal and expected for some of them to be this way at all times.)
- ic4l 1y agoThis broke our builds since we rely on several public Docker images, and by default, Docker uses docker.io. Thankfully, AWS provides a docker.io mirror for those who can't wait: FROM public.ecr.aws/docker/library/{image_name} In the error logs, the issue was mostly related to the authentication endpoint: ▪ https://auth.docker.io https://auth.docker.io → "No server is available to handle this request" After switching to the AWS mirror, everything built successfully without any issues.
- anon7000 1y agoI manage a large build system and pulling from ECR has been flaking all day
- firloop 1y agoI wasn't able to get this working, but I was able to use Google's mirror[0] just fine. Just had to change FROM {image_name} to FROM mirror.gcr.io/{image_name} Hope this helps! [0]: https://cloud.google.com/artifact-registry/docs/pull-cached-dockerhub-images https://cloud.google.com/artifact-registry/docs/pull-cached-...
- ic4l 1y agoWe tried this initially FROM mirror.gcr.io/{image_name} We received failed to resolve source metadata for mirror.gcr.io/ So it looks like these services may not be true mirrors, and just functioning as a library proxy with a cache. If you're image is not cached on one of these then you may be SOL.
- da768 1y agoDuring the last Docker Hub outage we found Google mirrors lost all image tags after a while. Image digest references would probably work
- CamouflagedKiwi 1y agoMild irony that Docker is down because of the AWS outage, but the AWS mirror repos are still running...
- deleted 1y ago[deleted]
- sschueller 1y agoWhat are good proxy/mirror solutions to mitigate such issues? Best would be an all in one solution that for example also handles nodejs, packigist etc.
- bravetraveler 1y agoPulp is a popular project for 'one stop shop', I believe. Personally, always used project-specific solutions like 'distribution/distribution' for containers from the CNCF. This allows for pull-through caching with relatively little setup work.
- conradfr 1y agoIs there a built-in way to bypass the request to the registry if your base layers are cached?
- edoceo 1y agopull: never?
- wolfgangbabad 1y agoeven reddit throws a lot of 503s when adding/editing comments
- throw-10-13 1y agoreddit is always going down, thats the least surprising thing about this
- wolfgangbabad 1y agohttps://www.bbc.com/news/live/c5y8k7k6v1rt https://www.bbc.com/news/live/c5y8k7k6v1rt
- dd_xplore 1y agoDoes it decrease the AWS's nine 9s ?
- speedgoose 1y agoThe marketing department did the maths and they said no.
- nobleach 1y ago"MOST of the time" we're nine 9s.
- helpfulmandrill 1y agoI wonder if this is why I also can't log in to O'Reilly to do some "Docker is down, better find something to do" training...
- p0w3n3d 1y agoJust install a pull-through proxy that will store all the packages recently used.
- phillebaba 1y agoShameless plug but this might be a good time to install Spegel in your Kubernetes clusters if you have critical dependencies on Docker Hub. https://spegel.dev/ https://spegel.dev/
- deleted 1y ago[deleted]
- CaptainOfCoit 1y agoThere is a couple of alternatives that mirrors more than just Docker Hub too, most of them pretty bloated and enterprisey, but they do what they say on the tin and saved me more than once. Artifactory, Nexus Repository, Cloudsmith and ProGet are some of them.
- phillebaba 1y agoSpegel does not only mirror Docker Hub, and works a lot differently than the alternatives you suggested. Instead of being yet another failure point closer to your production environment, it runs a distributed stateless registry inside of your Kubernetes cluster. By piggy backing off of Containerds image store it will distribute already pulled images inside of the cluster.
- CaptainOfCoit 1y agoI'll be honest and say I hadn't heard of Spegel before, and just read the landing page which says "Speed up container pulls and minimize downtime with a stateless peer-to-peer OCI registry mirror for efficient image distribution", so it isn't exactly clear you can use it for more things than container images.
- scuff3d 1y agoWhat exactly does "stateless" mean in this context?
- phillebaba 1y ago
- Zekio 1y agowhat good options are there for container registry proxies / caches to protect against something like this?
- phillebaba 1y agoI build Spegel to keep my Kubernetes cluster running smoothly during an outage like this. https://spegel.dev/ https://spegel.dev/
- PhilipRoman 1y agohttps://docs.docker.com/docker-hub/image-library/mirror/ https://docs.docker.com/docker-hub/image-library/mirror/ ?
- l2dy 1y agoRecovering as of October 20, 2025 09:43 UTC > [Monitoring] We are seeing error rates recovering across our SaaS services. We continue to monitor as we process our backlog.
- darkamaul 1y agoFor other people impacted, what helped me this morning was to use the `ghcr`, albeit this is not a one-to-one replacement. Ex: `docker pull ghcr.io/linuxcontainers/debian-slim:latest`
- TimWolla 1y agoThat image is over one year old: https://github.com/linuxcontainers/debian-slim/pkgs/container/debian-slim https://github.com/linuxcontainers/debian-slim/pkgs/containe... Google Container Registry provides a pull-through mirror, though, just prefix `mirror.gcr.io` and use `library` as the user for the Docker Official Images. For example `mirror.gcr.io/library/redis` for https://hub.docker.com/_/redis https://hub.docker.com/_/redis.
- remram 1y agoIs there a way to configure alternate mirrors in containerd?
- jabiko 1y agoIts impressive that even though registry-1.docker.io returned 503 errors they where able to keep a the metric "Docker Registry Uptime" at 100%.
- lbruder 1y agoWell, the server was up, it was just returning HTTP 503...
- deleted 1y ago[deleted]
- cloudking 1y agoI'm fairly new to Docker. Do folks really rely on public images and registries for production systems? Seems like a brittle strategy.
- edoceo 1y agoYes, 1000s of orgs. Larger players might use a pull-through-cache - but it's not as common as it should be. Similar issue for other software-supply-chain (NPM, pyPi, etc)
- theanonymousone 1y agoIt's quite funny/interesting that this is higher in HN front page than the news of the AWS outage that caused it.
- deleted 1y ago[deleted]
- mcintyre1994 1y agoNot on the real secret front page! https://news.ycombinator.com/active https://news.ycombinator.com/active :)
- pknopf 1y agoWhat does the "active" page sort by?
- mcintyre1994 1y agoAccording to https://news.ycombinator.com/lists https://news.ycombinator.com/lists it's "Most active current discussions" I find that it better surfaces the best discussion when there are multiple threads (like in this example), and it keeps showing slightly older threads for longer when there's still discussion happening.
- cakeday 1y agoThat's informative, I wasn't aware of that way to view HN, thanks.
- 2OEH8eoCRo0 1y agoThe internet was designed to be fault tolerant and distributed from the beginning and we still ended up with a handful of mega hosts.
- gjvc 1y agomirror.gcr.io is your friend
- tj_591 1y agoHi all, Tushar from Docker here. We’re sorry about the impact our current outage is having on many of you. Yes, this is related to the ongoing AWS incident and we’re working closely with AWS on getting our services restored. We’ll provide regular updates on dockerstatus.com . We know how critical Docker Hub and services are to millions of developers, and we’re sorry for the pain this is causing. Thank you for your patience as we work to resolve this incident. We’ll publish a post-mortem in the next few days once this incident is fully resolved and we have a remediation plan.
- tonyabracadabra 1y agopls bring it back
- Bayko 1y ago[flagged]
- freedomben 1y agoPart of me hopes that we find out that Dynamo DB (which sounds like was the root of the cascading failures) is shipped in a Docker image which is hosted on Docker Hub :-D
- tj_591 1y agoWe’ve published an incident report outlining what happened and the steps we’re taking to strengthen resilience in the face of upstream service interruptions. - https://www.docker.com/blog/docker-hub-incident-report-october-20-2025/ https://www.docker.com/blog/docker-hub-incident-report-octob...
- m463 1y agothis is by design docker got requests to allow you to configure a private registry, but they selfishly denied the ability to do that: https://stackoverflow.com/questions/33054369/how-to-change-the-default-docker-registry-from-docker-io-to-my-private-registry https://stackoverflow.com/questions/33054369/how-to-change-t... redhat created docker-compatible podman and lets you close that hole /etc/config/docker: BLOCK_REGISTRY='--block-registry=all' ADD_REGISTRY='--add-registry=registry.access.redhat.com'
- anon7000 1y agoSadly doesn't help if you were using ECR in us-east-1 as your private registry. :(
- compootr 1y agoI still think this is an acceptable footgun (?) to have. The expressiveness of downloading an image tag with a domain included outweighs potential miscommunication issues. For example, if you're on a team and you have documentation containing commands, but your docker config is outdated, you can accidentally pull from docker's global public registry. A welcome change IMO would be removing global registries entirely, since it just makes it easier to tell where your image is coming from (but I severely doubt docker would ever consider this since it makes it fractionally easier to use their services)
- scuff3d 1y agoThis is a huge stretch. Even if you could configure a default registry to point at something besides docker.io a lot of people, I'd say the vast majority, wouldn't have bothered. So they'd still be in the same spot. And it's not hard to just tag images. I don't have a single image pulling from docker.io at work. Takes two seconds to slap <company-repo>/ at the front of the image name.