6 ms·
Letsencrypt Service Disruption
- tersers 5y agoWell this saves me a half hour of troubleshooting!
- corentin88 5y agoI’m not sure if this is affecting websites already having a certificate provided by Let’s Encrypt. I believe not.
- pierrebeaucamp 5y agoJust the issuing of new certificates
- remram 5y agoThe OCSP responder being down could affect websites. https://en.wikipedia.org/wiki/Online_Certificate_Status_Protocol https://en.wikipedia.org/wiki/Online_Certificate_Status_Prot...
- LinuxBender 5y agoFWIW people can mitigate that issue by using OCSP stapling [1] assuming their load balancer / web server supports caching the response. [1] - https://en.wikipedia.org/wiki/OCSP_stapling https://en.wikipedia.org/wiki/OCSP_stapling
- hannob 5y agoTraditionally OCSP stapling implementations in the major webservers (nginx+apache) have been pretty bad and are not capable of mitigating OCSP server downtimes. Apache has a new stapling implementation since a few versions that can be enabled with "MDStapling On", I don't think it gets enabled by default yet.
- mholt 5y agoCaddy has had a robust stapling implementation for years.
- dlkmp 5y agoDoes the caching reliably work in common webservers by now? I remember having read a couple years ago that Apache would simply clear its cache if the connection to the ocsp provider breaks (or did something similar unhelpful, resulting in an error to the enduser).
- hannob 5y agoVery likely you read my blogpost. See the update at the top: https://blog.hboeck.de/archives/886-The-Problem-with-OCSP-Stapling-and-Must-Staple-and-why-Certificate-Revocation-is-still-broken.html https://blog.hboeck.de/archives/886-The-Problem-with-OCSP-St...
- mholt 5y agoCaddy does OCSP stapling + caching reliably (and has for years). It will even auto-renew your certificate if it gets an OCSP response of "Revoked".
- LinuxBender 5y agoIt's hit-and-miss per LB/server and I have not seen this become a priority since it's not a super popular feature. Here [1] [2] are a couple articles on the topic. My experience has been with HAproxy and F5 load balancers. HAProxy uses an out-of-band process to lay down a .ocsp file and load it via the API. This in effect acts like a cache assuming the script creating the .ocsp file has error handing to avoid clobbering the file if the upstream OCSP endpoint can not be reached. F5 load balancers will cache the response in memory. I have not tested Apache with OCSP stapling/caching recently so I can only assume based on feedback from others here that they have not improved it. I would expect nginx to improve now that they are owned by F5, maybe, eventually. I am a fan of OCSP stapling/caching for the privacy aspect. No need for browsers to leak to the OCSP end-point what domain you are visiting. There are enough nosy people sniffing our traffic already. [1] - https://www.keycdn.com/support/ocsp-stapling https://www.keycdn.com/support/ocsp-stapling [2] - https://blog.cloudflare.com/high-reliability-ocsp-stapling/ https://blog.cloudflare.com/high-reliability-ocsp-stapling/
- sneak 5y agoI imagine that, for still-valid certificates, it will almost always fail open.
- pgporada 5y agoWe have been serving OCSP responses out of another datacenter since the initial status.io went up.
- tetha 5y agoHeh. We shutdown a couple of legacy mystery servers today for a second scream test. I hope this wasn't us! (and no, we are not affiliated with Lets Encrypt. It's just a funny coincidence. but who knows what those mystery boxes all in all do?)
- jchw 5y agoI think Let’s Encrypt is great, but it is worth noting that other ACME providers do exist, for example, ZeroSSL. Worth noting since it seems like LE is quickly becoming a gigantic SPOF for the internet!
- zoobab 5y agoRecipe for disasters. Who designed this SPOF?
- stingraycharles 5y agoNobody “designed” it and as the parent said, there are alternatives. It’s just that SSL certificates need to be generated by an authority, and LE is the authority chosen by most people that use LE’s technology.
- throwaway290232 5y agoA couple people. Working groups that decide internet standards. Industries and private companies that control how the internet works. Web browsers that determine how the world wide web works. Organizations made up of people from each of these categories. And of course, every single user, developer, vendor, etc that hungrily adopts and subsequently becomes dependent on everything they put out. There is no God, but there are working groups of angels discussing Existence Standards.
- Shorel 5y agoOr a single point of entry for a powerful enough government.
- sneak 5y agoEvery root CA (your browser and OS have dozens built in, including those explicitly issued by governments, including ones hostile to you) is a point of entry for a government to compromise TLS. There's no "single" one, hence all of the varied efforts (DANE, CT, et c) to reduce this attack surface.
- tomjen3 5y agoThis really shows why we need more than one competitor in this space.
- aaomidi 5y agohttps://zerossl.com/ https://zerossl.com/ https://www.xf.is/2020/06/30/list-of-free-acme-ssl-providers/ https://www.xf.is/2020/06/30/list-of-free-acme-ssl-providers...
- pbiggar 5y agoOh good. Something was going wrong issuing new customer certs for Darklang. Glad it's not my bug!
- throwaway290232 5y agoThere's a particular class of bug that affects all datacenters in a globally distributed service. Almost invariably it comes down to knock-on effects due to multiple things happening "the wrong way" in tandem. Those individual failures typically happen because somebody said "well yeah we should have a better way to do this, but the other components have redundancy, so shrug". If a single domino isn't tested well and has all potential failure modes well documented and tested for, it can still take down the entire chain of dominos.
- jedberg 5y agoOr just a bad global configuration change. We used to take down Netflix globally with config changes until we changed the system to have config changes scoped to single regions. And even then sometimes the failure wouldn't show up until multiple regions got the config change.
- mrkurt 5y agoAlways bet on a config change. Or DNS.
- pgporada 5y agoNeither of these apply to this situation though. We're working on it.
- mrkurt 5y agoI lose many of my bets.
- terom 5y agoIt sounds like it was a hardware issue [1]. First rule of electronics applies: thou shalt check voltages. > [Update] We're conducting follow-up maintenance to address power supply issues from our earlier service disruption. All production API services may be down for up to 30 minutes. [1] https://letsencrypt.status.io/pages/maintenance/55957a99e800baa4470002da/60f5f2460451310998105e7f https://letsencrypt.status.io/pages/maintenance/55957a99e800...
- jedberg 5y agoIn theory this should have no effect, right? Since the certificates should be renewing a few days before they expire, in theory it could be down for a couple of days. Good way to make sure all the software is resilient to renewal failures though!
- mrkurt 5y agoThis affects us! But only because we do certs-as-a-service. Existing certs won't break but it has wrecked the UX for a lot of customers.