8 ms·
Fastmail 30 June outage post-mortem
- xeromal 3y agoRedundancy is the name of the game. Glad they realize this.
- lern_too_spel 3y agoThey will still be in only one datacenter. This is worse than its competitors.
- hedora 3y agoThat doesn't seem to be true from a durability perspective: https://www.fastmail.help/hc/en-us/articles/1500000278242 https://www.fastmail.help/hc/en-us/articles/1500000278242 Read the section on "slots". They keep two copies in New Jersey, and one in Seattle. However, based on the post mortem, it sounds like they're not willing to invoke failover at the drop of a hat. They allude to needing complex routing to keep their old good-reputation IP addresses alive. That might have something to do with it. (They were "only" 3-5% down during the outage, which is bad for them, but not unusual by industry standards.)
- withinboredom 3y agoYou might have multiple datacenters, but until everyone isn’t in the same office, you still have the same problem (office can burn down, or fall down and bury everyone). Also, it’s email. It was literally designed to work in such a way that you can be down for days and still get your email.
- lern_too_spel 3y ago> You might have multiple datacenters, but until everyone isn’t in the same office, you still have the same problem (office can burn down, or fall down and bury everyone). Fastmail has multiple offices. > Also, it’s email. It was literally designed to work in such a way that you can be down for days and still get your email. Delivery is designed that way, not storage. Once Fastmail tells the sender that a message was delivered to an inbox, Fastmail cannot ask the sender to redeliver it if its inbox storage is lost.
- nik736 3y agoCrazy that they were single homed. Edit: After seeing the network diagram I have even more questions. What happens if CF is down? This all seems cobbled together and very prone to failures.
- OhNoMyqueen 3y agoThey're still singled homed, right? They just added redondancy to in/outbound routes.
- stuff4ben 3y agoIf CloudFlare is down, a significant portion of the Internet is down. Not that it's an excuse, but this isn't Microsoft or Apple. I'm sure funds have to be allocated to take into account the likelihood of something being down. But by all means write a blog post and tell them what they're doing wrong and how you'd fix it. Maybe they'll hire you...
- theideaofcoffee 3y agoAnd you don't have to have the resources of Microsoft or Apple to plan and build for the eventuality that a provider becomes intermittent or unavailable. There are fundamental aspects of running an internet-facing service and they failed at one of the most basic.
- stuff4ben 3y agoLOL ok, they "failed". They haven't had an outage like this in decades and this one only affected a small number of their clients. But sure, let's spend money on providing a backup for CF. Armchair QBs are the worst.
- theideaofcoffee 3y agoYet they still had the outage. I take exception to being called an 'armchair QB' when most of my career has been spent being called in to repair failures like this, providing postmortem advice to weather future ones and fix technical and cultural issues that give rise to just this type of thinking: oh, it won't happen to us because it has never happened to us.
- dizhn 3y agoNYI. Good people.
- theideaofcoffee 3y agoNetgear switches? In an environment like this? I’ll give them the benefit of the doubt that that is maybe a provider-owned thing, and that they have an 'enterprise' line, but, really... Netgear. The firewall brand isn’t revealed in the network diagram, but what is it, a $100 sonicwall? Should I be concerned keeping all my email, business and personal, there about what other parts of their infrastructure they are cheaping out on? When you are running a service like this, redundancy among transit providers is the most basic, table-stakes thing you can do. It's almost negligent to not have that.
- dktnj 3y agoNetgear do some half decent fully managed switches. It’s not all blue crap off Amazon. The worst switches I ever used were HPE ones in the old C5000 blade chassis. Absolute turds. Packet loss, constant port failures and complete hangs. HPE’s solution was to tell us to buy new ones.
- bluedino 3y agoThe worst switches I've ever used would probably be various 'Cisco' switches from their small business line, usually ones that ran the same OS used when they were sold under different names like 3Com or Linksys.
- incahoots 3y agooh god, when I worked for a small telecom in the midwest they heavily used the 3Com switches. They were the bane of my existance, things would power loop, or my favorite, continue to work but prevent any sort of access to them.
- publicmail 3y ago> It’s not all blue crap off Amazon. To be honest, those little blue unmanaged Netgear switches aren’t bad at all. We have dozens of them in our lab at work running 24/7 for like decades and have never had a failure that I remember.
- JohnMakin 3y agoI'm not clear from the post-mortem why the outbound packets were having issues. Was it cloudflare? Did someone accidentally delete an outbound route? Why couldn't they see the issue themselves? I only have more questions now.
- ilyt 3y agoIt looks like their upstream provider fucked something up so until they release that info we can only speculate.
- scandox 3y agoI'm surprised they don't own their own IPs. In the email world I would say that's quite important. Seems a tad casual to say "luckily they are willing to lease them"...
- withinboredom 3y agoFirst, someone has to be willing to sell them…
- tivert 3y agoYeah, me too, but it sounds like they have good reason and are working to fix it: > Thankfully, NYI were willing to lease us those addresses, because IP range reputation is really important in the email world, and those things are hardcoded all over the place — but it has caused us complications due to more complex routing. Over the past year, we’ve been migrating to a new IP range. We’ve been running the two networks concurrently as we build up the trust reputation for the new addresses. Might be one of those things that they did when they were small, but then got hard to change. Hopefully they will fully own their new addresses. The thing I'm surprised about is that they only have "single path for traffic out to the internet."
- gwright 3y agoI'm a little rusty in this area, but I'm pretty sure you'll have a hard time getting a direct IP allocation (vs from your transit provider) unless you are multi-homed, which apparently they were not. There is a nice tick up in complexity when you go to advertising your address space via BGP to multiple providers.
- electroly 3y agoI'll admit that seeing they are single-homed is sketchier than I assumed Fastmail's infrastructure was. I've been using Fastmail for years and I like them, but they are clearly big enough to have a second transit provider, and have been for many years. I'm amazed it took an outage for them to decide to get one. I appreciate the post-mortem but I felt better before I had read it.
- majkinetor 3y agoJudged by comments and upvotes here, and previously, I am actually amazed how almost everybody thinks how perfect and unfallable he is and how he surrely deserves 100% perfection from the day he is born to the day his grand-grand-... kids die.
- daenney 3y agoIncidents happen, that's life. You can hedge your bets, but some things are out of your control. Communicating with your customers however is entirely within your control. Fastmail did a poor job of it. Their status page was useless beyond an initial "we found an issue" and then nothing for almost 11hrs. Their Twitter account was the same story, didn't bother with the Mastodon account at all. Unfortunately they don't seem to realise or recognise that they dropped the ball on this and that goes entirely unaddressed. I'm also not really charmed with how they try to minimise the importance of the incident by repeating it only affected 3-5% of the customers. That might very well be. But those are real people and real businesses that rely on your services that were unavailable for the whole of the EU workday and a significant part of the US workday. Everyone I know who was affected is a paying customer, none of us have received so much as a communication or apology for it. For a company that's been on the internet since 1999, the single-homed setup is a little shocking. But fine, it's being addressed. But both the communication during and after the incident don't inspire a ton of confidence.
- luuurker 3y agoOn one hand, I too like better communications in situations like this. On the other, you knew they were aware of the issue and working on it. Updating the status page with a "we're trying to fix it" every hour or so wouldn't speed up the process of fixing the problem or help you in any way.
- daenney 3y agoI didn’t ask for an update every hour, but an 11hr stretch of silence is not cool. An update every 3 would’ve been fine. Even if there isn’t much new to share, reiterating you’re working on it is useful to reassure your customers. I’m fairly certain they could’ve found something more meaningful to communicate than “still twirling the thumbs”. For example, the update after 11hrs included the tip to switch to a VPN. That would’ve been useful to communicate.
- bananapub 3y ago> I'm also not really charmed with how they try to minimise the importance of the incident by repeating it only affected 3-5% of the customers. it's not meant to be charming, it's meant to convey the scale of the outage. as a customer, even if I'm not effected, there's a world of difference between 3-5% and 99%.
- Lk7Of3vfJS2n 3y agoI don't know if IP range reputation is part of SMTP but I wish sending email worked without it.
- _joel 3y agoIt's not part of SMTP, it's part of the cludge of anti-spam measures we bolted on top of it.
- micah94 3y ago...bolted on top of DNS.
- _joel 3y agoSure, in addition to this though. IP reputation comes from spamhaus.org and the like and implemented in spamassassin or something similar, at SMTP(ish) level. What you're talking about there is SPF/DKIM etc, which is an anti-spam measure but not IP reputation :)
- N0RMAN 3y agoAm I the only one who thinks that this post-mortem for a margin of email users worldwide is a pure marketing gimmick to attract more "technical" users?
- Obscurity4340 3y agoI don't think the Lastpass approach is how anyone gets any new customers, to say nothing of the technical ones