23 ms·
Tarsnap outage postmortem
- mplewis 3y agoI always appreciate seeing a professional, courteous, and honest postmortem like this one.
- cperciva 3y agoblinks Ok, I really wasn't expecting this to land at the top of HN. I'd love to stick around to answer any questions people have, but it's 10PM and my toddler decided to go to bed at 5PM... so if I'm lucky I can get about 4 hours of sleep before she decides that it's time to get up. I'll check in and answer questions in the morning.
- gfv 3y agoHow long do you keep the transaction logs before rewriting them? I too had a few EC2 instances go down with signs of being severed from the EBS in the recent couple of weeks; mine were in eu-west.
- cperciva 3y agoThere's a continual background cleaning process which depends on the amount of storage which can be reclaimed -- there's a tradeoff between cleaning too slowly (and paying for wasted storage) and cleaning too fast (and paying for lots of S3 operations). I think it averages a couple weeks right now.
- throwawaaarrgh 3y agoAre you gonna switch to us-east-2?
- stigz 3y agoWhy would I use your service over restic? God bless you Colin, but reading this, it appears you're the only one in charge of the infrastructure for this service. I'm glad you're clear about no SLA, but this seems like a big liability between me and my backups.
- deleted 3y ago[deleted]
- IntelMiner 3y agoI'm curious how the prices shake out against services like Wasabi, since it's just dumping to an AWS S3 bucket Wasabi does $7/TB with no ingress/egress fees. My NAS is set up to rclone to it about once a day and I've yet to have any problems
- crossroadsguy 3y agoI haven’t checked the pricing in a long time but you can use Tarsnap also if you have to backup only 7.3kb (okay I might ne exaggerating here but you get the drift) and pay for only that much. You can’t do that with Wasabi et al. Also it’s really simple and does what it says it does, nothing more, nothing less. In today’s everything convoluted and bloated world this is a luxury imho. The GUI app is also quite good and functional. Support is prompt (that is if you need it). You don’t have to worry about file being deleted just because your machine didn’t connect or backup for some time even if you keep paying (hello Backblaze) etc. I mean there’s no circus, melodrama , and cliffhangers involved. I personally would never use it backup my entire laptop, due to price alone. But I have a subset of VVI files and Tarsnap is one of more than one backups for those files. So for that use-case Tarsnap is perfect for me, so far.
- Aeolun 3y agoBackblaze has kept my ‘shutdown two years ago’ machine data without issue. What problems did you have with them (or did others have)?
- terinjokes 3y agoBackblaze has a policy of allowing backups of external disks, but the disks have to be connected at least once every 30 days, or they'll delete the backups. I understand they want to avoid abuse, but the lack of any grace period, or ability of support to ad an override, really soured the service for me.
- mike_d 3y agoThis was an extremely well written and thoughtful postmortem, but I hope to never see one from you again. :)
- Tepix 3y agoIt was a postmortem without the mandatory "how can we prevent this in the future" steps…
- bluehatbrit 3y agoI think that's a little unfair given what was in the postmortem. It may not be a separate section with the key points, but the information is all there of what the issues were and what the solutions are. I think it's fair to assume they're actually acting on those without them needing to be reiterated at the bottom of the page.
- Tepix 3y agoWell, for sure he has fixed several bugs, but he didn't say that he would be testing his disaster recovery procedure every year in the future for example.
- cperciva 3y agoYes, rehearsing the process every year is the main lesson learned. Sorry, it was getting late and I wanted to get the email out so I cut it short.
- acedTrex 3y agoI agree, we don't really need a "key points/future actions" section that boils down to "The service will be geo redundant"
- tptacek 3y agoIn 15+ years of running this service, this is one of two (2) postmortems he's ever published, and the first in eleven (yes, 11) years.
- dharmapure 3y agoThank you for the post-mortem Colin and I hope you get some sleep!
- cperciva 3y agoThanks, I did! My long suffering wife was up at 3:30 though. :-(
- bombcar 3y agoTime to get your toddler providing round-the-clock support! ;) Have been having some luck reading https://www.amazon.com/No-Cry-Sleep-Solution-Toddlers-Preschoolers/dp/0071444912 https://www.amazon.com/No-Cry-Sleep-Solution-Toddlers-Presch... - available everywhere libraries (blockbuster for books!) are found.
- cperciva 3y agoShe's generally a wonderful girl. Right now she's dealing with her second molars coming out and just picked up a cold though, which is throwing off her sleep schedule.
- jacquesm 3y agoIn future postmortems (of which I hope there will be very few or even none) you may want to spell out your 'lessons learned' to show why particular items will never recur.
- idlewords 3y agoIt always amuses me how people want reassurance that the next crisis will be a fresh, new problem, and not one the person can demonstrably solve. A lot of 'lessons learned' analysis boils down to this: in order to prevent a recurrence of X, we introduced complex subsystem Y, the unexpected effects of which you can read about in our next post-mortem.
- bityard 3y agoThat's an overly cynical take, post-mortems are not for anyone's reassurance, they are a learning opportunity. The airline industry is as safe as it is because every accident gets thoroughly investigated with detailed reports ("post-mortems") including what to do differently going forward. These are taken as gospel among all players in the industry and as a result, you very rarely see two different accidents caused by the same thing anymore.
- jacquesm 3y agoThat was entirely not what I was getting at and is a cheap shot that is well beneath you, especially because I suspect that you know that that wasn't what I was getting at.
- idlewords 3y agoMy comment wasn't intended personally; your words about "will never recur" just reminded me of this peculiarity of software systems, where it's often error handling/monitoring/backups/etc. that cause cascading failures in the systems they're intended to safeguard. I'm sorry if I misconstrued your meaning, but I am flattered that you think there are things beneath me!
- 3y ago
- nodesocket 3y agoSome recommendations on the AWS front (not sure if some of these are already implemented since the postmortem does not go into AWS details). - Setup nightly automatic snapshots of EBS volumes (this is supported natively now in AWS under lifecycle manager). - Use EBS volumes of the new GP3 type, and perhaps use provisioned IOPS. - Setup a auto-scaling group with automatic failover. Of course increases cost, but should be able to automatically failover to a standby EC2 instance (assuming all the code works automatically which the blog post indicates is not currently the case).
- e63f67dd-065b 3y agoCan you say a bit more about the log-structured S3 filesystem? I wrote something very similar recently (https://github.com/isaackhor/objectfs https://github.com/isaackhor/objectfs) and I'm curious what made you settle on that architecture. The closest thing I know of that's similar is Nvidia's ProxyFS (https://github.com/NVIDIA/proxyfs https://github.com/NVIDIA/proxyfs)
- nextaccountic 3y ago> the central Tarsnap server (hosted in Amazon's EC2 us-east-1 region) What prevents you to distribute load among other regions? (Also: did you ever think about abandoning AWS?)
- rlt 3y agoNice write up. A couple questions: - The use of “I” begs the question: what’s the “bus factor” of Tarsnap? If you were unavailable, temporarily or permanently, what are the contingency plans? - Will you be making any other changes to improve the recovery time, or did the system mostly function as designed? For example having a hot spare central server?
- LinAGKar 3y agoWhat I'm wondering is, I had data on Tarsnap, why am I only hearing about this now?
- switch007 3y agoNot to be that guy, but it’s unreadable either zoomed in or in reader mode either horizontal or landscape on iOS. Colin, could the website be updated to the 2010s? :P
- ehPReth 3y agoThis should work in reader mode: https://pastebin.com/raw/hanm8mgG https://pastebin.com/raw/hanm8mgG
- switch007 3y agoThanks!
- memefrog 3y agoIt's a mailing list archive. Use a real computer.
- switch007 3y agoHow dare I use my phone before I get to my computer.
- Semaphor 3y agoJust FYI, Firefox Reader mode works great with it.
- kevincox 3y ago"Great" is a bit of a stretch as there are random short lines where they were originally hard-wrapped. But it certainly does make it readable.
- TheDong 3y agoIt's not Colin's fault that you're using a browser that can't render an html rendition of an email which has been widely in use since before iOS existed. This is entirely Safari's fault for not having good compatibility with a common existing webpage format. Anyway, if you're the intended audience (someone using tarsnap), you also received a copy to your email address, where you can read the text with your email reader of choice.
- RockRobotRock 3y agoAren't these storage prices absurd? Please let me know if I'm misunderstanding.
- GhostWhisperer 3y agoyes, people have been saying they should "charge more" for over a decade
- RockRobotRock 3y agoCan you not be snide and please help me understand? It seems 50 times more expensive than B2. I'm genuinely curious about the product.
- stigz 3y agoThe service is a layer on top of S3 if that helps answer things. https://www.tarsnap.com/faq.html#is-tarsnap-reliable https://www.tarsnap.com/faq.html#is-tarsnap-reliable
- yjftsjthsd-h 3y agoThen the obvious question is "why would I use this instead of something else over S3" (ex. rclone), to which I think the answer is ease of use (don't need to deal with AWS yourself, encryption/deduplication/compression handled for you, nice interface), which isn't everything to everyone but is certainly useful.
- RockRobotRock 3y agoThank you.
- alanfranz 3y agoYou need a remote service that keeps backup readonly. You’re not covering attack scenarios if you just use raw object storage from your client machine. I have written about this some time ago if you’re interested: https://www.franzoni.eu/ransomware-resistant-backups/ https://www.franzoni.eu/ransomware-resistant-backups/
- zokier 3y agoBased on the description it sounds like it should be relatively easy to test this recovery process on a regular basis, to catch any lingering bugs and evaluate the recovery time. As they say, the only backups are the ones you have tested.
- baz00 3y agoAs someone who just discovered my DR process does not work by testing it, 100% this. The only plan that is likely to work is a repeatable tested one.
- tialaramex 3y agoIdeally, the thing you do in an emergency is largely routine, so that it happens by instinct rather than being a special case you need to remember. It should not be different in arbitrary ways. For example in both trains and cars, thanks to anti-lock braking, the correct way to stop the vehicle ASAP is to brake just like normal but as hard as you can, the computers will automatically solve the much trickier problem of turning your input into maximum deliverable braking force by periodically releasing brakes on sticking wheels. If you run a fire drill, it's surprisingly difficult to get employees to use fire doors that they're used to finding alarmed and unusable. Even though intellectually they know that, say, the door at the bottom of the stairwell is a fire door, with crash bars and leads directly to the outside world, and this is a fire drill, they are likely to (for example) exit on a higher floor and go through a chokepoint lobby, as they would normally, instead of following this safer path that is emergency only. Sadly it is hard to fix buildings after construction if they were designed with such "unused" emergency exits. For a backup process, having restoring machine images be a service that is sometimes, though not constantly, used anyway for some other reason, is a good way to be comfortable with how it works, that it works, etc. At work for example we routinely test upgrades on test servers restored from a recent backup. Restore serviceA to testA, apply upgrade, discover upgrade completely ruins the service, throw testA away and report this upgrade is garbage. But in the process we gained confidence in the restore process, infrastructure people instead of trying to recall something they only ever did in a drill, when things go badly wrong are very used to this procedure because they do it "all the time".
- aborsy 3y agoTarsnap is undoubtedly expensive, but it also donates to various efforts! Neglecting the pricing, does Tarsnap have any advantage over Restic? Restic also deduplicates, using little data.
- mattbee 3y agoThe deduping in restic is just on the edge of acceptable for me, making me think I'd have trouble with a lot more data. Basically the one a month "prune" operation takes about 36h (to B2) . I feel I could be tuning something but also it works and I don't want to touch it.
- sandgiant 3y agoCurious how much you backup, which version of restic you're running and why you think the deduplication is borderline unacceptable. There were several major (orders of magnitudes) improvements made to pruning within the past ~1 year, that's why I'm interested.
- mattbee 3y agoA straight upgrade, that I can do :) It's been running for years without one. I was only edgy about it because when it takes 36h it blocks the next daily backup, and I wondered whether that was going to get worse (it hasn't).
- chibea 3y agoThe max-unused percentage feature is well worth it to 80/20 the prune process and only prune the data which is easiest to prune away (i.e. not try to remove small files big packs but focus on packs which have lots of garbage). In general, there's an unavoidable trade-off between creating many small packs (harder on metadata throughout the system, inside restic and on the backing store but more efficient to prune) versus creating big packs which are more easy on the metadata but might create big repack cost. I guess a bit more intelligent repacking could avoid some of that cost by packing stuff together that might be more likely to get pruned together.
- verytrivial 3y ago(caveat: I may be running on old tarsnap company info but) I must say, the ONLY thing that has ever made me shy away from seriously using tarsnap was the prospect of an unexpected Colin Percival outage. i.e. key person risk. I'm guessing I'm not alone in this.
- saalweachter 3y agoI mean, if you are on HN, you will probably learn of a Colin outage within 24 hours, so practically speaking you would really only have a problem if your primary data storage, Tarsnap, and Colin all failed in the same 24 hour window or so before you had time to switch to a new backup provider.
- koolba 3y agoPretty sure his brother works on tarsnap too. They should take separate buses to ______.
- cperciva 3y agoPretty sure his brother works on tarsnap too. Yes, I hired him in 2015 IIRC. If you look at tarsnap's GitHub you'll see a lot of commits from gperciva.
- koolba 3y agoNice. Being able to work with your family is great.
- pkx166h_ 3y agoOh! Do say Hi to Graham for me. He mentored me so that I was able to contribute and eventually help maintain and manage LilyPond's Documentation and Patch Testing in a meaningful and rewarding way - all without any programming experience.
- bombcar 3y agoI would never consider a backup provider to be more reliable than that, because if you depend on it, it will fail you at the hardest time. Better to have multiple layers of backup, of which tarsnap and friends are only one, and verify regularly.
- dinner 3y ago[flagged]
- defrost 3y ago> This post-mortem just lists mistake after mistake, but gives no indication as to what the maintainer will do to prevent this in the future. Each to their own - I myself wouldn't expect that from a comprehensive "what didn't go smoothly" list such as this. Clearly Colin is aware of every point listed and no doubt is already mentally dot pointing procedural changes and additional guard rails to ease recovery in future outages and to ensure no data is lost (which appears to be the primary goal here).
- mst 3y agoThere are multiple comments in the post-mortem about what should - in hindsight - have been done instead and I think it's fair to expect that those things -will- get done reasonably soon. Pretty much all ops problems come down to the interaction of multiple mistakes that hadn't previously been an issue - GCP and AWS post-mortems tend to show exactly that, although usually with somewhat less detail. So I'd expect that any equivalent service has a similar number of gremlins hiding in their infrastructure and procedures, and I'd suggest to anybody reading this that a 43 minute old account that was created just to post the comment I'm replying to is perhaps not the most reliable judge of competency or otherwise on the part of M. Percival.
- duckmysick 3y ago> It was an honest post-mortem that revealed far too much incompetency to trust this service. That's why post-mortems are heavily sanitized. Or not posted publicly.
- jacquesm 3y agoWhat you see is someone who is actually willing and able to learn from any mistakes during this outage, no matter how small. That degree of attention to detail is exactly what I would expect from Colin. Novelty accounts created with the express purpose of slinging crap however are the equivalent of heckling in a theater, they don't contribute and in this case seem to be motivated by malice. I do tech DD for a living and pretty much every company could do better if and when something goes wrong, but rarely do companies extract the maximum of learnings from an outage. That is what should impress you rather than to perceive it as a negative. Note that most companies don't make any information about outages public and note that if and when they do it is usually heavily manipulated to make them look good. Colin could have easily done the same thing and the fact that he didn't deserves your respect, not your scorn. Consider the fact that even the best make mistakes. I'm aware of a very big name company that lost a ton of customer data through an interesting series of mishaps that all started with a routine test and not a peep on their website or in the media. Tens of thousands of people and hundreds of customers affected. And yet, you probably would trust them with your data precisely because they are not as honest as Tarsnap.
- abiro 3y ago> The second step failed almost immediately, with an error telling me that a replayed log entry was recording data belonging to a machine which didn't exist. This provoked some head-scratching until I realized that this was introduced by some code I wrote in 2014: Occasionally Tarsnap users need to move a machine between accounts, and I handle this storing a new "machine registration" log entry and deleting the previous one Recommend writing a TLA+ model to catch stuff like this
- hightrees2023 3y agoThe downtime could have been much shortened if you had properly setup and _tested_ disaster recovery steps. Create a full fledged separate staging system which you can bring down and recreate and periodically test various failure modes + document all detailed steps of system restore etc. Also I would suggest to think about the business long term and seeing if you can increase the revenue enough to enable you to hire a part-timer who can be of great help in case a similar event happens. We are also a small cloud solution provider (we focus on ML API's) and over the years it has become clear to us that when you use cloud hardware (either dedicated or virtual), from time to time the outages periodically happen. RAM, HDD or other parts of the hardware just can malfunction anytime. So this is something which 100% needs to be taken into consideration when running any high availability online service over long-term.
- zetalyrae 3y ago>The process of recovering the EC2 instance state consists of two steps: First, reading all of the metadata headers from S3; and second, "replaying" all of those operations locally. (These cannot be performed at the same time, since the use of log-structured storage means that log entries are "rewritten" to free up storage when data is deleted; log entries contain sequence numbers to allow them to be replayed in the correct order, but they must be sorted into the correct order after being retrieved before they can be replayed.) Far be it from me to tell anyone how to write software, but why build a database on top of S3 when you can just chuck the metadata into RDS with however much replication you want? The backups themselves should be in S3, but using S3 as a NoSQL append-only database seems unwise. This would benefit from being further from the metal.
- amluto 3y ago> This would benefit from being further from the metal. How, exactly, is that a good thing?
- zetalyrae 3y agoHow is not rolling your own database a good thing? Mainly because the business of tarsnap is 1) encrypted 2) backups, not building a database storage engine.
- catiopatio 3y agoIt’s cute that you think implementing client-side encrypted, deduplicated backups doesn’t involve building a database storage engine.
- catiopatio 3y agoImplementing client-side encrypted, deduplicated, snapshot-enabled backups with server-mediated access control inherently requires building a minimal storage engine to represent your opaque log-structured data. Embedding the log-structured representation of user data in Postgres would increase complexity and overhead without offering significant resiliency or recoverability advantages — in fact, quite the opposite.
- viscousviolin 3y agoUnrelated to the outage, but I'm curious nonetheless: would it be possible to hook up Tarsnap's encryption software to a Dropbox folder? I'm not sure if it even makes sense to use Tarsnap for this, but I'd love to have an easy setup that allows me to use Dropbox's servers but only let them see encrypted data so they can't snoop.
- matthiaswh 3y agoYou probably want something like https://cryptomator.org/ https://cryptomator.org/
- ivoras 3y agoDoesn't plain old Duplicity (https://duplicity.us/ https://duplicity.us/) do that already? (except for de-duplication)
- colonwqbang 3y agoWhat would be the benefit of tarsnap over using something like restic+backblaze at order(s) of magnitude lower cost? What specific need would motivate you to pay $3000 per TB-year?
- jpgvm 3y agoExtremely good deduplication means that for the core set of very important data I backup to Tarsnap the costs are negligible. I imagine the math is probably different if your data is changing more frequently. I for instance use other services to manage my video and photo libraries but my accounting databases, critical documents, etc are backed up to Tarsnap. I have been using Tarsnap for a decade and not only has there been minimal availability issues there have been almost no issues of any kind that I can recall.
- carapace 3y agoSome of us have lots of extra money and like an excuse to give some of it to cperciva so he doesn't have to work a shit job and can apply his skills and talents to bigger, better things? (People here asking about the low Bus Factor: you don't keep your backups in one service/location, eh? You use Tarsnap and Restic with Backblaze, Rsync.net, S3, etc. right? "Backups are a tax you pay for the luxury of restore.")
- idlewords 3y agoHats off to you for an honest postmortem and your capable handling of a difficult situation. The only remark I would offer is with respect to sleep deprivation—when you're the only person who can fix a problem, there's no shame in trading some additional outage time for a fresh mind. Though it feels weird to go nap when all the klaxons are blaring, problems are too easy to compound under the combination of adrenaline and inadequate sleep.
- cperciva 3y agoDon't worry, I had a couple naps in there. "This seems to be running smoothly but it will take several more hours; I'll set my alarm to wake me up in two hours and have a nap" is part of why I didn't notice the second step was unnecessarily I/O bound.
- deleted 3y ago[deleted]
- gus_massa 3y agoIIUC the process had a few steps were you only had to wait while data was transferred or processed for long times. They were probably useful to take a nap or eat or just drink more coffee.
- deathanatos 3y ago> Following my ill-defined "Tarsnap doesn't have an SLA but I'll give people credits for outages when it seems fair" policy, on 2023-07-13 (after some dust settled and I caught up on some sleep) I credited everyone's Tarsnap accounts with 50% of a month's storage costs. This speaks volumes to me about what kind of person Percival is; that credit would appear to be generously on the "make customer whole" side of the fence, and unlike the major cloud providers, he didn't make each customer come and individually grovel for it. And a clearly written, technical, detailed PM, too. This is how it ought to be done, and done everywhere. Thanks for being a beacon of light in the dark.
- rsync 3y ago"Thanks for being a beacon of light in the dark." That's well put. It makes me very happy to live in a world where tarsnap exists and is priced in picodollars.
- cperciva 3y agoFor the record, I'm happy to live in a world where rsync.net exists. I've pointed quite a few customers in your direction over the years, when tarsnap hasn't been suitable for their needs for a variety of reasons.
- jpgvm 3y agoThey make a good pairing. I backup my ZFS NAS to rsync.net for all my media and Tarsnap for all my documents/critical things.
- cl3misch 3y agoI am only using rsync.net at the moment, more specifically with the discounted "borg" mode without an explicit full shell. Your comment sounds like tarsnap is more secure (in terms of longevity) than rsync.net. Is this true? If yes, why? Genuine question, because I'm using rsync.net for my critical stuff and would gladly move to tarsnap if appropriate.
- akashshah87 3y agoUnfortunately, looks like https://www.tarsnap.com/infrastructure.html https://www.tarsnap.com/infrastructure.html will have to be updated. >> So far such an outage has never occurred; but over time Tarsnap will become more tolerant of failures in order to minimize the probability that such an outage occurs in the future.
- mherrmann 3y agoIt sounds like most of the 26h downtime was spent restoring backups. Incidentally, this is exactly the reason why Tarsnap is unusable for me for production environments. Backup restoration (as a user) is excruciatingly slow. When my systems are offline, I have no patience to wait for hours for my backup service. Maybe things are better now; Last I tried was a few years ago when Tarsnap took on the order of magnitude of one hour to restore a backup of a few GBs.