32 ms·
The absolute worst scenario happened
- WJW 6y agoThat is an amazing story. Who knew that putting everything in a database without backups was risky?
- brailsafe 6y agoDidn't they only say that their backups are fucked, but not that they didn't try to make them? I've definitely run into situations where the business is only willing to pay for the most rudimentary conception of a backup strategy, and then it turned out to have been subject to corruption.
- pedrovhb 6y agoCan't find it now, but I've often seen a phrase to the effect of "If you don't regularly test that you can restore from backups, then you don't actually have backups", and it seems to ring quite true here.
- arethuza 6y agoI'd go one step further - actually restore the databases and check that you can bring up the relevant applications that use those databases.
- regularfry 6y agoIt's up there with "unless you test power supply failover, don't rely on that generator to save you". I know of one case where power failed, the "24 hour" generators kicked in, power company says "We'll need 12 hours" and the outage happened 6 hours later when they ran out of oil.
- arethuza 6y agoI contracted at one place for a bit where they shut everything down in each of their data centres once a year and power everything back up. When I was there this didn't go too well and they couldn't get one their data centres online again - failover to their other centres did work though. This was ~20 years ago in the finance sector.
- grumple 6y agoThat sounds like their testing worked out for them. Better than a random failure.
- myself248 6y agoYup. Have the problem when all the right people are awake and on-site to handle it. I was in a building when someone inadvertently powered off the wrong equipment, which had been running for several years, and several of the power supplies failed to come back up. It was 1+1 redundant though, so we could quickly shuffle packs around to bring it back up without redundancy. Then, jogging through the building and asking if anyone had spares, we found a field tech in the lunchroom who had a pile of stuff in his van. Whole thing was back to 100% in less than an hour, and we let the beancounters sort out the field spares being used for office equipment. If that same failure had happened during the overnight maintenance window (when volatile work was supposed to be performed), there certainly wouldn't have been the same resources around.
- ewindal 6y agoBackups were extant, but inaccessible, and non-functional due to causing a kernel panic when applied. [1] [1] https://www.reddit.com/r/sysadmin/comments/ma4mwl/the_absolute_worst_case_scenario_happened_what/grqlh5q/ https://www.reddit.com/r/sysadmin/comments/ma4mwl/the_absolu...
- M2Ys4U 6y agoThey said they knew they haven't been able to restore backups for 2 years. Which is another way of saying that they don't have backups at all.
- brailsafe 6y agoAh, didn't spot that
- PeterisP 6y agoThey explicitly mention that they have backups, but that the "backups are fucked", whatever that means. Ransomware scenarios come to mind, where the attackers will look for and actively destroy the backups you have if they are remotely accessible with e.g. AD admin credentials, and only trigger the ransomware when the backups are gone. If you have offsite backups in some cloud system, that helps you against natural disasters but not against malicious activity. The need for backups that are not just offsite but also offline is somewhat recent and very, very many companies do not have them.
- fuzzy2 6y agoThey backups are probably fine (as in “the data is in there”) but there may be no plans on how to restore them. Or they may be incomplete. I doubt any external factors are to blame.
- Nextgrid 6y agoThis doesn’t sound that bad. Unless I’m misunderstanding something, the ERP was used as a source of truth for DNS records, and that mechanism broke. There’s no indication of actual data loss so it should still be in the ERP DB, and it’s a matter of reverse-engineering how it was all put together and then extract the data into a zone file (or into a managed DNS service such as Route 53) at least for the core domains which should at least bring their internal services back online and allow them to proceed further.
- arethuza 6y agoI guess if you are in the business of selling services that include custom domains that isn't quite as crazy as it initially sounds... You are, of course, assuming they have a backup and the backup can actually be restored and that the restore contains the required information.
- WJW 6y agoIt does say in the original post that: > Our backups are fucked. So I assume that they have already tried and they didn't work.
- arethuza 6y agoEarly in my career (>30 years) ago I nuked a companies salary database (a missing $ in a shell script moved everything to the file i) No problem they said - we have backups. They had three tapes. First tape failed. Second tape failed. Third tape worked. In retrospect that was a useful, if rather stressful, lesson.
- lupire 6y ago> nuked a companies salary database > a missing $ Bravo!
- artursapek 6y ago>a missing $ in a shell script moved everything to the file i I've always had anxiety when running a shell script, this is a perfect example of why
- IgorPartola 6y agoWhat exactly happens if you break a three month resignation period? I don’t advocate doing this in generally but as a hypothetical exercise if this person just stopped showing up for work, what exactly could the company do?
- brailsafe 6y agoWas curious about that as well. I've never heard of such an obsurd amount of time, and it seems doubtful it would hold up. "You need to give us 3 months heads up no matter what, but don't worry, we don't have to give you any if you oversleep for a meeting."
- nerbert 6y agoThis isn’t particularly crazy in Europe. It indeed goes both ways.
- ratww 6y agoIn my experience it's also the same in Latin America, but the notice period is normally of 1-2 months rather than 3-6 months.
- arethuza 6y agoI had a 6 month notice period in one role, 3 months isn't uncommon for senior roles, in the UK at least.
- xchaotic 6y agowhile I agree that this is common in Europe, if the company ceases to exist, what good is the 3 months notice for the employee? If he/she has any options, I'd leave asap and not worry about amicable terms with a soon to be non-existent entity.
- arethuza 6y agoFair point - in that case I suspect that would be highly dependent on whether there are any assets remaining in the company and how the local legal system treats employees as creditors. Mind you - if the company is going down the tubes they are highly unlikely to enforce any job contracts.
- aidos 6y agoI’ve seen a few data disasters in my time but I think this takes the cake. It’s a good reminder of why you just need to prioritise the sort of work that has no immediate payoff.
- michaelhah 6y agoHow many people had to log into their own work systems and check that things were running in order to be sure it wasn’t their employer?
- HourglassFR 6y agoNah, I'm not saying my employer's IT infra is rock solid but at the very least we have backups.
- loloquwowndueo 6y agoI know it wasn’t mine because we still have one guy who knows how everything works :) (unless he quit over the weekend oh crap)
- tclancy 6y ago“let the boat sync.” Sure, now you start thinking of backups.
- atmosx 6y agoI loled, although non-native speaker who does similar mistakes all the time :-)
- sombremesa 6y agoLoathe to be "that guy" but a sync is not a backup.
- le-mark 6y agoIt’s easy to imagine there are a lot of these type of time bombs out there, particularly in really old legacy systems (> 20 years for example). I was at one place where even building the application for deployment led to a one week outage when some core people were laid off.
- protomyth 6y agoInterestingly, some of these legacy systems run on things like IBM i (AS/400) which have an easy backup story and vendor support for recovery. I can have a new i machine receiving a backup and deployed fairly quickly (day or two) and I'm out in a rural area.
- abruzzi 6y agoUp until 5 years ago, my employer's primary system (government, property tax assessment, billing, and collection) was from a company that had gone belly-up in the early 90's. We fortunately had the source, but it ran on a VERY legacy database called Unidata because it had originally be developed on the Pick operating system. Unidata was/is a Multivalue Database with a built in BASIC programming language. By the time 2014 rolled around there was one programmer in the organization that understood the system, and he was responsible for all code changes. I half-joked with management that we needed an insurance policy on his life/ability to work, because if we lost him, it would put us in a very difficult position, and that software accounted for 90% of our revenue (as well as a lot of the revenue that went to all the local towns and school districts for their operations.) That employee was also nearing retirement (he is retired now.) Fortunately we were able to successfully migrate off of that application, though we keep a minimal license pool to the old system so people can validate pre-migration data against post migration data.
- tsomctl 6y agoimproved link: https://old.reddit.com/r/sysadmin/comments/ma4mwl/the_absolute_worst_case_scenario_happened_what/ https://old.reddit.com/r/sysadmin/comments/ma4mwl/the_absolu...
- benlumen 6y agoThanks. Reddit is such a good example of how "progress" hasn't really been progress in web development. The new front-end framework site is so much slower and more buggy.
- xchaotic 6y agowow, it was so much better before. What the heck? Who made the decision to switch?
- mkl 6y agoThey want people using their app, for some reason I don't understand. Maybe to make ad-blocking harder by getting people off computers where it's easy (most mobile browsers don't support plugins, so it can't be about them).
- the_only_law 6y agoIronically, every time I click their “open in app” link it’s drips me straight to the App Store despite me having the app installed.
- perlgeek 6y agoThe old design was optimized for usability, the new is optimized for ad revenue and/or for pushing people on the app.
- fsflover 6y agoEven more improved link: https://teddit.net/r/sysadmin/comments/ma4mwl/the_absolute_worst_case_scenario_happened_what/ https://teddit.net/r/sysadmin/comments/ma4mwl/the_absolute_w...
- newsbinator 6y ago> we don't even have a listing of our ~300 employees SMS numbers
- elif 6y agoThis surprised me. If ever there was a use case for an organizational structure, this is surely it.. just have every manager reach out to their direct reports? Following the tree communication should happen at something like nlogn speed
- teddyh 6y agoThey might check on robtex.com and other DNS historical information sites.
- mperham 6y agoTechnical debt can lead to non-technical bankruptcy.
- gillesjacobs 6y agoThe managed service provider company built a custom DNS inside an ERP without backups or bootstrap strategy in place. House-of-cards core infrastructure is a sign of negligence and OP should have reconsidered working there a long time ago.
- hinkley 6y agoDivision of responsibility can make this sort of thing very difficult to know. If you aren’t running drills then transparency may not be in everybody’s self interest and someone somewhere is covering up for the fact that they only half know what they’re doing.
- sorokod 6y ago"Every person who knew how the custom DNS system worked has left the company years ago." Keeping the people who know how the system works is really the last line of defense. From, the point where such last person left without any compensating mechanism in place, the clock started ticking
- hinkley 6y agoSomewhere right now there is a disgruntled former employee drinking a glass of Scotch and smiling.
- ChrisMarshallNY 6y agoThis is the snow on the tip of the iceberg. I'm sure there's plenty more where this came from. There's so many things wrong with the scenario described, that I can't even begin to talk about it. My experience is that folks never want to talk about backup or DR, as they are expensive, and the idea is that they should never be used. No one ever has a problem with insurance premiums, though...
- hinkley 6y agoI keep wondering if there’s an underwriter angle here where you offer insurance and audit the company to set the premiums based on how broken their IT situation is.
- scandox 6y agoHow people handle this level of disaster is always very revealing. In every case I've ever experienced someone kept their head and a solution was found. Usually there was reputational damage but surprisingly most times the actual business bounced back. But it was always the actions of someone in particular in those first hours that made the difference. Thinking you're fucked is unhelpful, because you start talking about resumes and expecting some deus ex machina to appear and resolve the situation (for good or ill).
- colmvp 6y agoThinking you're fucked is unhelpful, but it's pretty normal to go full Alien Hudson in a dire situation and need someone else to assure you that things will be okay.
- 4e530344963049 6y agoIf you have been that person before, it can be fun and exciting to think up solutions and implement them under pressure. Definitely a real test in your ability to concentrate. But don't expect the credit you deserve for it, and don't even expect others to even bother to put in any effort towards fixing things. Also, make sure everyone knows you are going on vacation once you resolve things.
- paulz_ 6y agoThese types of situations always wind up being the most fun I have at work and leave me feeling most fulfilled. I wish I could do that all day instead of going to standups.
- 4e530344963049 6y agoA "firefighting" consultancy might be fun, no? And you would give people a lot of time off!
- paulz_ 6y agoI've thought about that before. You get the call. "It's 30 minutes away. I'll be there in 10" If anyone happens to know of a career path or company that does that sort of work I would be interested to hear about it. Bonus points if the pay is half decent.
- xchaotic 6y agohas anyone figured out who the company is by now? Surely there must be some reports of the outage by now?
- viraptor 6y agoThere's a small chance it's just a made up story for fun. If not, some outlets should pick it up soon.
- the_duke 6y agoNot necessarily, there are many small companies in this space. But knowing Reddit, I'd give this a solid 50/50 chance of being fake.
- avaldeso 6y ago> Surely there must be some reports of the outage by now? Not everyone is AWS. There's a lot of obscure software providers nobody cares when they're down.
- protomyth 6y agoWell, it looks like someone better track down the people who knew how it worked and offer some serious cash. The only actual backup is one that has been tested. With the cost of hardware, buy 2 or have an agreement with your vendor to have a machine ready to receive a backup.
- sethammons 6y ago> The only actual backup is one that has [recently] been tested
- erikstarck 6y agoThis reminds me of when the Danish mega-corp Maersk was attacked by ransomware and had all of their computers locked and decrypted with no access. All computers except one, which due to a power failure had been offline the whole time. This computer had the critical DNS information needed to restore the network. Only problem: it was in Africa, in a country that required Visa which would take weeks to apply for and receive. So someone from the African office had to bring the hard drive to an airport that someone from the London office could travel to and pick up the critical hard drive, then bring it on a plane back to London. A sweaty and nervous trip, I can imagine. Whole story here: https://www.wired.com/story/notpetya-cyberattack-ukraine-russia-code-crashed-the-world/ https://www.wired.com/story/notpetya-cyberattack-ukraine-rus...
- driton 6y agoSounds very interesting, but the story seems to be behind a paywall. I found some other articles[0][1] related to the incident, however none of them seem to mention the flying of a hard drive across continents. 0: https://www.zdnet.com/article/ransomware-the-key-lesson-maersk-learned-from-battling-the-notpetya-attack/ https://www.zdnet.com/article/ransomware-the-key-lesson-maer... 1: https://portswigger.net/daily-swig/when-the-screens-went-black-how-notpetya-taught-maersk-to-rely-on-resilience-not-luck-to-mitigate-future-cyber-attacks https://portswigger.net/daily-swig/when-the-screens-went-bla...
- erikstarck 6y agoSorry about the paywalled article. Here’s a summary of the same article: https://www.chrislouie.net/blog/2018/9/10/better-to-be-lucky-than-good-after-not-petya-shipping-company-maersk-saved-by-power-outage https://www.chrislouie.net/blog/2018/9/10/better-to-be-lucky...
- Eremotherium 6y agoOr if you want the story as a podcast I can wholeheartedly recommend this episode (and the podcast in its entirety) of Darknet Diaries: https://darknetdiaries.com/episode/54/ https://darknetdiaries.com/episode/54/
- viraptor 6y agoThese kind of posts are awesome tabletop scenarios. (and fuel for https://twitter.com/badthingsdaily https://twitter.com/badthingsdaily ) Coming up with potential fixes is a good exercise. Things I would do that I haven't seen mentioned yet: - Grab images from the relevant servers to have current backups before messing with anything. - Try tcpdump-ing and decoding the custom KV protocol - other databases are complicated, but things like memcached/redis/... have trivial wire protocol. Maybe it's possible to recover some data by listing all entries. - Run `strings` on the KV binary or `file` on the data to make sure it's not just some open source project with a proprietary interface. With any luck it's just something like bdb that can be accessed externally. - Hook up a debugger to the KV and try to recover the keys (or maybe just strace to see if they keys are in any files) - Start tcpdump on the old DNS IPs. They mentioned they have no idea which domains they own anymore - as long as NS on customer domains wasn't trashed as well, they could collect queries coming in and recover a list from there. - See if the customer domains can be collected from billing data / some email store of notifications. (as much as I try to design for no disasters in production, the rush of "this is so FUBAR, what insane thing I can do to fix it" for me is an amazing feeling :) )
- flaxton 6y agoCan’t you just pull the drive from the failed server, extract the DNS entries and build a new DNS server? It doesn’t sound like everything is broken, just that it is failing because the DNS server isn’t answering. Just install Linux on a new box and set up a DNS server on it. It could be done in a day I would expect. Or am I missing something?
- Kranar 6y agoIts encrypted and the encryption keys are lost.
- capableweb 6y agoIt would work if they knew what the DNS entries were but seems they lost all DNS records and don't even know what records they have/had, so spinning up a new DNS server won't help as they don't know what it should serve. > All domain information was wiped out and records became null [...] Our records are wiped from all domain servers out there [...] We don't even know what domains we own, the listing was hosted in the ERP which is now busted
- drummojg 6y agoThat's part of what bothers me about this whole story--it says they were running BIND, which configs on text files for goodness sake. These critical records were tiny and could fit anywhere & transfer in the blink of an eye. That such a simple thing is buried under & dependent upon an entire complex and untested/maintained DR plan is mind-boggling.
- capableweb 6y ago> That such a simple thing is buried under & dependent upon an entire complex and untested/maintained DR plan is mind-boggling I find this to pretty common in many setups where the engineers don't focus on simplicity and removing layers of abstraction. If the workforce is young and inexperienced, over-engineering tends to happen everywhere and you end with situations like this. I have seen worse in my years in the industry, that's for sure.
- PaulHoule 6y agoYou should sue the LTO tape conspiracy that means you need to spend $3000 to back up the contents of a $300 hard drive.
- arnaudsm 6y agoAre the CTO and sysadmins liable in such scenarios? Can they be sued ?
- viraptor 6y agoAnyone can be sued. But it wouldn't be a good idea if they kept the receipts: https://www.reddit.com/r/sysadmin/comments/ma4mwl/the_absolute_worst_case_scenario_happened_what/grqkzus/ https://www.reddit.com/r/sysadmin/comments/ma4mwl/the_absolu... > We tried really hard to make management aware and actually succeeded in that, only for management to follow up and say we don't have enough money to fix it, and they understood the possible outcomes.
- ReptileMan 6y agoNo keyboard found, press f1 to continue type of situation.
- ExcavateGrandMa 6y agoThere is always a solution... but don't redo the same errors :)
- tw04 6y agoI actually don't see any issue that isn't simply a matter of money. Everything they list that's an issue seems to revolve around the fact that the employees with domain knowledge for the system have left the company. Reach out to them and hire them on as contractors. If they left under bad terms because the business was a bunch of dicks, expect to pay 10x market rate. If this is truly "fix this or the business is out of business" - then it shouldn't be a tough decision to make.
- blunte 6y agoYes. And after the boat is righted, fire every exec from IT Director upward through CEO. That last part may not be necessary though, as customers will jump ship and the company will probably sink anyway.
- Nextgrid 6y ago> customers will jump ship and the company will probably sink anyway. Depends, they might have a captive market or a customer base that's non technical and doesn't understand how bad this is. If their systems are like this, chances are their offering isn't anything groundbreaking and the competition already provides a better service for possibly even cheaper, so the customers who are willing to jump ship would've done so long ago and the fact they're still in business suggests they have a complacent customer base that's likely to stick to them even despite this incident. The same reason this technical debt was left unchecked for ages applies to customers. The task of migrating to another provider will most likely rot forever in a Jira board somewhere and they will keep paying their bill in the meantime.
- viraptor 6y agoThere's the possibility that those employees are happily retired and don't care. Or they're just not available in many ways. I know of a hospital which is still refusing to migrate off of a system which is not supported for a decade. Pay more money? 3 people worked on it: Adam had a stroke, Bob retired with enough money, Charlie left the country.
- 6y ago
- alistairSH 6y agoTangent Alert! I have a 3 months resignation period, so I can't even leave... Maybe it's my (US) American experience, but being unable to quit a job strikes me as a terrible policy. This guy(gal?) is really required to submit 3 months notice or face some sort of repercussions (financial, I assume)? Does this 3 month period apply in reverse (employer can't fire employees without a full 3 month warning)? Not that the American system (no cause firing/quitting with no notice) is perfect, but 3 months? Yikes.
- polote 6y ago> Does this 3 month period apply in reverse Yes (in France at least) it is on both sides, unless both sides agree to shorten it > face some sort of repercussions (financial, I assume)? I'm speaking for France. This is a tricky subject, you almost can't get any repercussion. But it can sucks, like the employer can refuse to fire you and suspend your contract, in the mean time, you are not paid but are not allowed to work for someone else.
- CaptainZapp 6y ago> Does this 3 month period apply in reverse (employer can't fire employees without a full 3 month warning)? Of course. Different notice periods for employers and employees would be outright illegal in most of Europe.
- Symbiote 6y agoIt's fairly common to have an X-month notice period to resign, and a 2X or X+1-month period to be made redundant.
- detaro 6y ago> Different notice periods for employers and employees would be outright illegal in most of Europe. AFAIK more common is "lower notice period for employers is illegal".
- CaptainZapp 6y agoI stand corrected. See my comment below.
- deleted 6y ago[deleted]
- mysterydip 6y agoA good reminder for everyone to check your backups work, and if you don't have any, start!
- hinkley 6y agoQA departments used to be good for this. Setting up a test cluster using real backups can be good. If you have the crosstalk problem sorted out that is.
- hef19898 6y agoAnd that is why I choose to host "my company" at the big names: homepage directly with Wordpress instead of a smaller local provider, office directly with MS and so on. Because, even there might be cheaper solutions, at least I can be sure to have a huge org in place to make sure stuff is up. And to find ways to get stuff back up and running in case something goes south. I value that peace of mind a lot.
- D-Coder 6y agoInteresting. I choose to keep a copy of everything on my home machine (it's only personal stuff, not a company), plus backups to two removable disks, one in a fire-resistant box. Because I don't trust a big company to decide that my account is bad, or something, and I can't ever get my data back.
- airhead969 6y ago50% of companies who lose all of their data go out of business within 6 months. They clearly failed to test their backups with regularly-tested restores. And, a separate alternative system should've had multiple revisions of critical pieces of information like customers, assets, inventory, bank accounts, debts, and employees. A data recovery shop could attempt to forensically-reconstruct SMS numbers. Look for logs and other places critical deets might be found. Try your best to get things back and make alternatives processes to piece together what you can. A phone carrier or SMS send provider might have a log of phone numbers. The paycheck company may even have the address of everyone. Send them a $0.01 paycheck with a note on them to call in. I would ride it out for a paycheck because it's not like it could get worse than going under. Just be sure to get paid.
- lopatin 6y agoThey didn't fail to test their backups. They tested them, they didn't work, and then they said "this is fine".
- airhead969 6y agoOh my, I missed that, thanks. I can envision the dog in the burning house meme. Ouch!
- ghaff 6y agoReaching people seems like the least of their problems to be honest. If they use a payroll service, they'll have addresses--don't know about Europe but does just about any company do their own payroll in the US at this point? In any case, I assume most people know how to reach most of their direct reports and many of their co-workers. Especially at a fairly small company, I would think you could reconstruct contact info on employees fairly quickly. Of course, that may not help with the bigger problem.
- airhead969 6y agoYep. It's going to be a massive problem figuring out what's what, who owns what, and who owes what. I don't even see how they'll keep enough customers to stay afloat.
- TheRealDunkirk 6y agoOne of the comments in the thread said, simply, "Prepare 3 envelopes." Being a 25-year sysadmin, immersed in thinking about the problem, it caught me totally off-guard. I don't think I've ever seen a more perfect response in a Reddit thread.
- etripe 6y agoWhich 3 envelopes would those be?
- TheRealDunkirk 6y agoIt's an older reference, which is why it was so funny to me. I hadn't seen it in a long time. GIYF.
- lmkg 6y agoIt's an old story. New executive starts the job. On his desk are three envelopes, left by his predecessor, with a post it note saying "When shit hits the fan, and you don't know what to, open one of the envelopes." Yadda yadda, exposition, things are normal and exec is successful for a year or two and then things go pear-shaped. Completely out of options, he opens the envelope, finds a letter which says "Blame me." He blames his predecessor for leaving the state of things in such a poor shape that this misfortunate was unavoidable, everything blows over, his success continues. Blah blah, narration. Another shitstorm. He opens the second envelope, finds a letter which says "fire all the managers." He promises a significant shake-up, says he has the leadership to turn this around, purges the old guard. Things go well. Hurly-burly story, and so forth. Catastrophe strikes. Having blamed his predecessor and fired everyone and installed his own managers, he has no one else to turn to for casing the blame. He opens the last envelope, hoping to find salvation. "Write three letters."
- unnouinceput 6y agoMy recollection of 2nd letter is "blame it on current economy situation"
- simonebrunozzi 6y agoBest thing to do for this poor sysadmin, before all else: talk to a lawyer, understand what you should do from now on to minimize risk of being sued by company/shareholders, etc.
- comeonseriously 6y agoChoice 1: Leave. Just leave. Choice 2: Put your head down and fix this. Your worth (to another company because you really should leave after you fix this) will increase dramatically.
- shireboy 6y agoOn a personal level, this may be a little self centered, but I use cases like this to put my own problems in perspective. Recently I had to troubleshoot lots of IT issues during a winter storm. It was a bad outage for the organization during high load important scenario. I distinctly remember thinking during the thick of it “this is bad, but imagine being the poor engineers in TX responsible for the power grid.” This helped me not panic and focus on the problem at hand. I feel like this ability to step back and not take a problem more seriously than necessary can be an asset. In this case, it’s pretty bad, but “at least people aren’t in immediate danger of their life.” Think it to yourself, but probably not best to share with others until after the problem is mitigated ;)
- devit 6y agoDefinitely not the "absolute worst", they merely lost all DNS records.
- iJohnDoe 6y agoI’m guessing the IT department reported to the CFO. Bean counters have zero IT knowledge but like to think they do. They also by default like to say no to any expenditures. Combine the two and you have a disaster waiting to happen. IT reporting to finance/CFO is always the worst decision.
- undefined1 6y agogood idea in that thread: If you don't have a backup of your zone files, a good way to quickly pick up the main domains being used would be to turn on logging for your DNS server and output into a log collector. That way you can quickly build queries on which IPs are asking for which FQDNs and start rebuilding a list of your lost zones and can maybe do some guess-work on which IPs those FQDNs were likely to resolve to. I would expect irregular or extremely infrequent processes that use specific FQDNs will pop up from time to time as errors/failed processes and should prime IT teams on what to look out for. https://www.reddit.com/r/sysadmin/comments/ma4mwl/the_absolute_worst_case_scenario_happened_what/grr3tgo/ https://www.reddit.com/r/sysadmin/comments/ma4mwl/the_absolu...
- simonw 6y agoSo many beautiful details if you search for more comments by the original author of the post: "The company that built this ERP solution went bankrupt 6 years ago, and we don't have the source code. It uses PostgreSQL and MySQL, and also has a built-in key-value database for which none of us has any credentials. We don't have the source code for this software (it's built in C and delivered to us as compiled packages..)."
- matwood 6y agoExtremely poor planning/lawyering on the companies part. When larger entities deal with smaller entities, it's common to have a source code escrow as part of the agreement which grants the purchasing company rights to the source code in the event of this exact situation. Obviously we all know that source code alone isn't enough, but it is something to start from. If the ERP solution was not SaaS and installed on site, the purchasing companies IT should have gotten credentials as part of the contract. Again, normal type stuff when buying software from smaller companies that have somewhat higher risk of going out of business.
- deleted 6y ago[deleted]
- gist 6y ago> So, everything is fucked. We have (had) a custom DNS system built into our custom ERP software (don't ask). It had an integration to our old (bind) nameservers, which was stragith up awful and we've been trying to replace it for years. Well, on Friday it all broke down. All domain information was wiped out and records became null. Our company's domain is down, as well as our customer domains (we're a medium-sized MSP). Every person who knew how the custom DNS system worked has left the company years ago. Unclear why HN wastes time on this type of 'story' with no attribution and scant details. (I have expertise in this area let's say). To start this implies that they have not added a customer or made a change in several years by this statement: "Every person who knew how the custom DNS system worked has left the company years ago." This also contradicts: "and we've been trying to replace it for years"
- blhack 6y agoIt's dark, but I kindof love situations like this (although luckily I haven't had something like this happen to me in digital space since I was a teenager). Some places I would look: 1) Are you SURE that everything is actually gone? Is there a logfile, a billing trail, sent emails trail from your email provider, something like that? If it was me I'd be looking for a secondary system (like alerts, logfiles, billing, etc.) that had some of this information stored in it as secondary function, and writing some scripts to parse those things and start rebuilding what can be rebuilt. 2) You said you lost all the emails and SMS of your team. Do you have an HR department? Surely they have this information. Your team is likely going to start trying to send emails to each other about this. Can you dump logs off of your mail server to rebuild at least the naming schema? For phone numbers: can you try an old school phone tree style? Start with somebody you know, and ask them for everybody that THEY know and so on. 3) Was there a staging environment that might have had some of the data you're missing on it? 4) Are you really really sure that everything was wiped out? Is it possible that your ERP system is broken (lost its DB connection or something?) and that caused a cascading failure that makes it LOOK like everything is gone? 5) Can you get at the BIND server somehow? Are there logfiles there? I'd love to know more specifics about what "wiped out" means in this case. It seems unlikely that it's properly "wiped out" unless something malicious happened, and even then I think you could rebuild a lot of this from sortof "secondary" data. Think like a hacker, but your system is your own target. How can you steal the data you're looking for from places it isn't really "supposed" to be?
- CobrastanJorji 6y agoI get where you're coming from. It's a little bit like a real world, high stakes escape room. Come to think of it, this could be a fun online challenge. I picture a series of VM images, each of which describes the design of the system and the terrible thing that has gone wrong, and you are challenged to get it fixed and/or get the data out of it.
- hinkley 6y agoStart connecting with everyone on LinkedIn for one...
- mikewarot 6y agoAs a system administrator, I never, EVER encrypted data at rest. Stupid policy I know, but encrypted stuff tends to be unrecoverable when you need it. I had an Exchange Server that had issues, and I didn't have enough hardware to try a backup, I kept telling management... it died, and I was able to recover most, but not all of it. Shortly thereafter they outsourced IT, and I assume the new crew was able to get the resources to actually fix it.
- kodah 6y agoI'm a systems engineer and a software engineer. A couple tips that are relevant to this story: 1. Applications do not need DNS servers, ever. Hosting platforms do. The way I read this is that it's some odd form of split horizon where one DNS server seems to replicate to larger DNS servers, yet the single server is what carries most of the critical information. 2. State should be transferred from where you gather input to where it becomes actionable by an eventually consistent process. ie: If you store your customer CNAMEs in ERP, then something like webhooks to a service that maintains the state in machine readable form should verify, accept/deny, and do something with that change (like update DNS). 3. Customer systems and core/critical systems should be separate. The fact that critical contact systems for employees were not separate from 4. Have a DR plan and artificially enact it at cadence, doing so even in a sterile environment is better than nothing. Successful DR's in production are best. If the first time you live test your DR strategy is when all is lost, then all is most likely lost and at best you have a lot of hours of work ahead of you. 5. The holy grail is replayable audit frameworks. This could mean capturing diffs but it can be as simple as logging the current state of a given $thing and sending it to an off-local system.
- chmod600 6y agoThere would be a lot of value in independent organizations that can evaluate engineering practices. Non-technical businesses could just ask if you have cert XYZ before paying you. It wouldn't have to get into the weeds too much. I mean really basic stuff like finding documents and following them to restore a backup or release a new version or fix a simple bug injected in the system.
- pjmanroe 6y agoI worked for a newspaper group back in the 1990s. I did all their programming in dBase III & IV LAN. I was putting in 80+ hours a week. I wasn't getting paid squat, but then they wanted to put me on salary (at less than I was currently making) and I said no way I'd give my notice first. So that's what I ended up doing. They didn't believe me. But I left. After a few months of haggling I went back at $50 an hour on contract. I did this for a few weeks (about 25 hours a week), then started to not want to pay me. So I quit for good. Unless they made an upfront deposit. Which they wouldn't do. They ended up spending close to $250,000 on equipment and programmers and still couldn't do what I could. They were duplicating and tripling everything. Wasted so much money. Now I'm going to college for my BS in Computer Science with a specialty in Software Engineering. I'm currently running Windows 10 Pro Insiders Edition DEV and VMware with Kali Linux.
- AtlasBarfed 6y agoIf it's an ERP solution, where is the ERP support? What is the "E" in ERP without support?
- detaro 6y agoThe E obviously refers to the enterprise using it, not the "Enterprise-y-ness" of the software.
- AtlasBarfed 6y agoThe discussion on reddit said it actually was another vendor's ERP, but they "went out of business". Honestly there is so much WTF-ness to the whole stack I'm starting to suspect this is a fabrication.
- Cooper95 6y agoThanks for the update and quick reply. I'll be sure to keep an eye on this thread. Looking for the same issue. Bumped into your thread. Thanks for creating it. Looking forward for solution. https://www.tellpizzahut.one/ https://www.tellpizzahut.one/
- Cooper95 6y agoThanks for the update and quick reply. I'll be sure to keep an eye on this thread. Looking for the same issue. Bumped into your thread. Thanks for creating it. Looking forward for solution. https://www.tellpizzahut.one/ https://www.tellpizzahut.one/
- dredmorbius 6y agoTechnical debt, meet technical bankruptcy.