11 ms·
The anatomy of a 2AM mental breakdown
- mmastrac 2y ago[flagged]
- jakubmazanec 2y agoThank you for this. The article started interesting, but then devolved in boring all caps recap of panic thoughts which killed its momentum. I was still interested in the origin of the bug though.
- LipSchilling 2y agoI'm not entirely sure what you expected, given the title
- charles_f 2y agoIt's not your style ; I still found it an entertaining read that conveyed well the prolonged stress the author was feeling.
- gpvos 2y agoYou can always skip to the end if it bores you. I liked it.
- ARandomerDude 2y agoI don't think that's the right TL;DR here. The point is the FUD we sell ourselves is often just not true. Take a breath and face life's challenges without telling yourself "it's over" along the way.
- deleted 2y ago[deleted]
- deleted 2y ago[deleted]
- riiii 2y agoI heard a sleep expert say that during the night your logic and reasoning abilities are greatly reduced. I think it was in relation to dreaming, you don't want to apply much logic to that stuff. That's why trying to solve problems in the middle of the night just ends up in stress.
- SketchySeaBeast 2y agoIt also makes sense why I'll wake up in the middle of the night terrified of things that don't bother me during the day. My fears all are much more immediate in the witching hour, and I can't talk them away.
- riiii 2y agoI'm the same. The only thing that has a chance of working is to remind myself that it's the middle of the night and however terrifying this is I'll deal with it in the morning.
- Waterluvian 2y agoFunny enough, I almost exclusively do my best problem solving at night. I'll wake up with solutions, or I'll stay awake and get things done. I'd say 90% of my Master's work was done between midnight and 3am. But indeed, regardless of time of day, if I just wake up, or am woken up, I'm basically a big, dangerous toddler when it comes to problem solving. Context and nuance is important, of course. We're all so different.
- jrgoff 2y agoThat often happened for me in grad school as well. Generally the questions I had trouble with on a take home exam would yield to late night inspiration. And if they didn't yield by a semi-reasonable time, I would go to bed and many times, as I was drifting off, an insight would come to me. One memorable time though, that didn't happen and I woke up several times in the middle of the night from stress dreams where I was trying to solve the problem. And when I thought about the dream, nothing I had been doing in it made any logical sense to actually help me with a solution. Fortunately I woke up early and was able to figure it out in the morning. It was not a very restful night of sleep though.
- rockyj 2y agoAll monitoring comes at a cost and adds complexity. I wish people realized that, I struggle with this in my own team, we keep adding layers upon layer of monitoring, metrics etc.
- htrp 2y agoYou need to insulate your metrics/monitoring from the critical path so that failures in these providers don't take out your app.
- macNchz 2y agoI’ve definitely come across this genre of issue on many sites in the past when checking out the console when a site is broken. Page is just a blank white screen? Oh, looks like the render function was placed after the init for some 3rd party user monitoring, which crashed because the script didn’t load properly. “Complete Checkout” button just does nothing at all? Oh, looks like the code to take my money runs in a callback to some analytics script that my ad blocker blocked. Oops.
- ejs 2y agoI had these issues before for plenty of things, it just hurts the most when it's something non-essential. I've had outtages because silly system updates to slack broke and took things down. I run metrics and such out through logs these days because UDP don't care.
- ssiddharth 2y agoHa, that was a stressful yet funny read. The self flagellation bit hits too close to home though. I run a somewhat successful iOS/MacOS app and pushed a release that completely broke about 350k+ installations. Not entirely my fault but doesn't matter as it's my product. The cold sweats and shame I felt, man... Plus it's on the App Store so there's the review process to deal with which extends the timeline for a fix. Thankfully, they picked it up for review 30 minutes after submission and approved it in a few minutes.
- cqqxo4zV46cp 2y agoMy first employer as a developer, due to their incompetence, not intelligence, ‘let’ me break our customers’ shit from a young age. As I’ve progressed through my career, and through my transition to leadership, I’ve realised that it was a very valuable experience. I may have stressed about it in the past, but those memories are too distant for me to even reach now. I certainly don’t stress about it now. I’ll, maybe controversially, sometimes allow my (early-career) team members to break prod, if I can see it happening ahead of time, but am confident that we’ll be able to recover quickly. It’s common knowledge that being given room to fail is important. But many leaders draw the line at failures that actually hit customers. If one finds oneself in the very common and very fortunate position to be building software that isn’t in charge of landing planes or anything similarly important, they should definitely let their team experience prod failures, even if it’s at the expense of so-and-so from Spokane Washington not being able to use their product for a few minutes.
- mynameisvlad 2y agoWithin my first 6 months at a FAANG, I accidentally took down our service for half the world for 30-60 minutes. That was honestly one of the best things that could’ve happened to me and I still use the story for new hires to this day. It’s a very humbling reminder that nobody is perfect and we all make mistakes. And I’m still here, so it’s never the end of the world.
- sandspar 2y agoNice writing. Art can be useful for helping us cope.
- curiousllama 2y agoLove the detailed emotional reaction to scrambling to fix an outage. Nothing quite like attempting calm, dispassionate debugging while actively wrestling your own brain
- tempfile 2y agoEvery external service you integrate is adding a small, non-zero, compounding probability of finding yourself in exactly this situation.
- datavirtue 2y agoNot to mention the performance hit to your actual customers when it all "works."
- shadowgovt 2y agoThis also serves as a cautionary tale to small-business web people. You can start a web service business solo (or with a small handful of folks). But the web doesn't shut down overnight, so either have a plan to get 24-hour support onboarded early or accept that you're going to lose a lot of sleep. (And if you think that's fun, wait until you trip over a regulatory hurdle and you get to come out of that 2AM code-bash to a meeting with some federal or state agent at 9AM...)
- JohnMakin 2y agoWorking as a SRE for a year in a large global company broke me out of this "panic" mode described in this post. To a business, every problem seems like a world-ending event. It's very easy to give in to panic in those situations. However, in reality, it's rarely that bad, and even if it is, you'll probably survive without harm. The key in these situations, and what I try to do (totally relate to breaking out in a sweat, that still happens to me, just happened yesterday) is to take 5-10 minutes before doing any action to try to fix it and sketch it out, think about it as clearly as you can. Fear interferes with your ability to reason rationally. Mashing buttons in a panic can make your problems spiral even worse (seen that happen). Disrupting that fear circuit any way you can is important. Splashing my face and hands with extremely cold water is my trick. Then after you go through a few of these, you'll realize it really isn't too bad and you've dealt with bad situations before and you'll gain the confidence to know you can deal with it, even when there's no one you can reach out to for help.
- H8crilA 2y agoAnd ultimately you carry none of the risk. It's not your company, and the company can (and will) cut you off at a random time. Unless it is your company :)
- minkles 2y agoThis is the entire risk: being fired and financially in trouble. I spent 5 years eliminating that risk. Now I don't give a single fuck. There is no fear. They get better, more rational work and I have security.
- toomuchtodo 2y agoWas fired, worked out but only through luck (the firing was humane, so credit where credit due). Your advice is spot on: derisk financially, do your best work, but don't care when you get walked out. It's just a job, it doesn't matter. If it matters because you need the job, treat your financial situation like an emergency until you don't need the job. It will be a cold day in hell before I am ever on call or in a pager rotation again. "We're not saving lives, we're just building websites." as a wise old man once told me early in my career.
- ldayley 2y agoAuthor: Thank you for writing this! I love reading about how people overcome challenges like this, especially under pressure (and usually overnight!). I am better for hearing not just the technical post mortem but also the human perspective that is usually sanitized from stories like this. This is the kind of technical narrative only a small/solo dev or entrepreneur can share freely.
- deleted 2y ago[deleted]
- adamc 2y agoGreat reminder of the people behind services, as well as a nice account of the debugging process. The reality is that pressure doesn't make you debug problems any faster... usually, it interferes with your thinking. You have to try to ignore the consequences and stay as calm as possible. But most of us have been in some situation similar, if not quite as bad. (Running your own company is going to be uniquely stressful.)
- uaas 2y agoHaving a(n accurate) service graph of all your (internal and external) dependencies is a game changer in troubleshooting issues like this.
- thruway516 2y agoDo you need another dependency for that? A dependency to manage dependency hell.
- uaas 2y agoOptimally, that’s already part of your observability stack, and might well be a built-in feature. In a certain scale/landscape it easily pays off.
- jonnycat 2y agoGreat post, but kind of buries the lede: PostHog is having a CrowdStrike moment.
- timgl 2y agoPostHog cofounder here. This affected users that did not have a specific version of the JS library pinned and deployed a new version, or were using the snippet, and had network capture enabled, (a feature we introduced very recently and is only enabled on 3% of projects), and had recordings enabled on that particular session (for most customers, only a small percentage of sessions are recorded due to sampling or billing limits) This outage was definitely disruptive and we shouldn't have let this happen. We will be doing a full post mortem write up, but this affected a small percentage of our users, so the comparison with Crowdstrike isn't fair.
- Etheryte 2y agoYou're trying to phrase this as if those conditions make it any less bad, but they don't. This affected users that were using the latest version and used... features? Give me a break. Every product has bugs, but trying to downplay the issue after you've just read a distressed user of yours struggle with it is definitely not what you should be doing.
- mrweasel 2y agoThere's certainly a failure to test properly from PostHog, as in they have production features that aren't being tested before a release. On the other hand the author of the article did the exact same thing. They either pushed a release without testing, or they automatically just pull in the latest version of an external library, without any testing or verification. Now I lean towards this being the latter, as if they pushed a release and then the site broke, they would have considered a rollback. Kinda hard to blame others for failing to do testing that you also didn't do. Edit: So others have pointed out that PostHog will just pull down the latest version on it's own, unless you actively disable that feature. That seems like a brave move.
- 2y ago
- misja111 2y ago"Maybe PostHog, I have the api_key blanked out locally to reduce costs" Come on, if POST requests work locally and not on PROD, isn't this an obvious place to start?
- mrsilencedogood 2y agoThe author mentioned significantly more notable prod/dev differences than the posthog API key, which I suspect is where they looked first and second. So no, not an obvious place to start.
- charles_f 2y agoSuch an entertaining read, conveyed very well the sense of stress and abandonment felt by the author. It adds to it that this was written fresh and right in the moment, and feels as an expiation. I'm struggling to find the lesson to take out of that. Limit your dependencies? Have a safe mode that deactivates everything optional?
- foodevl 2y ago> This is no good. Let me just try reverting to a version from a month ago. Nothing. Three months ago? Nothing. Still failing. A year ago? Zilch. Reverting your own code, but still using a broken PostHog update from that same day? For me, the lesson is to make sure that I can revert everything, including dependencies.
- roywiggins 2y agoIt seems that PostHog just always loads the latest version of this piece of itself: https://github.com/PostHog/posthog/issues/24471#issuecomment-2298298849 https://github.com/PostHog/posthog/issues/24471#issuecomment... Though you can opt to bundle it yourself: https://github.com/PostHog/posthog/issues/24471#issuecomment-2298786491 https://github.com/PostHog/posthog/issues/24471#issuecomment...
- slashdave 2y ago> I definitely want to figure out in detail what happened here so I can add a test to prevent a similar change in future! Whoa! Good idea! Could have been worse. At least the change didn't expose a hidden exploit.
- deleted 2y ago[deleted]
- ricardobeat 2y agoOuch. That just adds insult to injury.
- phkahler 2y ago>> It seems that PostHog just always loads the latest version of this piece of itself: Now there's a supply chain attack vector...
- philsnow 2y agoYears ago, IT at the company I was working at force-pushed a browser extension that did this same trick, but the extension vendor in question didn't even bother loading over https. Edit: the extension's manifest gave it nearly every permission, on every web site, including internal ones
- bluepnume 2y agoLooks like the bug was in a monkey-patched `window.fetch` https://github.com/PostHog/posthog-js/blob/759829c67fcb8720f315b620365004d5fec9e2a8/src/entrypoints/recorder.ts#L481 https://github.com/PostHog/posthog-js/blob/759829c67fcb8720f... The biggest lesson here is, if you're writing a popular library that monkey-patches global functions, it needs to be really well tested. There's a difference between "I'll throw posthog calls in a try/catch just in case" and "With posthog I literally can't make fetch() calls with POST"
- pbasista 2y agoThank you for pointing this out. I did not read the post in detail but I was wondering how could a monitoring library cause the entire application to go down. At worst, I thought, it should have failed to process the monitoring events, assuming that it was integrated in a reasonable way. PostHog patching a very important global function is a feature that should be well-documented so that the people who are using it are aware of it and can be reasonably expected to have it in mind when debugging these seemingly unexplainable issues.
- ataru 2y agoI think it worked as defined, it hogged the POST requests?
- ricardobeat 2y agoI was poking around to understand how this was not caught in a test - any ordinary fetch call could have triggered the error, and besides how poor coverage it has for all the ways `fetch` can be used, it seems excessive mocking may have played a part: https://github.com/PostHog/posthog-js/blob/main/src/__tests__/request.test.ts https://github.com/PostHog/posthog-js/blob/main/src/__tests_... The whole fetch and XHR functions are mocked and become no-ops, so obviously this won't catch any issues when interacting with the underlying (native or otherwise) libraries. They have Cypress set up so I don't see why you'd want to mock the browser APIs.
- sethammons 2y agoI have seen so many mocked tests where you end up asserting the logic in the mock works; effectively testing 1=1. The number of issues that can be prevented with an acceptance level test that has a user log in and do one simple interaction is amazing. Where I can convince the powers that be, PRs to main are gated by a build that runs, among others, that simple kind of AC test. If it was merged to main, you _know_ it will not totally break production. We had regular outages with our internal emailing system at a small e-commerce shop. I stepped in and added one test that actually sent an email to a known sink that we could verify and had that test run pre-deploy. We went to zero email outages. Tests had the occasional flake that auto-retried. Also, if your acceptance tests are flaky, how do you know your software isn't? Bad excuse to avoid acceptance level testing
- refulgentis 2y agoThe author needs to relax and Posthog needs more discipline. (and a rename)
- f1shy 2y agoI would love to read from the author what are the lessonS learned. Use better tools? Know better your tools? Know better how to debug? Add yet another tool to detect the error? In all big companies where I worked, at the end of such an event, it boiled down to answer the 3 questions: - what happened? - why did it happen? - what do we do so it does not ever happen again?
- dangsux 2y ago[dead]
- mobeigi 2y agoThese breakdowns happen to everyone and its really bad when its just you against the world. I've been lucky that my last few major outages have all been team efforts with anywhere from 2-10 people working on the issue. Albeit, this is a perk of working in a large enterprise. With more than one person you can bounce ideas off each other and share the pain so to speak. It's highly desirable.
- simpaticoder 2y agoThis person's stress was caused by a single line of code in PostHog. This is the reversion: https://github.com/PostHog/posthog-js/pull/1371/commits/759829c67fcb8720f315b620365004d5fec9e2a8 https://github.com/PostHog/posthog-js/pull/1371/commits/7598... Highlights two lessons. 1. If you ship it, you own it. Therefore the less you ship, the better. Keep dependencies to a minimum. 2. Keep non-critical things out of the critical path. A failing AC compressor should not prevent your engine from running. Very difficult to achieve in the browser, but worth attempting.
- M4v3R 2y agoThese are valuable lessons for sure, but then someone from the marketing comes and demands you add PostHog or any other tracking script to the site and won't take no for an answer.
- simpaticoder 2y agoThen communicate clearly the trade-off the project leader is making.
- DavidPiper 2y agoWhile I agree with you, this sounds like it could lead to a classic case of "that escalation worked, in the sense that it was heard". Unless someone with decision-making power is going to be woken up at 4AM for a problem, they have very little incentive (intrinsic or extrinsic) to block a project on such nebulous claims as "Reducing dependencies" or "Less things on the critical path", if a business leader has come to them with a request.
- threecheese 2y agoEven worse, it appears that PostHog dynamically updates their part of their code at runtime - not bundling it at build time. Their docs note an Advanced Option where all dependencies are bundled in the build. I mean, I get why, and maybe I am misunderstanding but as a user I would expect lazy loading of executable code to be an optimization rather than the default. And used only if fully bundling was a serious delivery delay.
- __MatrixMan__ 2y agoFrom the PR: > fetch() broken on August 19: TypeError: ... Not broken at this version, broken on August 19. This is why I'm terrified of putting anything on the web. It is a dark scary place where runtime dependency on servers that you don't control is considered normal. One day I'll start my own p2p thing with just a bunch of TUI's and I'll only manage to convince six people to use it each for less than a month and then I'll have to go get a real job again but at least I won't have been at the mercy of PostHog.
- scottlamb 2y ago> Not broken at this version, broken on August 19. This is why I'm terrified of putting anything on the web. It is a dark scary place where runtime dependency on servers that you don't control is considered normal. Yeah, that is terrifying. In a nearby comment [1], a PostHog co-founder wrote this affected sites which "did not have a specific version of the JS library pinned and deployed a new version, or were using the snippet". I gather from that is it possible to pin the version, and this incident highlights the value of doing so. [1] https://news.ycombinator.com/item?id=41301008 https://news.ycombinator.com/item?id=41301008
- __MatrixMan__ 2y agoI'd prefer a something where such references are resolved by cryptographic hash so that there's never any ambiguity re: what you're actually getting. Unison does this I believe.
- linuxrebe1 2y agoBased on the way you were troubleshooting it. You can tell you're a programmer first. You went to your code, you went to your logs. Both reasonable, both potential causes of the problem. Both ignore the primary clue that you had. It worked on localhost. As an SRE/devops/platform engineer or whatever the title of the day is people want to give. I would have zeroed in on the difference between the working system. And the non-working system. Either adding and then removing, or removing and then adding back the differences one at a time. Until something worked. What I see is two things. 1) you have an environment where it does work. 2) the failing environment was working, then started failing. Is my method superior to yours, no. It just is being stated to highlight the difference in the way we look at a problem. Both of a zero in on what we know. I know systems, you know code.
- 9659 2y agomany years ago, i was working as an electronic technician. we had a stack of processor boards (from a Perkin Elmer 7/32) that were removed from service. broken. many different revisions, and only schematics for one revision of each board. i thought it was hopeless. an older wiser tech taught me how. plug a good board on an extender. run a diagnostic that fails in a loop. using a scope, look at every pin on the connector. write down what you see. replace with a bad board. repeat. which signals are different? chase them back. if the schematic does not match, get out a voltmeter and your eyes and draw a schematic that reflects how the board is wired. he called this "good card - bad card". and it worked. not going to make any claims about cost effectiveness, but we fixed every board. and my troubleshooting skills in digital electronics were greatly improved. this was a 'fireman' kind of job. waiting for the system to break, so it didn't matter if 2 techs put a week into 1 circuit board.
- tlarkworthy 2y ago... and thats why you pin dependancies
- sillysaurusx 2y agoFor what it's worth, I'm not sure this is a mental breakdown, and it might give the wrong impression to people who do legitimately suffer breakdowns related to stress about tech. For me it's only happened once. It was an anxiety attack, and I'm very lucky my wife was there to talk me through it and help me understand what was happening. She's had them many times, but it was my first (and thankfully only). It turns out that this sort of thing happens to people, and that there's nothing wrong with it. It doesn't mean you're defective or weak. That's a really important point to internalize. Xanax is worth having on hand, since that was what finally ended it for me and I was able to drift off to sleep. I guess my point is, there's a difference between having intrusive thoughts vs something that debilitates you and that you legitimately can't control, such as an anxiety attack or a panic attack. You won't be getting any work done if those happen, and that's ok.
- hughes 2y agoYeah I was actually disappointed to find an article about a routine dependency debugging story. I've felt on the precipice of a breakdown a few times before and was really hoping this would be a more relevant article.
- mynameisvlad 2y agohttps://www.webmd.com/mental-health/signs-nervous-breakdown https://www.webmd.com/mental-health/signs-nervous-breakdown Not all breakdowns come in the form of panic and anxiety attacks. Those are certainly a way that breakdowns can manifest, but it’s not the only way. Stress manifests in wildly different ways for different people and even stressors. You weren’t in his head, experiencing what he was experiencing, so it’s pretty much impossible to “diagnose” from the outside. It certainly sounds like he was functionally paralyzed for hours, even if he didn’t have a full blown panic attack.
- sillysaurusx 2y agoThat's a good point. I didn't mean to gatekeep, which is what I ended up doing. Thank you. If he was functionally paralyzed for hours, then that absolutely qualifies. I was reading through it and thinking "If you can debug stuff, you're probably not having a mental breakdown" and wanted to highlight that some people do break down (which is why it's called a breakdown) and that it's ok.
- deleted 2y ago[deleted]
- dangoodmanUT 2y agowrite my own everything gang rise up
- dbacar 2y agoSince you were able to think and act, I would not call this mental breakdown. That kind of thing is very, very different.
- jppope 2y agoLove the story. The question for the author of course... what did you learn and how can you keep this from happening again
- robinhouston 2y agoAs others have mentioned, the bug that led to this late-night stress was a one-line change to the PostHog library[0]. I take this as a reminder of the importance of giving precise names to variables. The code res = await originalFetch(url, init) looks harmless enough. But in fact the `url` parameter is not necessarily a URL, as the TypeScript declaration makes clear: url: URL | RequestInfo The problem arises in the case where it is not a URL, but a RequestInfo object, which has been “used up” by the construction of the Request object earlier in the function implementation and cannot be used again here. It would have been more difficult to overlook the problem with this change if the parameter were named something more precise such as `urlOrRequestInfo`. (A much more speculative idea is the thought that it is possible to formalise the idea of a value being “used up” using linear types, derived from linear logic, so conceivably a suitable type system could prevent this class of bug.) [0] https://github.com/PostHog/posthog-js/pull/1351/commits/24979cb0a97f86acd247a3027a61ffb2bf74f055 https://github.com/PostHog/posthog-js/pull/1351/commits/2497...
- AlotOfReading 2y agoThe problem with linear/affine type systems is that they have an incredibly high barrier to entry. Just look at ownership semantics in something like Rust. They're not impenetrable (especially with experience), but they're severe enough to be the number one complaint for learners.
- adamredwoods 2y agoSo Zarar needs to keep the localdev as close to prod as possible, or have a separate pre-prod environment that can run integration tests to catch vital function disruptions.
- sltkr 2y agoYour advice is useful for detecting bugs in the code that you release, before your push it to production, but it would not have helped here, because the bug was in Posthog code that was pushed to production asynchronously. In fact, it would have made debugging this particular issue harder, because the difference in Posthog configuration between dev and prod is what clued the author in on Posthog causing the problem. To avoid this kind of problem, the solution is to avoid “live” dependencies which can change in production without your testing. Instead, pin all dependencies to fixed versions and host them yourself.
- adamredwoods 2y agoAh, thanks, I didn't know it was a live dependency. >> Instead, pin all dependencies to fixed versions and host them yourself. Indeed.
- deleted 2y ago[deleted]
- coolhand2120 2y agoSide loading 3rd party scripts in a critical path is asking for problems. Try https://partytown.builder.io/ https://partytown.builder.io/ runs 3rd party scripts like this in a web worker. I'm not sure it would help in this case. Maybe? Probably couldn't hurt to try.
- lifeisstillgood 2y agoWeirdly I think this is heavily related to social anxiety / shame - as in “everyone will knowingness me and point”. This is buried so deep in our brains it’s almost certain to do with herd behaviour. And it’s (IMO) why anonymity online is usually a bad idea - we need to learn, deep in our bones, that what is said online is the same as standing up in front of the church congregation and reading out our tweets - if you would not in front the vicar, don’t in front of the planet.
- gosub100 2y ago> you would not in front the vicar, don’t in front of the planet. Every tyrannical government in history would love this maxim.
- Apocryphon 2y agoNot really. One can be civil without self-censorsing oneself. Candidness is not the same as rank rudeness.
- deleted 2y ago[deleted]
- codexb 2y agoRelevant [XKCD](https://xkcd.com/2347/ https://xkcd.com/2347/)
- languagehacker 2y agoI know that it's not a real post-mortem, but this is the opposite of what a good post-mortem looks like. It includes: * Blaming the tools (and the author) * Not focusing on facts in the timeline * Not considering improvements But that doesn't make for engaging content, right? > * At $TIME we observed HTTP POST calls failing > * At $TIME customers reported inability to make changes to ticket prices and promo codes > * $PERSON took the following steps to debug... > * Root cause: an update to a vendor library resulted in cascading failures to the site > * 5 whys (which might include lack of defensive programming, the use of a CDN without a fixed version, etc. etc.) > * Next steps: pin the CDN version or pull the dependency into the build, etc. Actually, that still looks like a pretty good story to me without any of the associated mania.
- photonthug 2y agoThe trouble with post-mortems is always the things that you can't say. Are we going to talk about how someone wanted analytics everywhere even though it's expensive and they don't use it, that they chose the vendor and didn't provide time to evaluate it, that they wanted it added now-now-now to production even though it hadn't gone through dev, even though it was unplanned work in a busy sprint with other risky work, and that they wanted it on the critical path despite specific objections from some in engineering? Not saying any of this was the case here, but this kind of thing happens frequently. While engineering is saying "let's use blameless post-mortems to highlight problems in the system" a lot of other people are thinking "now we get to see who could or couldn't fix a problem we caused, focus on that, and downplay our own role in causing trouble". Devs talking about CDNs and HTTP-POST and engineering processes are super helpful for steering the conversation away from broken business processes. Easily 80% of PMs I've been involved in probably come down to "we had rules for moving tested code through dev/qa/prod, but we were forced to ignore the process we had agreed to because $PERSON / $DEPARTMENT said so". No one can actually talk about the elephant in the room though, because it's career suicide after you're branded as not a team-player. For all the kids out there.. it's really important to learn to be able to recognize the difference between a PM that's going to be an earnest effort to understand things, vs a PM that is going to be an unhelpful but necessary ritual of humiliation.
- begueradj 2y agoThat's a display of endeavour and persistence. Congratulations.
- flumpcakes 2y agoI've been on my end of plenty of operational outages. I don't want to be harsh but this could have been written by one of my colleagues, the type of colleagues that I really wish I didn't work with. Console logging for hours? Randomly disabling things? Sometimes when you feel "imposter syndrome" you shouldn't ignore it and maybe up your game a bit. In fact, I have dealt with an extremely similar situation where a bunch of calls for one of our APIs were failing silently but only after they had taken card payment transactions. Dealing with the developers of this system was like pulling teeth, after we got them to stop stammering and stop chipping in with their ideas (after half a day with this issue ongoing) it took 10 minutes to find the culprit by simply going through the system task by task until we got to the failing task (confirmation emails were unable to send so the API server failed for the entire order despite payments being taken etc.). This only required 2 things: knowledge of the system, and systematic process to fault finding. You would think that developers who have at least the first, being the ones who wrote it, but sometimes even that is a big ask. Maybe I'm just burnt out from this industry and incompetent people but... come on... no excuses really.
- sleazebreeze 2y agoAgree with everything you said. And then I’d add: Start with reading the error message. In his panic state, he seems to have thought it was a red herring. Error messages are gold. It gives you a concrete thing to work backwards from.
- deleted 2y ago[deleted]
- vvoruganti 2y agoA mantra I heard recently that has been helping me with my own 3AM panics was "None of this matters and we're all gonna die". A bit Nihilist maybe, but has been helpful and just kind of removing the weight of the situation
- dennis_jeeves2 2y agoGood thinking. Sometimes I have to remind my nerdy colleagues about what a true emergency is. A true emergency is when your granny is having a heart attack and has to be rushed to the hospital. When a big corporate website, that does millions of dollars of transactions is down at night it is NOT an emergency. They have to hire staff specifically for night shifts if they are so concerned about the money that they may lose.
- deleted 2y ago[deleted]
- gcommer 2y agoI know solo projects always have an infinite list of "nice to haves". But personally I never skimp on vendoring dependencies. In my experience, not vendoring has _always_ led to breakages that are hard to debug and fix. Meanwhile, vendoring is quite easy nowadays. Every reasonable package manager, and even npm, can do this near-trivially.
- sethammons 2y agothe argument is always "the pr that pulls in the dependency is gross to review with dependency updates" -- and there are ways to mitigate that. I vendor dependencies. My customers want stability and that means a bit more process in managing dependencies. Easy win.
- paxcoder 2y ago[dead]
- kayo_20211030 2y agoI feel your pain. Someone else shot me in the foot. No fun whatsoever.
- ThinkBeat 2y agoWhenever I am dealing with a 3rd party services, I like to write a small adapter for it that bridges the connection point and can keep an eye on a few things. Primary I use a code generator to write most of it. For huge services it may not be practical, but for most it usually provides a heads up if something stops working. with an integration.
- Octabrain 2y agoI can relate to this. I used to be on call for many years and honestly, it destroyed my mental health. In the last company I did it, it felt like falling in a meat grinder for a week. I remember once spending a whole weekend giving support on an bug that was introduced by a recent release. 72 hours of working non stop. Because of that among others, I got a severe burnt out that took me to the deepest dark place I've ever been. To this day, I simply refuse to do on call. There's no enough money you can pay me that would make me to suffer that again. PS: Fuck you, Rackspace.
- tristor 2y agoDuring the mid-early part of my Ops/SRE career I had a senior who was a mentor to me. I noticed as we dealt with outage after outage he was always calm, cool, and collected, so one day I asked him "<name redacted>, how do you always stay so calm when everything is down?" That's when I found that before he'd been in tech, he'd be a Force Recon Marine and had been deployed. His answer was "Nothing we do here is life or death, and nobody is shooting at me. It'll be alright, whatever it is." While I have never experienced anything similar myself, it really helped me to put things in perspective. Since then, I've worked on some critical systems that were actually life or death, but I no longer do. For the /vast/ majority of technology systems, nobody will die if you let the outage last just a few hours longer. The worst case scenario is a financial cost to the company who employs you, which might be your own company. Smart companies de-risk this by getting various forms of business insurance, and that should include you if it's your own company. So, do everything you can to fix the outage, but approach it with some perspective. It's not life or death, nobody is shooting at you.
- wmichelin 2y agoPin your dependency versions! The rollback should've fixed things :(
- wonderwonder 2y agoNothing good comes from working at 2am. Company I worked at pushed a major waterfall upgrade one day. Completely tanked the database while attempting to update the schema. I spent hours on a call with the clients sr. engineer and we eventually came up with a script to fix it. It was after midnight, my director said, good job, you are tired, I'll run the script, call it a night. An hour later director ran the wrong script... and then called me. Clients sr. engineer was legitimately flabbergasted, only time I have ever seen that word apply in real life. Was a not good, very bad day.
- recursive 2y agoB&H Photo Video shuts down intentionally for a whole day every single week, and as far as I can tell, they're one of the top retailers of pro/prosumer A/V electronics.
- seanthemon 2y agoWe once planned an international trip with my dad to new york around visiting b&h to buy electronics about a decade ago. Best memory was walking in there and seeing all the conveyer belts above me shooting packages across.
- merek 2y agoDuring these situations, a prompt message can be very reassuring for users, for example: "We're investigating an issue affecting $X". As a user, I can rule out that the issue is at my end. I can focus on other things and I won't add to the stack of emails. This is one of my biggest frustrations with AWS being slow to update their status page during interruptions. I can spend much of the intervening time frantically debugging my software, only to discover the issue is at their end.
- darepublic 2y agoI remember as a junior dev getting a call at around 3am from Indian tech support about a failed deploy I had been heading. I stressed myself about it so much and only later reflecting back do I realize nobody but me cared. Also funny that the culprit was posthog since I have some past experience with it.
- KennyBlanken 2y agoThe thing that struck me the most was how wildly unprofessional Paul D'Ambra's comments are in response to the bug report. And then he rolled out a fix that was broken, too - showing incompetence in development, understanding the problem, and a total failure to do proper QA on the fix. Royally fucked the pooch twice and he's all "gee golly whillikers!"
- gsora 2y agoThis post resembles how I’ve been feeling lately at work, too bad it’s been months like this now!
- hoseja 2y agoI'm so glad I don't have to work with these byzantine JS monstrosities.
- nurettin 2y agoMore like "The anatomy of a calm and collected response while facing a dire situation thanks to years of expertise, eminem and my sweet wife."