14 ms·
We see this pattern at PagerDuty over the majority of our customers. There is a definite lull in alert volume over the weekends that picks up first thing Monda
by kenrose 10y ago
We see this pattern at PagerDuty over the majority of our customers. There is a definite lull in alert volume over the weekends that picks up first thing Monday morning.
It's led to my personal conclusion that most production issues are caused by people, not errant hardware or systems.
- asimuvPR 10y agoI've come to question releasing often as a result.
- jon-wood 10y agoI can see the argument that if releasing causes things to break then don't release so frequently, but in practice the end result of that is lots of things breaking at once and having to unpick everything. Debugging is much easier if you're debugging a single change fresh in your mind.
- bryanlarsen 10y agoYeah, debugging 2 bugs that are shadowing each other often takes several orders of magnitude longer to debug than just a single issue.
- Thrillington 10y agoWhile having tooling for tight debugging loops allows short release cycles, it leaves that decision up to the business
- asimuvPR 10y agoYou are right. Making small changes that can be isolated for debugging purposes is a good approach. What I mean is that we should always question "best" practices and how we apply them to our development process. These days there is a tendency to drink the kool-aid (guilty of this as well). DevOps is something that we are still learning and developing as a profession.
- dozzie 10y ago> DevOps is something that we are still learning and developing as a profession. Not really. DevOps is simply a modernish label for system administration that we have for dozens years already.
- the_other 10y agoIsn't that just "Ops"?
- dozzie 10y agoMy point exactly.
- aurelianito 10y agoYes, but now we actually try to automate in a systematic way; as opposed to a single admin hacking some custom shell scripts. That's what the dev in devop mean. And that's new, at least it was not a common activity in the XX century.
- dozzie 10y agoIf by "systematic way" you mean a group of programmers clueless about system administration hacking together some custom Ansible scripts, then I wouldn't call it progress. Sysadmins have tools to automate their work for a long, long time (cfengine, bcfg2, even Puppet and Chef predates DevOps hype). DevOps didn't bring anything new to the table.
- aurelianito 10y agoIs not about doing it right or wrong. Is about the automation that is reused cross-project and the fact that a lot of things that used to require a sysadmin now we automated its job away. If developers implement it correctly or poorly is a different issue. It is also not about hype (or not). I do agree that the name is posterior to the beginning of the practice, but it is the name that we have.
- zaroth 10y agoIt seems like so many "best practices" are really thinly veiled attempts at exploding complexity with only tenuous potential business advantages. We create ourselves so many of the problems we are paid to solve.
- mLuby 10y ago"We create ourselves so many of the problems we are paid to solve" Sounds lucrative ;)
- jessaustin 10y agoIt works for doctors and lawyers; why shouldn't software developers try it?
- thomaslee 10y ago> It seems like so many "best practices" are really thinly veiled attempts at exploding complexity I'll bite. :D It's a bit more nuanced than that IMO, the "deploy often" mantra is only as good as the process around release + deploy. If you half-ass testing and push to production without without a process for verification -- or if, say, your deploy process is half-baked, or your staging environment is worthless -- you can probably expect "interesting" production deploys on a pretty regular basis. As much as we'd like to pretend we're all good engineers, this happens more often than you'd expect -- even with good engineers: at some point a company transitions from scrappy startup to a shambling beast, and the things that used to work for a scrappy startup (like skimping on testing and dealing with failures in production) are insufficient when you've got more eyes on the product. Further, the engineering culture remains stuck in "scrappy startup" mode long after the shambling has begun. And all that's ignoring the fact that less frequent deploys with more changes have their own set of problems. We actually got to a point with deploys of a certain distributed system such that we were terrified if we had more than a few days worth of changes queued up. So many things that could go wrong! :) > We create ourselves so many of the problems we are paid to solve. This, on the other hand, I completely 100% agree with: if not us, then who? :) EDIT: minor formatting change
- cortesoft 10y agoThe breakage rate per new feature is fairly constant; if you release 7 new features once a week or 1 new feature once a day, you will have the same number of issues. The question then is; is it easier to deal with all the issues at once, or a smaller number of issues every day?
- danek 10y agoIn addition to that, I would also worry about interactions between issues. I tend to lean towards spreading out the issues over time to make the eventual diagnosis easier. Occasionally we'll have a problem where we cannot deploy to production for several days (normally it's once a day). A massive inventory of ready-to-deploy features builds up. When we do finally deploy, this deluge of features and fixes creates new issues that force a rollback, delaying the deploy even longer...
- msoad 10y agoYou can roll back a small release that broke the world. A
- AYBABTME 10y agoOr the opposite: embrace failure, practice resiliency, work with it instead of against it!
- awj 10y agoFunny, I've taken it as justification for releasing often. If I have a hard time changing one thing without breaking the service, it's nearly impossible to change a hundred things without breaking something. Since it's a given that I will have to change things, I'll try to stick to a scope where I stand a chance of doing so successfully.
- selckin 10y agoDo they have the same load/users during the weekend?
- jweir 10y agoNo, in the video he states that the weekend is their busiest period.
- DuskStar 10y agoAbsolute numbers of alerts are probably a lot less useful than alerts/use, or alerts/users. After all, if people use the services PagerDuty covers less often on the weekend then you might see a lower alert volume even if issues are relatively more common.
- ex_amazon_sde 10y ago> most production issues are caused by people It's a well known fact, both for systems and networks.
- Yhippa 10y agoAfter working at various enterprises over the years (where deployments are slower in some cases) I've noticed you'd do a Thursday/Friday/weekend deployment, everything "looks good" and you'll still have a bunch of issues Monday morning due to users finally using the system en masse.
- jsprogrammer 10y agoThe user experience is pleasant as well. Nothing better than your dependency breaking at 4PM on Friday with no support until Monday.
- dnautics 10y agoPretty sure the user volume for Uber is higher over the weekends than during the week. Particularly system stressful times are 1:30-2:15am PT on Saturday and Sunday mornings.
- kasparsklavins 10y agoThere is less commuting on weeekends.
- rsanek 10y agoIn the talk, he specifically mentions that weekends are the busiest.
- chejazi 10y agoWhile a lot of people go out on Friday and Saturday night, I don't think the total volume surpasses commute volume. The surge in pricing you experience late at night is largely due to the number of drivers on the road, which would be less given the time of day. Edited out consideration: traffic in your area might be a good proxy for Uber volume. Traffic late at night is generally low.
- dnautics 10y agoYour statement is true in San Francisco, but Pacific time zone also encompasses Oakland, south bay, los Angeles, and san Diego, where relatively few people commute using Uber compared to going out using Uber. From personal experience, having driven for Lyft and Uber I'm two cities (Sf, sd) surge in SF does get acutely high in the morning, more broadly in the evening (work departure times are spread out over a wide range of hours), and acute around 2am. In California, bar closures statewide are 2am at the latest, so that accounts for unified departure times in multiple markets.
- anezvigin 10y agoWould PagerDuty consider publishing some anonymized aggregate statistics?
- employee8000 10y ago90% of outages are caused by configuration changes which is why change management was so hot for enterprise software.
- ChuckMcM 10y agoThat conclusion is well founded. We correlated issues at Blekko across a lot of different factors, the one that always held was code or configuration changes. Not too surprising in the large but definitely confirmed by the data.
- hinkley 10y agoSo I should fire all my developers is what you're saying... Man, I'm gonna save so much money.
- ryandrake 10y agoYou laugh, but a lot of time the best thing for a software product is to not make so many damn changes. People tend to over-estimate the value of features and tweaks and under-estimate the risk of things going wrong.
- hinkley 10y agoOh I know. I've been in a meeting with several bosses where I tell them flat out "don't give me more developers, give me more hardware and tools". They usually look like I have an extra arm growing out of the top of my head.
- ChuckMcM 10y agoThat does save a lot of money :-) The trick is managing the rate of change and the risk of disruption. If you manage it to no risk you end up changing too slowly, if you manage it to close to the risk you end up with unexpected downtime and other customer impacting events. Understanding where you are between no risk and certain doom really only comes with experience.
- hinkley 10y agoI'm bad at optimism, but maybe the takeaway here is that we're finally at the point where it's not the hardware or systems that are the biggest problem. The automation is actually working well enough that it's not the tall tent pole.
- Aeolun 10y agoThat doesn't surprise me at all. If you're not changing software, you can't create any additional bugs.
- pbhjpbhj 10y agoBet I can!;o)
- di4na 10y agoQuestion is : people at each level. Here most of the alerts happens at night... due to batchs... Most of the other alerts are due to network issues, moving servers due to hardware intervention that change the topology of the network, etc etc.
- mathattack 10y agoCouldn't this also just as easily be a change in usage patterns?