5 ms·
When that incident happened, Gitlab published an actual chat transcript of the incident response. It was a very interesting read. It had the real names of the e
by kunwon1 3y ago
When that incident happened, Gitlab published an actual chat transcript of the incident response. It was a very interesting read. It had the real names of the engineers involved for the first 24 hours or so, then they anonymized it. At some later date they seem to have removed it entirely, which is a shame, as it was an educational read
All that I can find left online is this [1], which is still informative, but not nearly as interesting as I remember the chat transcript being
[1] https://about.gitlab.com/blog/2017/02/10/postmortem-of-database-outage-of-january-31/ https://about.gitlab.com/blog/2017/02/10/postmortem-of-datab...
- ironSkillet 3y agoI thought the parent post was a joke, and then I read your link...
- jonatron 3y agohttps://web.archive.org/web/20170527000404/https://docs.google.com/document/d/1GCK53YDcBWQveod9kfzW-VCxIABGiryG7_z_6jHdVik/pub https://web.archive.org/web/20170527000404/https://docs.goog...
- rbancroft 3y agoThat's pretty cool. I notice how MS has become so much more respected as they have embraced open source. Is the next frontier for companies to become open process where the best companies attract the best people because of their open processes?
- sim7c00 3y agoi think ultimately this will not be entirely possible. if a company posts transcripts with real names of people that can be damaging for the people (they moght misbehave simply due to company culture, pressure not visible through the open process etc.) which then would give a company a cheap way out (emplyee scapegoat). since this is a possibility, i think employees amd unions etc. would resist. despite if it was an honest afair it would be awesome, and that (possible but improbable) future looks super cool :)
- hoofhearted 3y agoI’ve actually been exploring this concept myself with a blogging side project I’m working on. Things were “mehhh” for developer interest for the first month or so, as a I was working off of an internal todo list. I didn’t think anyone cared about my 300 line long todo.txt file, but then I started to wonder if I should find a way to put that doc out in the open for developers to follow, and possibly jump in and contribute. I had a hinkling that I could use GitHub issues to help with this, but I believed the title “issues” would hurt my project. I was under the impression that a new open source project with a single contributor, and a ton of open “issues” would look bad to developers. I started to inquire with devs on IH and hear about using issues for feature tracking. Much to my surprise, I got an overwhelming “yes, you need to use issues”. I was also told not to worry about the misleading “issues” title, and that enough developers were knowledgeable enough to know they weren’t just bug reports. As I started to open issues, and ask for help; surprisingly I started getting traffic and interest. The more issues I opened, and the more open I was online about my code and plans for it; the more followers and contributors I’ve gotten. My plan at this point is to just follow Gitlabs model, and go full open transparency with everything. My side project mentioned above can be downloaded here: https://github.com/elegantframework/elegant-cli https://github.com/elegantframework/elegant-cli
- 666satanhimself 3y ago[dead]
- grumple 3y agoWhat's interesting about that is that they have a relatively small db, and that they only had a single snapshot available (and WAL apparently wasn't available or backed up for a point in time restore). We use RDS at AWS, thankfully handles automated daily snapshots and point in time restore via binlog. Cost is relatively cheap too (for a company making millions), and our db is several times larger than Gitlab's was at the time. It's also pretty worrying that a single developer was doing work directly on production databases without a second person there to say "yeah, looks good". This is a big operational mistake, no matter how good you think you are.
- firecraker 3y ago>2017/01/31 23:00-ish YP thinks that perhaps pg_basebackup is being super pedantic about there being an empty data directory, decides to remove the directory. After a second or two he notices he ran it on db1.cluster.gitlab.com, instead of db2.cluster.gitlab.com >2017/01/31 23:27 YP - terminates the removal, but it’s too late. Of around 310 GB only about 4.5 GB is left I can't even imagine the sinking feeling..
- nofinator 3y ago> YP says it’s best for him not to run anything with sudo any more today, handing off the restoring to JN. Then in the post-mortem about lack of backups: > LVM snapshots are by default only taken once every 24 hours. YP happened to run one manually about 6 hours prior to the outage > Regular backups seem to also only be taken once per 24 hours, though YP has not yet been able to figure out where they are stored. According to JN these don’t appear to be working, producing files only a few bytes in size. I have had (and inevitability will have again) bad days like poor YP. All I can count on is to maintain good habits, like making backups before undergoing production work like YP did.
- capableweb 3y ago> like making backups before undergoing production work The specific part you mention also brings up a really vital part of a backup system, testing that the backups generated actually can restored. I've seen so many companies with untested recovery procedures where most of the time they just state something like "Of course the built-in backup mechanism work, if it didn't, it wouldn't be much of a backup, would it? Haha" while never actually tried to recover from it. Although, to be fair, I've only seen one time out of the untested 10s where it had an actual impact and the backups actually didn't work, but the morale hit that the company ended up having made my brain really remember the fact to test your backups.
- freedomben 3y agoIndeed, the feeling of dread when you do something that causes prod to go down is bad enough. I can't even imagine the feeling when accidentally deleting prod data...
- mdaniel 3y agois this the thing you're referring to? https://web.archive.org/web/20170527000404/https://docs.google.com/document/d/1GCK53YDcBWQveod9kfzW-VCxIABGiryG7_z_6jHdVik/pub https://web.archive.org/web/20170527000404/https://docs.goog...
- Shank 3y agoThey actually broadcasted a video call of them fixing it too, but they removed that as well.
- SushiHippie 3y agoIs this original the chat transcript? https://news.ycombinator.com/item?id=36635060 https://news.ycombinator.com/item?id=36635060
- kunwon1 3y agoI believe this contains excerpts from the original transcript, at least. I'm not sure if this is the same document I read, but if not then it's close!