6 ms·
Incident report for February 21st, 2024
- yatish27 3y ago"While building a feature, we performed a database migration command locally, but it incorrectly pointed to the production environment instead, which dropped all tables in production." This was scary.
- Rygian 3y agoSeparation between office and production should be the norm.
- hn_throwaway_99 3y agoExtremely so in my opinion. In this day and age, viewing something like this means this company fundamentally doesn't understand security. It should simply be impossible to "incorrectly point to the production environment", because whoever ran this shouldn't even have access to those credentials in the first place.
- ttul 3y agoThis implies that their production environment is mutable. As in, a command can run and change the production environment. That’s a no no. But I give them a pass because they are a young company. My company was similarly reckless early on, but as we scaled, we had to tighten things up and turning to an immutable deployment approach has saved our asses so many times.
- syslog 3y agoHow does immutable work in regard to databases? As in „we need to add a column“?
- inopinatus 3y agoSomeone writes the migration, commits it, it passes the build and unit test stages of the pipeline, then the application as currently running passes all function and integration tests with (and this is important) both the prior and the revised schema. Your commit is tagged as release ready! Not long after, the automation tooling confidently executes the now-tested migration under machine control during the next deploy, everyone goes home happy with your shiny new published_at column, and no-one has directly touched prod. Two days later the CTO sends everyone a stroppy email about "column bloat that should've been a table", ssh's into the personal instance that they've been keeping alive† since before you had funding and learned to launch servers as immutable black boxes, and whilst trying to prove a point by rolling it back manually, drops all tables by mistake when a cat treads on the keyboard -- † excuse: "it's for reporting"
- RowanH 3y agoRookie move having the cat on the desk while ssh'd into prod...
- seanwilson 3y ago> Someone writes the migration, commits it, it passes the build and unit test stages of the pipeline, then the application as currently running passes all function and integration tests with (and this is important) both the prior and the revised schema. Your commit is tagged as release ready! Not long after, the automation tooling confidently executes the now-tested migration under machine control during the next deploy, everyone goes home happy What happens if something goes really wrong after the production deploy? Is there a way to skip steps if you need to quickly push an emergency fix?
- kgeist 3y agoAt our company, we have "an immutable DB", too, but when there's a critical emergency (say, full downtime), we can apply fixes manually. In that case, we run the tests after applying the fix.
- ZephyrBlu 3y agoIs being a young company really an excuse when any half-decent engineer knows these things are bad? Being a young company doesn't mean you ignore all the mistakes other people have made and figure them out for yourself. I really surprised someone has access to the prod DB, and that it's possible for them to connect to it in dev (Meaning they have a copy of the credentials???).
- AYBABTME 3y agoKnowing it's bad and punting on it for later are both things that can be possible at the same time.
- disillusioned 3y agoThat single sentence contains multitudes: * Production should be immutable * No one doing dev in a dev environment should have such trivial access to prod * Are there still good reasons for a migration to drop all tables? I guess it's for the dev environment to etch-a-sketch to a known state? Yikes.
- anonzzzies 3y ago> No one doing dev in a dev environment should have such trivial access to prod It’s the new and hip ‘cloud’! Probably using planetscale or something like that, which (last I checked, maybe it changed but wasn’t on), doesn’t even have ip protections outside the mysql user settings (while bad, would’ve protected them). > Are there still good reasons for a migration to drop all tables? We haven’t found any.
- willsmith72 3y agoreally? i've definitely done it before on my local as a quicker alternative to cleaning up the docker container/volume, doesn't seem that bad ofc i'd think differently if i was also putting write-permission prod credentials into my machine, but luckily i haven't been in many places doing that
- rwieruch 3y agoPlanetScale has Safe Migrations which you can enabled for your production DB (branch). Wondering though whether this will protect against everything mentioned here. https://planetscale.com/docs/concepts/safe-migrations https://planetscale.com/docs/concepts/safe-migrations
- samlambert 3y agoSafe Migrations would prevent this completely. PlanetScale also allows you to restore multiple backups in parallel.
- viiralvx 3y agoI'd suggest reading up on what some of these new database providers are doing to help prevent or fix mistakes like this. Since you mentioned PlanetScale, I'll use them as an example. 1) PlanetScale has IP ACLs, which locks down passwords to specific IP addresses. [1] Additionally, with TailScale or another VPN solution, locking down based on IP isn't necessary foolproof. 2) They also have Safe Migrations. When enabled, it prevents DDL from being run directly on a database. [2] Additionally, using deploy requests for zero-downtime schema migrations also allows you to use reverts, which will revert the migration. [3] [1] https://planetscale.com/blog/introducing-ip-restrictions https://planetscale.com/blog/introducing-ip-restrictions [2] https://planetscale.com/docs/concepts/safe-migrations https://planetscale.com/docs/concepts/safe-migrations [3] https://planetscale.com/blog/behind-the-scenes-how-schema-reverts-work https://planetscale.com/blog/behind-the-scenes-how-schema-re...
- sergiotapia 3y agoI did this once early in my career, the ice cold feeling in the pit of my stomach permanently etched this lesson into my soul lol.
- JackFr 3y agoI did too, but it was 25 years ago and quite honestly there was a different attitude about developers accessing production. Today if a developer can bring down the operation accidentally, that’s a problem with the org more than the developer. (On the other hand if a developer screws up the shared dev environment, it is his or her fault and they deserve the wrath of their coworkers.)
- JackFr 3y agoAt a company with a modern, mature process this simply should not be possible.
- anonzzzies 3y agoThis sounds like one of those horrible tools like prisma which drop everything if something is not in sync on dev. We removed this type of stupid in favour of our own which, you know, fixes this actually instead of lazily dropping everything when they cannot resolve some trivial thing, for instance, a new required field without default when there are already rows and other crap which they call ‘opinionated’. No idea why we ever used that stuff as it caused so much grief (never on prod though); after prisma/drizzle and some other ‘modern’ horrors, we mistrust everything that’s ’hip and new’ so ‘everyone uses it’. One of those hip and new things I mistrusted was Resend even though ‘suddenly everyone uses it’. I’ll wait a few years before even contemplating it.
- satvikpendem 3y agoI like that Prisma drops the database. Developer databases should be idempotent and easily reseedable with sample data. It's the same concept as cattle, not pets, from the devops world but now applied to databases. There should be nothing special about a database in development that cannot be dropped and recreated.
- anonzzzies 3y agoI don’t care about dev databases; I care about why it does it; there is no need for it, at all. > I like that Prisma drops the database. ‘Like’ ; so if it didn’t but still worked perfectly, you would complain that it didn’t?
- satvikpendem 3y ago> ‘Like’ ; so if it didn’t but still worked perfectly, you would complain that it didn’t? Yep I would, because it maintains the idempotency invariant, plus it flushes out any bugs that may exist if databases are not idempotent from seed data, that is the "why it does it."
- kgeist 3y agoWe have a script which recreates the dev DB, but developers must run the command manually so that there was an understanding what's happening.
- walrus01 3y agoFrom the company's homepage: "deliver marketing emails at scale" Maybe this company doesn't need to exist, and shouldn't.
- vmfunction 3y agoIt is actually very open of them to post that! I'm sure things like that happens in larger companies, but I'm sure we won't get a incident report like that from them. I would rather work with that kind of company. Then we can also see if they learned from their lesson, and this doesn't repeat again.
- heyarviind2 3y agoMany companies are already solving the same problem, and then resend comes into the picture with the support of YC. They are doing everything except sending the mails.
- walrus01 3y agoThey do, or do not operate MX and deliver SMTP traffic to other mail servers?
- albertgoeswoof 3y agoThey are a wrapper over AWS SES
- kgeist 3y agoReminds of this classic thread: https://www.reddit.com/r/cscareerquestions/s/cqrama0L1z https://www.reddit.com/r/cscareerquestions/s/cqrama0L1z >accidentally destroyed production database on first day of job Same symptoms: while developing a feature locally, they accidentally pointed to the production DB, which was destroyed when running tests.
- doctor_eval 3y agoIf they’re small then you can see it happening where someone was logged into prod using some environment variables to sort out some issue - probably didn’t even update production, just a few queries - and then went back to work. Hours later they run some script that does DROP DATABASE from that same shell they used to troubleshoot, which takes a little longer than usual… Anyway I can totally see it happening to me in my little one man shop, but now i think I might look at removing some privileges from my prod account :)
- pistoriusp 3y agoOr just use https://www.snaplet.dev https://www.snaplet.dev
- doctor_eval 3y agoUnclear how that would have helped.
- pistoriusp 3y agoMy assumption is that they connected to production to run their local environment against production to debug something. They needed the data. Snaplet can capture a small subset of data related to a particular user, and remove the private information, which they can restore on their local machine. Does that make sense?
- doctor_eval 3y agoSure; but I assumed they logged into production to run some queries to debug an issue unrelated to dev. I mean, I know there are some bad practices out there - but connecting local dev environment to prod database server would be insane for any reason!
- pistoriusp 3y agoYeah, it's rather crazy, but it happens. Sometimes you need to mutate things in order to get the full experience of testing the software locally. So, really you need the data locally and this is what I'm suggesting. Don't mess with the developer experience, just give them tools to self service.
- laborcontract 3y agoMan, I use Resend, I really want to like them and yeah, it is really simple to get started which is great but man, this is, I think they really need to slow down a bit and maybe try to figure out how to put some processes in place to maybe intentionally slow things down. This is the second incident that can be characterized by a very hazy delineation between development and production environments. The first incident had to do with an attacker gaining access to private credentials due to devs leaving keys set in NEXT_PUBLIC environmental variables on their site.
- dbbk 3y agoI don't say this to be rude, but honestly it's because they are an amateur company. When you're dealing with critical infrastructure like email, you should stick to the battle-tested services like SES, Postmark, Sendgrid etc.
- hw 3y agoProduction database access must always be locked down from external traffic, and only allow traffic from the production application or within the production environment. Aside from mitigating local dev accidentally pointing to the prod db, if you have the db accessible externally means it’s susceptible to network attacks and password attacks
- globular-toast 3y agoSometimes I think the infra at the small company I work at isn't great, but we've not had direct dev access to prod DBs from day one. They're locked down at the IP level. You'd have to go through some serious hoops to accidentally connect to prod even if you had the keys. I remember as a child thinking adults had everything under control. That they know what they're doing. I guess I assumed a day would come when I too would know. That day never came. It's easy to think when you look at shiny websites and the first paragraph of this comment that other adults do know. But I'm often reminded of the truth: nobody knows. Everyone is always operating at least slightly outside their comfort zone or, in the case of the article, wildly.
- nbittich 3y agoincident report / post mortems could be the best way to promote a dev tool. Didn't know what resend was before someone deleted the prod database. Now I wonder if I need it / if I want it.
- wrftaylor 3y agoI love what Resend are doing and am a customer. I can also absolutely empathise as our lead engineer did exactly the same thing at a startup I was running a decade ago. It's a horrible situation. But yeah, both the incident and the report are really tough to read. It would be great if they can do a follow-up with further actions they're taking. There's a neo-bank called Revolut that allegedly at one point had just two teams: "go fast" and "don't screw it up". I feel like an infrastructure play needs some dedicated hires in camp 2.
- pistoriusp 3y agoUnfortunately these sort of mistakes are seen as a "right of passage" for many developers. I ran "`DELETE FROM users;` without a WHERE clause against production in my first year on the job. I felt absolutely terrible. I thought I was connected to a development machine. Fortunately we had backups available. Often this isn't a problem with the individual developer itself, but points to a problem with the organization. Frankly most developers shouldn't have access to a production database, let alone mutable access. One major concern is loss of data, but another is privacy. It's so frustrating to see this happening when there are tools that solve this like Snaplet (I'm a founder), and Replibyte that allow you to generate or obfuscate data for usage in dev-environments, and Neon that allows you to branch your database.
- tianzhou 3y agoWe are also building Bytebase, which enforces the change review process for such operations
- pistoriusp 3y agoIt looks like we’re getting downvoted, lol. Not sure why that’s a thing. Bytebase is awesome. We should write about each other in our docs. You do the migration and we’ll do the data.
- kgeist 3y agoI think DB engines should simply have some sort of a default option in the interactive mode where if you write "delete from table", it asks "Are you sure? You're going to wipe out the entire table! Y/N". Would've probably solved 99% problems.
- namaria 3y agoI think you mean 'rite of passage'
- pistoriusp 3y agoIndeed. Thanks for the heads up.
- rnts08 3y agoThis will keep happening as long as people are unable to learn from the past. Yes it's expensive to have a good experienced infrastructure engineer on the team, but at least you know there's someone testing your backups and procedures for when your eager dev team screw up.
- tedchs 3y agoI did this, around 2005, but I dropped ALL the prod tables. I was using a SQL GUI called Toad (awesome) and had separate windows open, for both "dev" and "prod". I was trying to reset the dev database, and used the wrong window. Thankfully, the "real" DBA at the time had 15-minute backups, and was able to restore it, and then I replayed a few transactions from logs. Lesson learned! > While building a feature, we performed a database migration command locally, but it incorrectly pointed to the production environment instead, which dropped all tables in production.
- JackFr 3y agoI deleted a years worth of trades on a production database for a bond trading system in 1997. They were restored within an hour, but it made a lasting impression. Different color scheme on prod consoles. Always use “BEGIN TRAN”.
- tetha 3y agoThis is why I've conditioned myself: Once I start running unsafe commands in a development-ish pattern (DDL changes, reboots, service starts/stops and such), I make sure to close any shell to productive systems and databases first. They are color-coded differently, normal users don't have these permissions and so on and so on, sure. But it's better not to have these shells open and available to accidentally choose them once my brain switches from ops-mode to dev-mode.
- forwardemail 3y ago[dead]
- dpcx 3y agoI did something like this nearly 15 years ago at a job. I was trying to use something like MySQL Workbench to export the database and generate a map of connections through foreign keys. Apparently I selected some option backwards and it deleted out the entire production database. Luckily for me I'd been working on some other things related to it and had taken a backup not long prior, but it was pretty nerve wracking to hear the CTO/CEO nearly running through the halls to find out what had happened. Pretty sure they never implemented more stringent access controls at that company, either.
- PeterZaitsev 3y agoYou should always be ready for your database to get trashed - application bugs, operator error, hacker intervention... What strikes me in the incident report they focus on failed migration, where the real issue is not planning or not testing for recovery if migration goes very wrong. Even if backup recovery would take just 6 hours would it be acceptable ?
- _andrei_ 3y agoI hate it that some companies focus more now on how their product's website looks than on the product itself, its quality and stability. I also hate it that users are so used to having issues, errors, and their data leaked, that a startup can cut all the corners and do things completely the wrong way, and still have success. I hate that because I couldn't do it. Last month Resend leaked a database API key [0] from their env, how is that even possible? Now a developer tried to run a migration locally and did it on production, again, WHY is that possible? How can you talk about enterprise plans, SLAs, ACL features, when you can't do the basics right? What's it gonna be next month? [0] https://resend.com/blog/incident-report-for-january-10-2024 https://resend.com/blog/incident-report-for-january-10-2024
- jonaswalker305 3y ago[dead]