26 ms·
I Accidentally Deleted 7TB of Videos Before Going to Production
- KingOfCoders 4y ago"I'm under an NDA" Don't write a blog post.
- JacobiX 4y ago> It involves bad practices and errors from multiple parties in a world that might seem > foreign to the "Silicon Valley" world but paints an accurate picture of what > development is for small IT companies around the world Everybody makes mistakes even in the "Silicon Valley" world, but such problems cloud be easily caught by testing (which he did but it was restricted to the first page) and performing a simple dry-run.
- crispyambulance 4y agoExactly, everyone makes mistakes. Sometimes huge ones. In hindsight or on the sidelines it's always easy to point out a few technical things that WOULD HAVE avoided catastrophe, but does that help? I think not (aside from a cautionary parable for interns). Things are complicated, people are human and forget things, there are pressures to "get it done" and override the guardrails. Everybody has horror stories. Some worse than others. Welcome to the OP's day of horror. I would think "Silicon Valley" dev-ops horror stories make this one seem like a triviality.
- legalcorrection 4y ago[deleted]
- JauntyHatAngle 4y agoI'm baffled by this too. Unnecessary bridge burning I'd call it. It's not even necessary to the story.
- dkersten 4y agoIts explained on the first line: " I'm a Junior Developer with less than one year of actual experience. Some of the things that might seem obvious to some might not be so for me". I guess it applies to this, too, not just the technical aspects.
- dsego 4y agoI might've missed it, but I don't think that line existed when this was first posted.
- thevinter 4y agoYou're right and I edited the company's name (might be too late but better this way). That said I'm not very happy with the experience of working for TheCompanyTM anyways so I'm in the process of switching jobs. Thanks for the comment :)
- legalcorrection 4y agoYou're welcome and good luck!
- KingOfCoders 4y agoTalking bad about your employer is great for finding a new job. Companies are eager to hire people who bad-talk them.
- breakfastduck 4y agoHe doesn't talk bad about his employer. He talks bad about his employers client.
- mkr-hn 4y agoTech is like any other human endeavor. People talk. People change jobs and still like the people in the place they left.
- breakfastduck 4y agoYes exactly. Which is why I wouldn't touch anyone who has no criticism for the systems OR culture of a place they've been before. Nowhere is perfect. If people can't be honest about the flaws then they're useless.
- KingOfCoders 4y agoOf course you can hire whomever you want. I would hire someone who has criticism about what he had done in the past and what they have learned. Nobody is perfect. But people with no self reflection blaming others and their employer? No thank you.
- dsego 4y ago> but at the time the code seemed completely correct to me It always does. > Well, it teaches me to do more diverse tests when doing destructive operations. Or add some logging and do a dry run and check the results, literally simple prints statements: print("-----") print("Downloading videos ids from url: {url}") print(list of ids) ... ... ... # delete() dangerous action commented out until I'm sure it's right print("I'm about to delete video {id}") print("Deleted {count} videos") # maybe even assert ... Then dump out to a file and spot check it five times before running for real.
- V__ 4y agoThis was my first thought too. Another think I like to do, is to limit the loop to say one page or 10 entries and check after each run that it was correctly executed. It makes it a half-automated task, but saves time in the long run.
- mipmap04 4y agoI do this, too, but I also take a count of the expected number of items to be deleted as well. If my collection I'm iterating over doesn't have exactly that number of objects I expect, I don't proceed.
- gilleain 4y agoYes, I find command line tools that have a "--dry-run" flag to be very helpful. If the tool (or script or whatever) is performing some destructive or expensive change, then having the ability to ask "what do you think I want to do?" is great. It's like the difference between "do what I say" and "do what I mean"...
- mmcclimon 4y agoThe rule we have is that anything that is not idempotent and not run as a matter of daily routine must dry-run by default, and not take action unless you pass --really. This has saved my bacon many times!
- 4y ago
- fedeb95 4y agoin my opinion any process that isn't preceded by another identical and automated process that varies only by the data involved is very risky to do in production. your management hopefully had a big reality check? or not because of backups?
- nix23 4y agoZFS -> Snapshot....always!! Before touching writable-data (my personal mantra) ;)
- hnlmorg 4y agoI love ZFS too but that's not really relevant to this discussion because the deleted items were on a video hosting platform and the company did already have local copies.
- nix23 4y agoYes and? Make a snapshot on live. Again, never touch data before snapshot.
- volume 4y agoThis reminds of some IRC threads. You post a question and someone's answer assumes you are going to rip out and replace your existing prod setup just so you can use their pet tool.
- nix23 4y agoPff, there are thousands of system's and filesystems that are capable to make snapshots, even a shadow disk from VM370 (1982) could be seen as one.
- hnlmorg 4y agoAt risk of sounding snarky, you do understand how video hosting platforms work? Customers, even enterprise ones, don’t have shell access let alone control over what file system is used. There are a hundred ways this problem could have been prevented but ZFS isn’t one of them.
- nix23 4y ago>you do understand how video hosting platforms work? No, no i don't.
- whiplash451 4y agoSo, “i am under NDA” but I reveal my client’s name and a lot of sensitive details about what we are doing. LOL.
- dewey 4y agoWhere do you see the clients name? I only see Vimeo being mentioned.
- ceejayoz 4y agoIt has been edited. https://news.ycombinator.com/item?id=31271836 https://news.ycombinator.com/item?id=31271836
- dewey 4y agoGot it. To be honest I'd be hesitant to publish a blog post like that with your name + current company name attached to it. It's a bit different to share a fun story a few years later about that time you almost wiped production.
- daniel-cussen 4y agoWell at least deleting the secret is a step back toward the NDA he left behind.
- Closi 4y agoIt still breaks the NDA: * Firstly, you don't have to name the company to break the NDA anyway (you are still disclosing information you aren't supposed to disclose regardless of if it can be linked back to the company). * Secondly, the client is still named on the front page of the website. * Thirdly, OP posted this with his real name that trivially links back to the dev shop he is working for. The site also has his CV which lists the client again, with a description of the project to link it to the post. * Finally, The client can trivially be identified by googling the description in the second paragraph (i.e. just search the named countries in operation plus the word Gym).
- Helitio 4y agoJust a note: being able to click yourself a server at Google, AWS etc. Might be cheap enough even paying for 15tb of traffic.
- ghoomketu 4y agoThe more I read about vimeo the more I wonder what's up with these guys. Only recently they made some god aweful policy changes for content creators(1), but it looks like they treat their enterprise customers just the same. Surely, there must be better alternatives for hosting videos than being at the mercy of a company who couldn't care less about big paying customers. (1) https://www.theverge.com/2022/3/18/22985820/vimeo-bandwidth-policy-changes-2tb-cap https://www.theverge.com/2022/3/18/22985820/vimeo-bandwidth-...
- pfista 4y agomux.com seems like a great alternative and is super developer focused.
- Peleus 4y agoUnder NDA but I'll give rough details of what's occurring while also naming my client and disparaging them to the public. Well that's a brave move...
- searchableguy 4y agoThey said they are a junior developer with not much experience. I'm afraid they may not know what is and isn't covered under NDA.
- KingOfCoders 4y agoMy tip would be: read what you sign.
- thevinter 4y agoJust to clarify, my company is under an NDA and not personally me. It also encompasses only the actual project details so a post like this is legally compliant. (Not a lawyer, might be wrong)
- KingOfCoders 4y agoSo you're not under an NDA as you wrote. I don't know your position but I would assume a NDA is part of your freelancer or employee contract.
- mkr-hn 4y agoOP might at least want to consult with a contract lawyer in Italy to make sure.
- bluehatbrit 4y agoIn every contract I've ever signed, part of the NDA clause with my employer is that I'm also bound by NDA's my employer is bound by, so if the employer signs an NDA with a customer, I would also be bound by that. It might be worth checking your contract, otherwise having a company sign an NDA doesn't hold much weight if their staff are free to go around sharing the information themselves.
- masswerk 4y agoControversial opinion: And this is why block syntax by white space is not for production.
- davbryn1 4y ago"What does this teach us? Well, it teaches me to do more diverse tests when doing destructive operations. It also should probably teach something to Vimeo and to my contractor but I doubt it will (and yes, the upload for some reason is still manual to this day. Go figure!)" So you wrote bad code, didn't test it properly, ran it on production on the Friday before a release and are blaming Vimeo and [name redacted]? And your resolution was yet another cobbled together script that you probably didn't test? This isn't a great article to have attached your name to
- gala8y 4y agoNot to mention that he _deleted_, but not _lost_ videos. Nothing to see here.
- chopin 4y agoI'd hire this guy if only being for this frank about his mistake. He owned it and that is what I would look for. After deletion, what should he have done? Postpone the go-live? That's often not a a cost-effective option. As for a risk-analysis the worst what could happen was deletion of the remaining videos. I don't think that that makes big difference in this situation. And to do the right thing, you have to have the infrastructure in place, if you are in a hurry. I doubt that's the case for a 10 heads shop.
- malexbone 4y agoAgree 100%. Acknowledged mistake, moved forward to find a solution. Reflected on lessons learned. Shared valuable lesson. To me this indicates intelligence, competence, integrity, grit and generosity. TechnicL proficiency is much easier to come by than integrity, grit and generosity. I would trust the author to deliver on commitments.
- davbryn1 4y agoOwning the mistake would be fine if he did that - he did'nt. He blamed the company he was contracting for. That's a big no from me
- wruza 4y agoCode without constant logging of “utc [who] does what exactly” is a no-go for me for a long time. Also, if you have to be destructive, replace the <rm/sell/halt> with log() for at least one time (aka --verbose --dry-run) and check your expectations. One-shot scripts like this are screaming disaster. (The problematic line lacks the closing ", probably a typo? I though it closed in an unexpected location)
- arein3 4y agoYou can automate using puppeteer or selenium
- dsego 4y agoThe author used Playwright in the end to automate uploads. Using e2e tools for automating tasks is clever, I'm not sure I would've thought of it.
- chopin 4y agoIt's clever, but also brittle. And might have disastrous error conditions (like hitting "Delete" instead of "Continue" if the wrong UI part has focus).
- BurningPenguin 4y agoI accidentally deleted a printer from the printserver by using a python script. The docs weren't exactly clear, so i thought it would only remove the local printer connection. After reading this post i feel better now. My fuckup wasn't that bad in comparison. :)
- NikolaNovak 4y agoHonestly, this is positively representative of any junior developer with comparable experience. Depending on their background and how much production work they had, there's an overwhelming sense of eagerness and enthusiasm. Quick to script and perhaps a bit too quick to execute. A friendly team will harness that enthusiasm and tame the quickness / encourage respect for production. We all made a massive doo doo and its how you proceed that'll define your career.
- RcouF1uZ4gsC 4y agoThis is one of those times that even if you don’t use a fully functional language, trying to make as much of your program logic pure functions would be helpful. It also makes it more testable. Instead of putting the delete call right in the loop, split it into four functions. function getAllVimeoVideos() function getAllDbVideos() function getVideosToDelete(vimeo_videos, db_videos) function deleteVideos(videos_to_delete) Your core logic lives in getVideosToDelete which is simply a set difference. Given that there are only a few hundred videos, it is easy to run the getter functions above and quickly verify they are returning what you expect.
- tomhallett 4y agoThis was going to be my exact recommendation. By “separating the concerns”, you make it easier on my pretty much every dimension: testing in unit tests, doing a dry run in production, ability to read the code (you and code reviews), and in some cases your code will be written in a more functional way reducing variable scoping issues.
- acutis_fan 4y agoYes that's fun. a List<Foo> getFoosToUpdate(List<Foo> foos, List<Bar> bars) function is the first time I thought about time complexity in my job. Say Foo and Bar have fields in common, such that you can say a Foo object "equals" or "matches to" a Bar object, like if they have name and dateOfBirth fields or something else that are the same (nothing like a common ID between the two). Now say there are some other fields too, like amountSpentThisYearOnDogFood that you know is always accurate for Bars, but might be out of date for Foos. How do you get the list of all the Foos to update? Initially I did the nested for loop solution that's like List<Foo> getFoosToUpdate(List<Foo> foos, List<Bar> bars) { List<Foo> returnList = new List<Foo>(); foreach (var foo in foos) { foreach (var bar in bars) { // check if "equal" or "matching" based on some criteria // if equal, update foo dog food expenditure with bar dog food expenditure, add to returnList, and break } } return returnList; } but that's O(n^2) right. The solution with a Dictionary is obviously better. All you need to ensure is that you have a method for both the Foo and Bar classes that will produce the equivalent hash for both, if they would be considered equal or matching by whatever criteria you are using. So you could have something like int GetHashOfFoo(Foo foo) { string firstName = foo.FirstName; string lastName = foo.LastName; DateTime dob = foo.Dob; return (firstName, lastName, dob).GetHashCode(); // convenient c# method } int GetHashOfBar(Bar bar) { string firstName = bar.FirstName; string lastName = bar.LastName; DateTime dob = bar.Dob; return (firstName, lastName, dob).GetHashCode(); } These two functions will return the same value if those fields are the same. So then you can do something like List<Foo> getFoosToUpdate(List<Foo> foos, List<Bar> bars) { List<Foo> returnList = new List<Foo>(); Dictionary<int, Bar> barsByHash = new Dictionary<int, Bar>(bars.Count); foreach (var bar in bars) { int barHash = GetHashOfBar(bar); barsByHash[barHash] = bar; } foreach (var foo in foos) { int fooHash = GetHashOfFoo(foo); if (barsByHash.ContainsKey(fooHash) { returnList.Add(foo.CopyWith(dogFoodExpenditure: barsByHash[fooHash].DogFoodExpenditure)) } } return returnList; } Which is faster cause you only have to go through the bars list once. I actually messed up something like OP with this, but with doing undesired additions instead of undesired deletions. You can think of it as having two endpoints, both expecting a .csv with rows being the things you were updating/changing/deleting. The problem was, there was a column to indicate (with a character) whether the row was for an edit, or addition, or deletion, but this was only with one of these endpoints. For the other, there was only addition functionality, but I thought changes and deletions were also options for the other kind of .csv due to some unwise assumptions on my part (thinking that the other .csv would have the same options as the other). That's how we accidentally put in over 100 additions that should have been changes that had to be manually deleted. Luckily I had a list of all the mistaken additions.
- iamben 4y agoI like these stories. I think they resonate well for 'the rest of us'. I've made plenty of mistakes like this - you learn and grow, right? One of the best things about HN is that so many incredible, talented people post. It's incredibly inspiring to raise your own game, to see what the best are doing. But sometimes it's equally important to realise we all fuck up, and for every unicorn dev there's another thousand of us grinding away. OP - well done for sorting the problem and telling us all about it!
- rossdavidh 4y agoAmen
- chanandler_bong 4y agoExperience is directly proportional to the amount of equipment ruined or data lost. Even though you were fortunate not to lose any data, you gained a lot of experience!
- LinAGKar 4y agoShouldn't that be `page={page}` rather than `page{page}`? Or better yet, use the requests `params` argument.
- andreagrandi 4y agoIt should really be something like: "a flaw in our system allowed me to delete 7am TB of videos". Not entirely your fault.
- mrkwse 4y agoSystem and/or development processes
- batch12 4y agoIt's like the first time you run rm -rf /path/to/delete/ * And realize it is taking too long...
- SnowHill9902 4y agoCan you explain? I feel like it removes / but not sure why.
- switch007 4y agoThe error is the space before the asterisk. The original intention was to delete the contents of the folder /path/to/delete/. Instead, the asterisk enumerates files in the current directory and they get deleted
- KarlKode 4y agoBesides recursively deleting /path/to/delete/ the command also deletes all (non hidden) content of the current directory (note the * at the end of the line). I assume the correct command would be /path/to/delete/*.
- Tesl 4y agoIt removes everything in the current directory
- pwg 4y agorm -rf /path/to/delete/ * Note the space between the last / and the * This will recursively remove the directory /path/to/delete and remove every file/directory that matches * in the current directory where 'rm' is being run. When what was most likely meant was: rm -rf /path/to/delete/* Note the lack of a space between the last / and . This will remove all files that match that reside in the /path/to/delete/ directory.
- lesgobrandon 4y ago
- SnowHill9902 4y agoRelated: is there any HTTP API model that supports transactions with commit and rollback? Also isolation levels? Usually one wants to set_stock(get_stock() + 10) but there may be competing from various clients between both calls, resulting in races. Usual web APIs seem vulnerable to this.
- jffry 4y agoWouldn't the model be to expose an increment_stock(10) type HTTP endpoint instead, and the backend can ensure it's atomic?
- p0d 4y agoFor many years I have had a private blog. I like to write but realised 99% of us are not interesting to read. This is a young guy processing his thoughts. Not "teaching" the rest of us as he frames it. This should have stayed in-house and personal. The company can then decide which clients, authorities to contact if necessary. There is a book in all of us as they say. For most of us it should stay there.
- johnklos 4y agoWe can all poke at this person for doing things incorrectly, but one has to wonder what mindset could lead to any programmer ever thinking that: 1) parsing a web page shouldn't be considered incredibly fraught with problems 2) that reloading web pages should be part of (1) 3) that this should ever possibly be run without validating the list of files that would be deleted So forget the specifics. Where are people learning these things, and what do we do to teach them better things?
- dboreham 4y agoCollege? Parents? In my experience it runs pretty deep so not sure it can be easily trained out. This mindset is probably quite useful in evolutionary terms: rush at the attacking bear without thinking, for example.
- plonk 4y ago> rush at the attacking bear without thinking, for example Would that work? I don’t see a bear backing down and I don’t see the human winning either.
- dncornholio 4y agoSome mistakes can only be learned by making them. Sometimes you can tell someone a hundred times something, they won't learn until they experience it. The point is not to prevent these mistakes, but to keep the consequences low. Have backups, have version control, etc.
- ufmace 4y agoTrue, and worth remembering why. Most of us are constantly getting warned about the dire potential consequences of huge numbers of things, most of which are either massively unlikely to ever happen or not actually that bad, or both. It's very difficult to tell which of the things we get warned about are actually high risk until something bites us.
- qayxc 4y ago> Where are people learning these things, and what do we do to teach them better things? Learn to learn and learn to work carefully. It starts in school and should be part of a proper college/university education or vocational training. There's several ways of learning the specifics: by experience on-the-job, which can be hard if mistakes can get you fired; or by putting in the work in your free time. If your job is to work with certain web frameworks and you're not very experienced, either ask senior devs to assist/review before going live with critical changes. Alternatively, practice at home. Unpopular, but you need to get experience from somewhere. OSS projects are a great way to do that - be that by creating your own or by contributing to an existing one.
- wumms 4y agoNot completely off topic (as one of my scripts deleted files recently which dates were off by one): > Fri May 06 2022 > I'm currently working [...] in Italy
- ricardobayes 4y agoAny process that makes a junior directly access prod codebase/database is flawed. No matter how small of a company you are, you can set up a proper CI/CD pipeline.
- thevinter 4y ago90% of IT companies in Italy don't even know what a CI/CD pipeline is. That said I don't think it's something we could've integrated in our pipeline as it's an error that originated from an external service!
- progx 4y agoNow you learned what a backup is.
- lpointal 4y agoHow can any enterprise only rely on such online services and not keep copies of their job on their own storage ? At least store in large TB hard disks connected with a SATA adapter when needed, and put them in a case in a safe place (better: two copies, stored in two places). What is the HD + copy time price relatively to production work ?
- unfocused 4y agoI'm currently working with FOIA software, and a regular user can only delete one document at a time from the information that they verify/redact before sending out. They can't even multi select! Only an admin can delete multiple documents at one time. I'm guessing users accidentally deleted multiple documents one too many times, and now it's baked in.
- rexreed 4y agoA big part of the reason for the problem in this post is because Vimeo made it impossible to move videos from one Vimeo product to another Vimeo product: "There were roughly 500 videos on VimeoOTT that had to be transferred to Enterprise and Vimeo doesn't provide an easy way of doing it." I have found working with Vimeo to be very frustrating, especially recently. They have a great video solution, especially for streaming, but they seem to put these unnecessary and frustrating roadblocks that make me constantly question my decision to use Vimeo. From in ability to move videos from one place to another, requiring complete uploads (resulting in problems like this post) to nonsensical limits and pricing, especially on their new webinar offering, which has a limit of 100 registered attendees. For anyone who has run webinars before, this makes no sense since 100 registered attendees usually means 20-30% of those people actually attend, so you're capped at 20-30 live attendees. They should price it like most event sites and charge per live attendance rather than registration. Regardless, I've been very frustrated with Vimeo since it could be so much better if they didn't have these roadblocks in place. If they could have easily enabled moving videos from one product to another, the post (and 7TB of lost videos) would never have happened. It wasn't always this way with Vimeo, but they went IPO in May 2021 and it's no surprise they're turning the screws on their product offering and pricing now.
- Reason077 4y ago> "What does this teach us? Well, it teaches me to do more diverse tests when doing destructive operations." I think it also teaches us that adversity sometimes leads to better solutions. I love that the OP made a hacky script that did in 4 hours what a guy was paid to do manually over several months!
- dncornholio 4y agoWhat is the f doing in url = f"https://api.ourservice.com/media?page{page}&step=100 https://api.ourservice.com/media?page{page}&step=100 ?
- fifticon 4y agoif it's python, it's the formatting/interpolation string marker.
- jraph 4y agof for format ("formatted string"). It does the same thing as `https://api.ourservice.com/media?page${page}&step=100 https://api.ourservice.com/media?page${page}&step=100` [sic] in Javascript, or "https://api.ourservice.com/media?page$page&step=100 https://api.ourservice.com/media?page$page&step=100" in Bash, PHP, Perl or Groovy (and other languages). It outs you into variable substitution / interpolation in the string literal. In Python these string literals are called f-strings if you want to look it up. They are defined in PEP 498 - Literal String Interpolation [1] and available since Python 3.6. [1] https://peps.python.org/pep-0498/ https://peps.python.org/pep-0498/ [sic] there probably would be a missing '=' in this url after "?page"
- qwertox 4y ago"f-strings", a (new) way to format strings.
- throwaway744678 4y agoIt's a Python f-string [0]. A way of formatting a string by directly including a Python expression between curly braces. [0] https://docs.python.org/3/tutorial/inputoutput.html#tut-f-strings https://docs.python.org/3/tutorial/inputoutput.html#tut-f-st...
- photon-torpedo 4y agoApart from all the advice on how to do such destructive operations more safely, I think there's also a lesson to be learned about communicating more actively: 1. Vimeo responds to the original request with "will look into it", then... nothing happens? This may depend on culture, but at least from my experience in the UK, this is a very non-committal response, and if you really want them to do something, you'll need to chase them. Wait a few days and inquire if they have any estimate for when it might get done, or if they need more information. I find that the "looking into it" response is sometimes used to gauge how important the request is to you. 2. Once you go with your own solution, just drop a quick message to Vimeo: "Hey, just wanted to let you know we've found our own solution for this, and won't require your help any more. Sorry if you've already committed any resources for this task. Have a nice day, yada yada." This not just avoids what happened here, but is also a courtesy to them.
- muglug 4y agoThe root of this particular issue was Vimeo's failure to do this migration for their customers. Vimeo OTT has a codebase written in Rails, whereas the main PHP application is written in PHP. At the time Vimeo acquired Vimeo OTT's codebase, the Vimeo OTT codebase was small — around 10,000 lines of Ruby. Rewriting that codebase inside the Vimeo PHP application would have been a tough technical challenge for the all-Ruby team, and they'd have likely lost some people along the way and missed out on some content deals, so they decided instead to maintain two separate codebases and two separate login systems. The video-playback and video-storage infra has since been unified, but all the business logic is still siloed.
- macspoofing 4y ago>The root of this particular issue was Vimeo's failure to do this migration for their customers. Yes and No. At the end of the day, you as a business have to insulate yourself from your infrastructure provider.
- notyourday 4y agoVimeo is the only infrastructure provider providing that service. It is impossible to insulate a business from it.
- macspoofing 4y agoYou're saying it's impossible to not accidentally delete 7TB of videos, and when you do, to blame it on Vimeo?
- conductr 4y agoHe wasn’t asking them to refactor their internal code bases. But they should be able to whip up the 20 lines of code needed to do this between APIs (or just directly on their servers). Essentially what author was trying to do when he screwed up. For the author this was disposable code, for Vimeo this would have been a reusable utility. I know how these things happen. Support ticket queues and all. And while I don’t fully know the difference in cost, I would assume a customer upgrading to an Enterprise plan would get a better support experience. Whoever within authors company negotiated the upgrade to Enterprise (or didn’t) and failed to embed some agreement around OTT to Enterprise transition assistance was the one who made the first mistake.
- BillyTheKing 4y agoFor larger 'live' production changes I've now started to rely on generative programming. I've got one script in some 'normal' programming language like javascript, or python, which in turn generates a script that contains a list of curl or other cli commands which do the actual deletion, modification, addition, etc. This allows me to run a small sub-set of commands and test those under a live-environment before running all commands at once. In addition, this also functions as a complete log of what has been changed manually in production.
- desarun 4y agoOh dude, we've all been there. 9 years ago I was working for a major broadcasting company in the arse end of London as a junior dev, building one of their Android apps. We'd roll features out months before & enable them with feature flags via a json file we'd manually push to a prod server at a later date. We'd just built a huge new feature letting you request content to be downloaded to your set top box remotely & it had a 250k marketing campaign to go along with the launch. Senior dev trusted me with prod deployment rights. I pushed the wrong json config to prod, launching the feature weeks before the marketing campaign. Thank god I was a junior perm, that was definitely a firing offence.
- hayd 4y ago> Senior dev trusted me with prod deployment rights. That part's crazy! If you think it was a firing offence wouldn't they've been fired? (I don't think it is, but obviously requires system changes/explanation.)
- shantnutiwari 4y agoWhat negativity and arrogance in the comments here. Jeez, it's like no one HN ever made a mistake, a bunch of 10xers ninja programmers here. Please read this: >I also want to preface this whole post by saying that I'm a Junior Developer with less than one year of actual experience. Some of the things that might seem obvious to some might not be so for me, thanks! It's just some kid sharing a mistake they made and owning up. Ease up on the "LOL what an idiot" attitude
- FunnyLookinHat 4y agoI was actually really impressed with this individual! For someone who has less than a year of experience, they're showing quite a bit of initiative, drive, and curiosity - which really are what make or break engineers as they develop. Taking the time to do a blog post (effectively a post-mortem) and share it is even better! And yes - I've literally done this exact same error (with TB of video data!). Spending the following week remediating all of that data loss was a great lesson in patience and attention to detail. :-) OP: If you're ever looking for a job be sure to send me a message. Contact info in profile.
- Moru 4y agoMy mistake was on floppy disc with source code, other text files and images. Was hand editing (in hex disc editor) the floppy to get back the data, sector by sector. Fun times. Not going back there though :-)
- nso 4y agoMine was a DELETE FROM Users; WHERE... Fun was had.
- codegeek 4y agoUsually the recommendation is to not start writing the DELETE query first. Write the SELECT query first and see the results. If you miss the WHERE clause, you will see that immediately. Then change SELECT * to DELETE. But I assume you have learned that lesson already :)
- uptown 4y agoJunior Dev: "I'm under an NDA" Also Junior Dev: "Here's my source code"
- esprehn 4y agoUnless you're Oracle that code is hardly critical to the business. Even as a Sr Dev I'd share stuff like that, it's code that'd appear on a stack overflow post anyway.
- thisNeeds2BeSad 4y agoThe only thing that I can remember helping against such actions, is the exponential need for confirmation by intent. Means, if you delete one small file you need one confirmation, if you delete thousands, you need a intent stating i expect thousand files to be deleted. Same goes for size. So not a okay button, but instead a form allowing you to enter the dimension of the intented outcome. 100 files max, 1 gb max deleted. If the request goves over the intent, the system aborts.
- kirillzubovsky 4y agoMistakes happen. Kudos to the author on taking it as a learning opportunity. I am friends with a lot of smart devs, and many of them have dropped a production db at least once, and if not then, then accidentally emailed 10k people …etc. It happens. Work to avoid it, but plan for what to do when it inevitably happens. ¯\_(ツ)_/¯
- aristus 4y agoHey, everyone, ease up. I have: 1) dropped a production database because I thought it was the test database. 2) screwed up a print job costing $100,000 in today’s money and had to do it again 3) crashed all of Facebook with a C++ bug. 4) crashed Facebook photo uploads, with a JavaScript bug, in my first month. 5) literally killed a startup’s cash flow and caused them to lose their merchant account because I over focused on the wrong bugs.
- hbn 4y agoAt my first development job (paid internship at a moderately-sized, though fast-growing business - maybe 300 people at the time?) I introduced a bug that didn't appear until a certain microservice stopped working (my code defaulted in the wrong direction when the ms failed) and as far as I can tell they may have lost or almost lost a pretty big account from it. In an after-hours meeting regarding the issue, one of the higher ups ended up storming out and never showing up again. In my defence, we had to get 2 PR approvals before anything was merged! But I definitely learned a thing or two from that experience
- deleted 4y ago[deleted]
- paintman252 4y agoYou worked at Facebook, we get it
- qwertox 4y agoAaaahhh, the feeling you get when you notice that you fucked up. Everything gets quiet, body motion stops, cheeks get hot, heart starts to beat and sinks really low, "fuck, fuck, fuck, fuck, fuck, fuck, fuck, fuck, fuck, fucking shit". Pause. Wait. Think. "Backups, what do I have, how hard will it be to recover? What is lost?". Later you get up and walk in circles, fingers rolling the beard, building the plan in the head. Coffee gets made.
- cntrl 4y agodamn, your description is spot on and reading this triggered PTSD in me... Last time I had this feeling was two years ago when I destroyed one of our development servers because of a failed application update. I know exactly how I wished Ctrl + Z to exist in real life... We had backups of the machine, but it was still kind of a humiliating feeling to tell everybody and ask for restore from backup (everybody was cool though in the end)
- AlwaysRock 4y agoGod the feeling of having your body temp rise based purely on realizing you fucked up is so relatable.
- sergiotapia 4y agoI lost 1hr and 30 minutes of a Slack like app (chat messages). Luckily at the time we were pretty small so not much data was lost but holy shit did that make me almost throw up. Thank God my automatic backups were so close to the mistake I made and I didn't lose 24 hours. Haven't made a mistake like that since and I don't destroy DB records like that anymore.
- Oarch 4y agoPoetic! Love it
- deltarholamda 4y agoPffft, it's not a real panic until you weigh the pros and cons of leaving the country with nothing but the clothes on your back and becoming a illegal immigrant shepherd in a nation with too many consonants in its name. (Your description is so, so, spot on.)
- macspoofing 4y ago> but at the time the code seemed completely correct to me I venture this kind of (misplaced) over-confidence is not atypical of many junior developers. As someone with a few years under my belt, I don't care how sure I was of the code I wrote that deletes important data, I would have gone through the code over and over again, and at least ran a simulation (by maybe logging the generated delete urls for manual verification). It's a rite of passage and we all went through something like this. It's how you learn and grow. >It also should probably teach something to Vimeo No. Even if Vimeo could have made things better, it's still your fault. You have to take responsibility for your business. At the end of the day, if this causes the closure of your company, Vimeo is still fine.
- ElCapitanMarkla 4y agoNice work :D I tend to always add a `--dryrun` flag to any scripts like this these days so that when we move it to production we can run an extra test there just to be sure.
- JasonFruit 4y agoI believe if we're honest, we've all done stupid things we should have avoided. I remember a group of about 3000 emails that went out to insurance agents saying that policy #123456789 for Someone Funky was going to be cancelled by underwriting. I also remember very quickly figuring out how to automate Outlook's email recall feature. We've all made big dumb mistakes. Recover and learn.
- 0xbadcafebee 4y agoThis is more common than you think. Not just losing data, but not having a good handle on where the important parts of the system are, and how close you are to catastrophe. I find diagrams really help. I can recall a visual map of the system when I work on some component, and think, "OH, I remember seeing this component connected to a really critical thing, I need to check something first." Start by creating one empty page for every component of your system. You won't remember them all, but over time you can add missing ones. Each page is the authoritative source of info on that component. If you need more pages for one component, put them in a directory of the same name as the page and add ".d" to the directory name, and link to them from the first page. Finally, create a diagram (however you want) that includes every component you have a page for. Add the count of components to the top of the diagram. If the count on the diagram doesn't match the number of documents, time to update the diagram. If you ever add, remove or rename a page, time to update the diagram. If you do this the same way for every different system you have, you can link them all together and get both small and large scale diagrams. (p.s. don't waste time automating this unless you find the system changing constantly or you have a very big system)
- bbbush 4y agoscary. maybe as well just pay vimeo to restore data.
- franciscop 4y agoThis is a great technical write up, I'd love to hear the human side of this story as well! When did you tell the higher ups that you deleted production? Was no one more senior on call to try to fix it? Did they want you to learn how to fix it? Or were you the most senior responsible for this whole area? Or did they don't know?
- thevinter 4y agoThe first part of my write up slightly explains it but the point is that HN is the top 1%. In my current company we have 10 developers, most of them without a technical degree. They know how to do what they've been doing for the past 10 years but (as with most small companies here in Italy) people don't know what best practices are used in the industry, what a pipeline is or what a dry-run is (I learned about it today myself!). What happened is that no one knew how to react and I was probably the best suited for it, we don't really have seniority in office. That said when I deleted the videos I immediately told my boss. He was kind of scared but his reaction was mostly "Well, now we have to re-upload them immediately, find a way. The people that uploaded them once won't be doing it twice". I was basically left on my own to find a solution (which I luckily did). Please note that I'm in no way blaming my company or accusing it of something, this is the standard knowledge base and way of dealing with things in many places, contrary to what working in big tech or reading HN might make you believe!
- franciscop 4y agoThanks for the explanation, that makes a lot of sense! > "HN is the top 1%" + "this is the standard knowledge base and way of dealing with things in many places, contrary to what working in big tech or reading HN might make you believe!" I'm in fact from Spain and now live in Japan, and I believe the practices in Spain would be as bad as Italy, and in Japan they are def worse (great at hardware, horrible at software), so I do understand a lot of what you are saying. FWIW, in Spain I've seen whole dev teams composed only of interns! > "we landed a big contract for one of the biggest gym companies in Italy, the UK and South Africa" + "we don't really have seniority in office" Maybe now that seems like you have the budget it's a good time to go to management and suggest to hire some senior devs who can mentor the rest into learning best practices? You can sell it like a reinvestment in the company to management if they want to take it as pure profit. If Italy is like Spain, many devs won't really even want to learn these things, but some will and then those will become seniors at some point.
- tomkwong 4y agoFirst, I want to say that this is a great post. You always grow stronger when you make mistakes. Writing it up solidify understanding in the learning process. This story resonates with many people here because many experienced engineers had done something similar before. For me, destructive batch operations like this would be two distinct steps: 1. Identify files that need to be deleted; 2. Loop through the list and delete them one by one. These steps are decoupled so that the list can be validated. Each step can be tested independently. And the scripts are idempotent and can be reused. Production operations are always risky. A good practice is to always prepare an execution plan with detailed steps, a validation plan, and a rollback plan. And, review the plan with peers before the operation.
- notyourday 4y ago> 1. Identify files that need to be deleted; 2. Loop through the list and delete them one by one. > These steps are decoupled so that the list can be validated. Each step can be tested independently. And the scripts are idempotent and can be reused. This is the most underrated comment. I'm saying it as someone who had the ultimate oversight of deleting hundreds of TBs per day spread of billions of files on different clouds and local storage.
- spiffytech 4y agoI've never regretted treating tasks like this as a pipeline of discrete steps with explicit outputs and inputs. Sending output to a file, viewing it, then having something process the file is such a great safety net.
- orange_puff 4y agoAs everyone else has already pointed out, better testing would have been very useful here. For instance, print(len(our_ids)) would have been a dead giveaway that that something was up I am also a junior dev and completely empathize with being given a lot of responsibility and potentially messing up. I think for someone with < 1 year of experience, to solve the problems you created as fast as you did is really impressive. Thankfully your story ends well :)
- aasasd 4y agoAfter having read about plenty of such cases over the years, I have a persistent dread of pulling something like that myself, to the point of being nervous with ‘*’ in the terminal, and generally checking everything twice. (And also have some kind of mild horror-high from corporate snafu stories, weirdly reminiscent of Ballard's ‘Crash’). So: I never feed the data straight from the gathering script into the modifying script, at least not in the first runs. Instead, I dump the whole list of items into a file, count them in there, gawk at them to see that they're right, and compare with the source data by hand until I begin to annoy myself. Then I feed that file to the second script.
- amtamt 4y agoA computer lets you make more mistakes faster than any invention in human history, with the possible exceptions of handguns and tequila.
- mindcrime 4y agoImagine coding while drinking tequila...
- _jplc 4y agoEveryone makes mistakes, juniors and seniors alike, but I consider you have the right mindset and resolutive skills that will make you thrive :)
- havkom 4y agoThe company was lucky to have someone like you that could actually sort out real problems efficiently. I would bring up this story when negotiating for a raise.
- deleted 4y ago[deleted]
- DeathArrow 4y agoThere's a thing called unit tests.
- DeathArrow 4y agoThis wouldn't be an issues if providers like Vimeo would soft delete and hard delete the items after a period of time, allowing recovery between. Everywhere I have to implement a delete operation, I never hard delete data on first call.
- DonHopkins 4y ago>... the "Silicon Valley" world ... To rebillionizing! https://www.youtube.com/watch?v=wGy5SGTuAGI&t=369s https://www.youtube.com/watch?v=wGy5SGTuAGI&t=369s ...yeah, the Tres Commas bottle was on the DELETE key. The corner of it was just, it juuuust got on there...
- alkaloid 4y agoDoes anyone else get that deep, dark, disturbing feeling in their gut when they know they have done something bad like this? This is why I use so many print statements and comment out destructive actions! Lots of experience with these feelings!
- furyofantares 4y agoGreat post and great attitude. I think I would reflect on why this is a script to begin with. It's run once and with only 500 items could be done manually, though 500 is certainly a bit much. But it's not a massive time saver; the point of the script should be almost entirely to increase accuracy. I think I would write one script to generate the list of videos to delete; that's the part that's actually difficult, and a human can then verify the list. I would probably just delete them by hand after that, but if I really wanted a script for that part too, it would be a separate script that uses a list that has been vetted by a human even if initially created by the first script.
- Fritsdehacker 4y agoThis is why you have backups. Good on you to have them! When I just started as a junior dev at a small company I made the classic mistake of emptying the prod db instead of my local dev db. This was a small and in hindsight insignificant project. But Google was our customer, so it didn't feel insignificant at the time. In this case my inexperience was partly my savior. All the data was inputted by people via a web form. Normally you're supposed to use POST to submit a form. But I was quite clueless at the time, so I had used GET. This meant all requests were still in the Apache logs. I could simply replay all requests. I still feel my hard pounding when I think about the moment I realized what had happened. I was really relieved when everything was back! What I learned from this incident: - make automated backups - no access to prod db from anywhere but prod
- cassandratt 4y agoYea, I’ve wiped out an entire government’s form library once. Backups are a career saver.
- stareatgoats 4y agoA great success story as far as I'm concerned, even if it doesn't reflect well on Vimeo support. But a good reminder to have someone doublecheck your logic if you aim to delete massive amounts of data from production. And to check if the backups are working (producing restorable data) on a regular basis. Sometimes they just seem to be working, as I have learned the hard way...
- deleted 4y ago[deleted]
- AtNightWeCode 4y agoThe conclusion should include that backup at separate locations is key. Also, that the backups are tested and work. I worked with clients that had everything from lightning strikes destroying servers to ransomware to people making mistakes. No problem with solid backups. There is a difference between a good process and skill.
- birdyrooster 4y agoIs 7TB a lot? Peers at personal arrays at orders of magnitude greater.
- deleted 4y ago[deleted]
- vjust 4y agoSo much wisdom in these comments, people have different styles of being careful, and each makes sense in a nuclear "go" situation
- deleted 4y ago[deleted]
- hanly_paul 4y agoI am also a junior with 1 year’s experience, just in Python but none with the requests module or web development. If the ‘page’ variable is being changed, was the error something specific to this module, not refreshing the page?
- beeforpork 4y ago> I Accidentally Deleted 7TB of Videos ... Spoiler: But there was a backup that could be reuploaded in time and everything was fine in the end.
- urbandw311er 4y agoWould you have had the courage to post this here if you hadn’t been able to fix it?
- dclowd9901 4y agoHis solution reminds me of how I used Cypress to generate test accounts on our local admin dashboard for Cypress tests, since our api was inadequate (it didn't do the billing signoff required to create accounts that last longer than a month... don't ask...).
- bufferoverflow 4y agoAlways do a dry run when deleting many things with code. - Captain Obvious
- mbostleman 4y agoRelated: The change is fine, it's only one line.
- donalhunt 4y agofwiw I would probably have turned to rclone.org for this. It doesn't have support for vimeo out of the box but the Vimeo API seems sane enough that it would be trivial to implement uploads quickly. Previously used rclone for doing massive transfers between cloud providers using "cheap" on-demand servers which provide unlimited data transfer (the public clouds make this very expensive).
- ge96 4y agoThe product I work on, I can watch the events occur afterwards (videos of people using it) and it's so embarrassing watching it fail. The wasted time. Ahh... I've gotten better to check deps and run a full automated E2E test everytime new code is deployed (before/after diff envs). Still things happen. Hopefully you have a large enough client base where some bad experience doesn't define the whole thing.
- IYasha 4y agoSo, apparently, vimeo has better support than youtube (not informative, but at least they DO something). Duly noted.
- RankingMember 4y agoI'm impressed you went with an automated solution (PlayWright) for 500 videos after all that, considering they could be cross-loaded from Google Drive almost instantaneously. I'm glad it worked, but coding around a screw-up under the gun seems like a high-risk operation compared to spending 4 hours doing the task manually (albeit being super bored the whole time), but with the benefit of knowing it's being done correctly instead of hurriedly writing a script to potentially do something else wrong very efficiently and dig your hole deeper.
- bruhbruhbruh 4y ago+1 to this. After the few major screw-ups I've caused at work, my self-confidence in my coding ability is rocked, and I tended to react by erring towards manual cleanup, rather than coding some scalable solution for fixing the issues
- leokennis 4y agoActually I was surprised reading that the person wrote a script to delete 900 videos. If you need to do it once, it’s probably 2-3 hours of work? That is identifying a duplicate video and then clicking the button(s) to delete it once every 20 seconds. Reminds me of https://xkcd.com/1205/ https://xkcd.com/1205/
- hexsprite 4y agowhen doing migrations/conversions I always write a script in dry-run mode first. I exhaustively check the results to make sure they are expected. Then try to do a real conversion/transfer of only the 1st file and make sure that worked. Then do a couple more. Etc. Only then do I feel confident to do the whole thing.
- mastazi 4y ago> Vimeo doesn't provide an easy way of doing it. I wrote to the support team around October asking them if it was possible to do a migration, and they told us that they "will look into it" without letting us know anything ever since. [...] At one point, without letting us know anything, Vimeo decided it was a great idea to comply with our request and dumped all the videos present on OTT onto the new platform. No questions were asked [...] they were duplicating videos that were already uploaded. Oh yes Vimeo, the crappy company that won't let you play videos unless you enable autoplay in your browser[1]. Selecting them as a provider was the actual mistake. [1] https://askubuntu.com/questions/777489/vimeo-video-not-playing-in-firefox https://askubuntu.com/questions/777489/vimeo-video-not-playi...
- lnxg33k1 4y agoBut are you a junior dev with less than one year of experience working by yourself alone at a company? No tech lead/help?
- mikotodomo 4y ago> Some of the things that might seem obvious to some might not be so for me, thanks! > my mind thought that url would refresh itself as soon as the page variable changed This is what I thought too when I read the code. I don't think it's obvious at all!
- xmprt 4y agoThat's actually surprising to me. In most languages that I've worked with, strings are immutable so the fact that url doesn't update is more obvious to me and I'd be surprised if it did update.
- peter303 4y agoHappened to Pixar Toy Story 2 too. https://thenextweb.com/news/how-pixars-toy-story-2-was-deleted-twice-once-by-technology-and-again-for-its-own-good https://thenextweb.com/news/how-pixars-toy-story-2-was-delet...
- notaplumber1 4y ago> .. physically backed up in a Google Drive folder ... That's not what a physical backup means.
- cat_plus_plus 4y agoYou are fine dude, you didn't delete any videos, only high availability cache of videos on a streaming site. If that was your master copy, you probably would have taken greater care, if not :-). Anyway, when working with caches that can be recreated in reasonable time, it's normal to take less care than when it comes to originals. The only concern is Google Drive as the only backup, please make sure you have a local copy on a local RAID drive and another one regularly archived and stored in a bank locker.
- brunooliv 4y agoKudos to you for "learning in public" by showcasing part of your learnings online!!! I think this is extremely important to do! Not everyone is an innate rockstar developer who provisions k8s clusters for breakfasts and delivers features for lunch! Being a developer is a really hard job and there are endless complexities and difficulties along the way and when we are more seasoned already. Don't let any negative feedback deter you from keeping doing what you're doing: learning from your mistakes and improving along the way!