6 ms·
This is why he is Data Engineering Coach and not actually responsible for production systems now. (apparently) Everyone loves reading tech horror stories and p
by psim1 5y ago
This is why he is Data Engineering Coach and not actually responsible for production systems now. (apparently)
Everyone loves reading tech horror stories and peeling off the take-away lessons. My lesson would be, don't hire this guy and don't use him as a coach. He's careless!
"But this kind of thing could happen to anyone" - sure, anyone who is irresponsible.
Am I ranting? Let me continue. I work on a team now with someone who seems to consistently forget to save router configurations and another person who yawns through meetings because he doesn't sleep. These guys make careless mistakes and we "learn" from them. Except we don't learn the most important lesson: Irresponsible people need to be off the team. They make more work for everyone else and make the team look bad.
But I guess the tech horror stories that amuse everyone on the forum or at the pub are better than saying, "I am responsible and do good work."
- whatever1 5y agoNone of the code for consumer production is verified formally. So please spare us the bs that you make no mistakes. You are just lucky. Be humble because the complexity of modern systems is insane, there is no way you have all cases covered, if you had you would had a formal proof. We all are just doing our best to cover most of the edge cases. That is why we need to keep learning from other peoples mistakes.
- sdhfjg 5y agoRead the article. We're not talking about formal verification. >It's Sunday morning and I just discovered that I've lost 3To of data and that all data pipelines have stop working because on Friday I ran for no reason hdfs dfs -rm /data This is profound incompetence.
- arpa 5y agoyeah this is not something to be proud of. At least make up a reason!
- snicker7 5y agoPossible explanation: lots of terminals paste on click. A single miss click can execute who know what from your clipboard.
- nicolas_t 5y agoIn my experience some people keep making the same careless mistakes, the first time, you let it pass, treat it as a learning experience. The second time, you start seeing that it's always the same person doing the same mistake. Copy pasting is not an excuse, before you run anything destructive, you double check what you're running. Anyone who is responsible will double check before running this kind of command, and if I can't get that person off the team, I'd severely restrict his access (in general, I think most people in a team should not have access to production data). And we do have disaster recovery plans so there's very little that could be done that would be catastrophic. But still, a lot of the disaster recovery plans call for downtime because it's not worth the cost benefit to engineer the system to be completely resilient to idiocy.
- ac50hz 5y ago+10
- teddyh 5y agoModern terminals marks pasted data as pasted, and similarly modern shells detect these marks and do not run the pasted data immediately, but shows it, highlighted, so that you can review it before confirming it.
- macrolocal 5y agoBack in 2014, my first tech job, I wrote a clean-up script to delete HDFS artifacts listed in some text file. One day I modified the list and left a blank line at the end. :) We had nightly back-ups though.
- hulitu 5y agoThat's why you have a backup. I also noticed that i'm incompetent on friday evenings so i avoid doing sysadmin work then.
- eftychis 5y agoI mean if people did not make mistakes we would all be writing directly in binary target or assembly. Should you ever delete folders named data -- probably not without knowing what you are deleting. Should you place your production data on a single machine -- definitely not if you know anything about production or complex (file) systems. I mean it could be literally your computer dying on you. You trust your company on that single hard drive not dying really? Or any system bug -- OS/firmware etc? Are you that crazy? Just short S&P long term -- it's less risky.
- lmm 5y ago> Should you place your production data on a single machine -- definitely not if you know anything about production or complex (file) systems. I mean it could be literally your computer dying on you. You trust your company on that single hard drive not dying really? Or any system bug -- OS/firmware etc? Are you that crazy? Just short S&P long term -- it's less risky. OP was using HDFS which is a distributed/clustered filesystem, their data was replicated across multiple machines. But "rm" still deletes it from everywhere.
- eftychis 5y agoYeah I should have made yhat a bit more explicit with the single hard drive comment. Thanks @lmm!
- arpa 5y agoMistakes happen. People forget which environment they are on. people forget where statement in their DELETE query. People misremember their cwd prior to running rm. Sometimes it's lack of experience, sometimes lack of sleep, sometimes just shit luck. Very rarely it's lack of responsibility. And regarding your team members: if you don't trust them, don't rant about it on public forum, man, move them away from important systems, implement personal improvement plan, and when that fails, fire them.
- winternett 5y agoExperienced solutions architects prevent mistakes from being costly with responsible infrastructure design, disaster recovery plans, and solid management workflows. Companies cut corners by not hiring architects because they are not cheap, and then they cut corners on salary for roles on mission critical projects because they are also cheap, and they often pay the price anyway just later-on because failure is expensive. Hire experienced people for the right money if it's a mission-critical system. The simple facts are all there. Nobody successful designs a good building without a good architect, and the project never turns out well when pay isn't right. Also important to note: The architect should not also be the one to physically build the "house", in order to maintain proper objectivity and in order to avoid conflict, even if the company insists on being cheap.
- ac50hz 5y agoHmmm, what you’ve described is lack of responsibility of recognising one’s own failings. There’s no excuse for that.
- arpa 5y agoI do not see where I said that. I am sorry, but you're projecting. Safety rules are written in blood; production safety rules are written in critical incidents and cold sweat. To err is human.
- short12 5y agoIn jest but personally I don't think you are a real programmer unless you have done something like this. It's kind of "welcome to the club" right of passage
- tomcooks 5y agoGood old "human error", be glad it happens for it means you still have a job that hasn't been automated (yet). Easy to mess things up when you routinely execute a bunch of commands everyday, many of which look similar in form but differ greatly in the pains they can create
- ricc 5y agoMaking mistakes like these in dev/test/sandbox/playground environments is one thing, but making them in prod over and over again is different. Yes, everyone should experience making catastrophic mistakes but please spare some effort to only do it in non-prod...
- short12 5y agoMany places don't have those separated environments and discipline
- tomcooks 5y agoFirst take the git log out of your eye, and then you will see clearly to take the specs out of your brother's report. Friendly advice: you don't sound like you rant, you sound like something else.
- hulitu 5y agoThat's why you have quality assurance. People make mistakes. Your process shall be able to detect and correct them.
- blef 5y agoI really like the effort you made in following more than 2 links to see what I'm doing. The issue is that you also don't have the full context on everything that happened. I don't know on which kind of teams you work in, but I personally worked a lot in teams where I was the only person responsible to do data related tech stuff for other people. I never said that I wasn't responsible but I was alone, so setuping backup was often the last of my priorities. I was and I did all my best to fix things as fast as possible and to remove any issues it could lead in the future. When I was setuping Hadoop cluster in 2014, almost no-one knows what it was in my local market. So let us try, then fail and learn. If you also think from 3 stories I'm writing on the internet that I'm careless so be it. But I also recommend you to fix your "router configuration" problem rather than ranting on posts on internet.
- ac50hz 5y ago+10 I’ve worked with similar people and had similar experiences. They are negligent, never question their own abilities nor actions, and to compound this they don’t learn from the experiences. Whilst no-one is perfect, it’s always possible to test before proceeding to cause uncontrolled havoc. Just remind oneself, check, check, check. Assume that there will be problems before proceeding; forewarned is forearmed.
- bravetraveler 5y agoAs they said in 'The Expanse'; welcome to the churn. I've been on both sides of this - the person that wants everyone to think twice before they type... but also the one that messed up. I'm really here just to say that... sometimes things aren't so clean-cut. For example: 'Business needs' often push myself and others to make calls that we'd normally never make. Coming from the top, down - executives rarely care about 'fact'. Just what they promised, and everyone suffers for it. In the end, the whole experience is a wash. We don't hit the target, and everyone is more stressed out. Repeat this enough, the diligence of most will falter.
- KronisLV 5y ago> Except we don't learn the most important lesson: Irresponsible people need to be off the team. This is false dichotomy. If people can screw something up, sooner or later they will. It doesn't matter how responsible or irresponsible they are - the processes in place should prevent mistakes from being made regardless of that, or to mitigate their consequences if they're unavoidable. Anything less and you're not addressing the root cause of the issue. These same processes should ensure that no one can create changes to the state of the system before them first going through another set of eyes and being validated. Most sane OSes at least prompt you before deleting a file - the very kind of safeguard that makes you double check whether what you're doing makes sense. Similarly, you'll notice that cars have seat belts and air bags, even if you're not going to crash daily. Get the irresponsible people off the team if you'd like, but if you can't put the appropriate processes in place to prevent mistakes, you probably need to rethink your priorities and provide the adequate pushback against "the business", when they expect you to SSH into prod and do anything. What that looks like in my current environment: - all server configuration is managed through Ansible and Git - developers have read only access to the servers when needed, no one can change the state of the system there - all changes are versioned and use merge requests and automated CI processes, need to be reviewed first and synced with the change management system, to know who is attempting to change something, why and also how (with the appropriate permissions) - furthermore, the Ansible processes are scheduled to run every day as well, just to make sure that the server status is as expected - of course, the changes to servers are also first done on development environments, then on accept testing environments, then on clients' testing environments (any number of them that's necessary) and then finally to prod - there is Zabbix for monitoring the infrastructure with alerting in place, as well as Skywalking for application APM, as well as some lower level tools - the apps are also shipped as containers, which have very similar processes in place, e.g. full CI run on every merge request, before it gets merged - those containers run in clusters, so if any of them fail for whatever reason, the load will be balanced as necessary and/or restarts will take place, each container also having health checks - lots of manual QA in there as well, to catch the things that aren't easily automated, as well as regression testing - automated integration and load tests done every now and then, to check that there are no regressions in a new release - testing the reproducibility of this can also be done by wiping any server and letting the CI processes restore a new one, complete with the appropriate access roles, firewall configuration, cluster membership, observation tools for developers to use etc. - everything also has backups and rollback strategies as needed - despite all of the procedures in place, delivering a new version can be as easy as clicking a button on a CI pipeline, the container images being delivered and becoming available on the clients' side in a few minutes for further processing Of course, that's still not good enough in my eyes and there are further improvements that can be done, as well as problems in place that prevent truly safe and easy development (e.g. no adherence to 12 Factor App principles due to historical reasons). If the projects used TDD and there was 90% test coverage, as well as ALL of the functionality was to be covered by integration tests (Selenium), then we could actually start talking about software engineering. Until then, i don't believe that the term is really apt, perhaps outside of the aerospace industry. Disclaimer: of course, the amount of effort that goes into designing a system and the workflows around it depends on a variety of factors, so not all systems need that sort of focus on quality. For most CRUD apps out there, you can probably just wing it, even though then you must be ready for the eventual outcomes and downtime that it will cause.
- deleted 5y ago[deleted]
- katepavlova 5y agoOh. Well why don’t you use your frustration into writing a message to the careless team member and firing him from the team? If you are working in a non professional team it doesn’t mean that others don’t learn from theirs mistakes. Also, how did end up in the company with such low quality team members?