4 ms·
This is news because later this year they are probably going to have problems. You have obviously never worked at a large internet service, but when you spend s
by evgen 4y ago
This is news because later this year they are probably going to have problems. You have obviously never worked at a large internet service, but when you spend some time backstage at such a venture you learn that there is a great deal of inertia within various projects and pieces of infrastructure. If 80-90% of the team for component X departs the system is unlikely to fail quickly (a frequently repeated mistaken belief when Twitter was shedding engineers so fast) but it is going to start having problems that will not be visible to the outside world. This is not a problem for a while because the former employees who built and maintained the system were fairly smart people and built something that was relatively reliable; they wanted to be able to sleep through the night without a page or go on vacation too.
Bugs exposed by other teams using that component will go unfixed and will need to be patched around by the caller. New features that might have used component X will be sidelined because no one wants to put more load on it or risk their feature on a component that people have started whispering about in the corridors. What is worse is that the people frantically spending their days patching and repairing this Titanic below decks do not see the iceberg and do not have enough manpower to turn the ship even if they did, the people who could help give warning or helped prevent the problem left a while ago. Eventually this component is going to get stressed, some lingering bug will get triggered, or some other system elsewhere in the architecture diagram will fail and shed load to our component X that is barely keeping up and it too will fail. The danger for Twitter is that the failure could cascade through multiple poorly-maintained systems and lead to a much bigger problem.
- gsatic 4y agoYou should visit big telcos. Things chug along not because there is a great competent army fixing anything. Institutional knowledge is regularly lost thanks to changes in tech, ever changing fortunes of the biz, divisions getting sold off etc etc. They chug along because people make do with what ever works. It sort of like your OS crashing. You keep backups, reload, reboot. You dont debug and fix. The older systems get that's pretty much the norm.