6 ms·
Why would Twitter be constantly in danger of crashing without human intervention? Isn't it just a bunch of software that just runs? If not, why was it not prope
by rdiddly 4y ago
Why would Twitter be constantly in danger of crashing without human intervention? Isn't it just a bunch of software that just runs? If not, why was it not properly engineered, and who needs to be fired for that?
Musk is embarrassing as always but he's right about one thing - engineers are costly, and for years now, the problem with Twitter as a business has been the huge outlay for engineering talent. And for what? No effect I can discern; the platform is clearly not a force for good in the world, first of all, but even looking at it strictly as a product, the product in fact gets worse every year and they're constantly discontinuing functionalities. So the whole "what has your code accomplished" show, it's inept but I know where it's coming from.
- david422 4y ago> Why would Twitter be constantly in danger of crashing without human intervention? Kindof true. Platforms need a ton of maintenance - but usually the fires needing to be put out are because of other human changes. Stuff like disk space, certificate expiration etc etc will bring down a platform without maintenance though.
- acdha 4y agoTwitter is a complex application which serves a ridiculous amount of traffic to people around the world. That kind of system is constantly changing in response to user activity - it’s done in the same way that a garden is done. Musk just laid off the landscapers so what’s about to happen isn’t immediate but is inevitable.
- dzikimarian 4y agoYou need to: * Apply security patches - there's thousands of dependencies. * Manage hardware/cloud resources according to volume of data you need to handle. * React to the changes in operating systems/browsers you run on. * Fix bugs - they are there, because simply there's not enough time, money, and need to write "perfect" software. Competition will not do that and will beat you to the market. There's no internet connected software, that can "just run", because world around is changing and you must catch up. Fells a bit weird to explain that on HN :-|
- site-packages1 4y agoThis is all needed to continue to iterate on software. It’s not as needed for the ossified state Twitter has been in for many years, as the parent pointed out. Acknowledge there is _some_ work needed, maybe refreshing certs (though software systems I have built have always been set up to do that automatically), apply security patches, keep an eye on dashboards. But this is a job for a skeleton team, not 7500 engineers or whatever. It’s honestly kind of sad it took that many people to basically “keep Twitter up” all these years.
- dzikimarian 4y agoI believe that 7500 is total headcount, not just engineers, but yeah software development can be terribly inefficient, when all problems are solved by throwing money at them, like VC funded startups like to do. We'll see what happens now, when there seems to be a drought.
- jrumbut 4y agoThere are also integrations with partners, new regulatory requirements, changes to support ownership's new priorities (and to undo the old priorities). I don't work at Twitter but these things happen at every business.
- mihaic 4y agoWhat your saying is true for almost every web app. I don't see any reason to justify Twitter's complexity that requires 1000 engineers. Scale is such a weak argument since their entire product paralelizez very easily.
- Xorlev 4y agoI don't think you understand the fanout/fanin problems their product has. Twitter's raison d'être is allowing folks with millions of follows to see their posts in near real time, then monetize it with ads. That means analytics, billing, and more. From pretty basic requirements you can create a huge amount of work, especially when you're managing your own infrastructure.
- sbergot 4y agoAt scale everything is more complex. The hardware that runs the software can fail. Ddos happen. Network issues. A building catches fire. Etc etc. Sometimes you have automatic recovery. Sometimes you don't.
- kweingar 4y ago> Why would Twitter be constantly in danger of crashing without human intervention? Isn't it just a bunch of software that just runs? If not, why was it not properly engineered, and who needs to be fired for that? I would encourage you to read about “site reliability engineers” and the work that they do, in case you haven’t heard of them.
- Barrin92 4y ago>Isn't it just a bunch of software that just runs? the fact aside that bug-free "just a bunch of software" of the size that serves half a billion people has yet to be invented. All software has the nasty habit of running on hardware and real infrastructure and that at the very least needs some serious maintenance. If your standard for proper engineering is that it runs without human intervention at all I can't think of any complex system in the world that fits that description.
- edgyquant 4y agoTwitter is hosted on aws they don’t maintain their own hardware. This is a cope because Elon is clearly right. An application where prod begins to get buggy just because it’s left alone is poorly engineered. There should be CI/CD and prod should be extremely stable with the exception of non-trivial bugs that should not just start popping up by the barrel because a couple of weeks without new pushes.
- detaro 4y agoTwitter is not "hosted on AWS". They have started to use various cloud providers (AWS and GCP, at least at various times), in combination with their own hardware.
- viraptor 4y agoEven running on AWS is not completely isolating you from hardware related changes. For example AWS is retiring both old instance classes for databases and Aurora 1 soon. This is a multi week migration project with a hard deadline for me and it's not even a huge company/service. For Twitter I suspect that kind of deprecation would take month(s) to fully roll out.
- Dunedan 4y agoThat's surprising to me, given how much effort AWS put into emulating old instances types using newer instances to still be able to offer instances with the characteristics of old instance types to customers. [1] Do you have additional information about that you can share (maybe an announcement regarding this made by AWS)? [1]: https://perspectives.mvdirona.com/2021/11/xen-on-nitro-aws-nitro-for-legacy-instances/ https://perspectives.mvdirona.com/2021/11/xen-on-nitro-aws-n...
- analog31 4y agoAll sufficiently complex control systems run in failure mode (partial or full manual override) approaching 100% of the time. For instance this is why there needs to be a multi tiered (culminating with HN front page exposure) process for fixing automated account cancellations and app store rejections.
- meaydinli 4y agoYou are making the same mistake Elon is making. You don't know enough about the subject to even be aware of how much you don't know and what you are missing. You might be very competent in another subject, but that doesn't grant you the right to write in such a demeaning tone.
- rdiddly 4y agoI don't need to be granted that right, I was born with it, just like you with your far more demeaning and ignorant comment about things (me) you know nothing about.
- jimjimjim 4y agoSorry, you aren’t born with any rights. Those that you may appear to have were granted to you by a society. Never forget that
- briandon 4y agoYour view on individual rights is far from universal. Never forget that. "We hold these truths to be self-evident, that all men are created equal, that they are endowed by their Creator with certain unalienable Rights, that among these are Life, Liberty and the pursuit of Happiness." https://www.archives.gov/founding-docs/declaration-transcript https://www.archives.gov/founding-docs/declaration-transcrip... Or, alternatively: "All human beings are born free and equal in dignity and rights." (from Article 1) https://www.un.org/en/about-us/universal-declaration-of-human-rights https://www.un.org/en/about-us/universal-declaration-of-huma...
- deleted 4y ago[deleted]
- jimjimjim 4y agoAnd if society collapses and whatever hellscape that replaces it withholds those “rights” from you?
- mmanfrin 4y agoA little shocking to see a comment saying 'why does it take work to keep a website running' at the top of a comment section here.
- vsareto 4y agoI think people are just surprised that it takes so much work because we thought things would get easier as time went on. It's actually the opposite though as more technologies have been developed. There's way more things involved in the stack and all of those can interact and break in their own interesting ways. There is also just frankly more traffic to serve because of mobile devices, so more things break at scale.
- acdha 4y agoI’d also note that expectations have increased, too. Consider for example how much people expect to just work when they type text into a search box, when the corpus is gigantic, in many languages, and constantly updating.
- nilsbunger 4y agoTechnical and support teams are needed to keep a site like this running globally. Good examples of the work from an SRE: https://twitter.com/MosquitoCapital/status/1593541177965678592 https://twitter.com/MosquitoCapital/status/15935411779656785... Other things which aren't necessarily needed to technically keep the code running, but without which you won't have an operating service: * Sales and account management teams to keep and grow advertiser revenue, and all the support services around them (marketing, integration, etc) * Moderation teams to keep illegal or dangerous content off the platform, and implement your moderation policy which you need for many reasons including appeal to advertisers * Abuse teams to respond to the latest DoS / hacking / impersonation / harrassment / etc. * Legal and ops teams to understand and implement compliance with laws, regulatory filings, etc across 150+ countries, and in some countries like the US, individual state variations. You tell me how many people you can do this with at the scale of Twitter. I think > 1000 people are a skeleton staff just to keep the lights on.
- csours 4y ago> Why would Twitter be constantly in danger of crashing without human intervention? Isn't it just a bunch of software that just runs? If not, why was it not properly engineered, and who needs to be fired for that? I can write software that runs today. I don't know about tomorrow. Disks crash, other services crash, networks go down. No matter how competent your Byzantine Generals are, sometimes you have to recover data that was in-flight during a crash because disks and networks don't just crash, they slow down and they lose packets and sectors. Add on security patches and changes in architecture (I'm looking at some Java 1.3 code that runs INSIDE an Oracle DB right now, ask me how I feel about not updating your architecture). You may say that architecture doesn't change that fast. If the business is successful, it will happen eventually.
- ineedasername 4y agoThis may come across snarky, but it’s unintended. Instead I’m trying to set out a sort of mental framework for answering the question from (slightly less than) first principles. That’s a bit like asking “why does any large internet service need infrastructure maintenance?” If you can’t immediately see the answer in twitter’s case, then you either haven’t yet applied your knowledge of how large cloud services function and need maintenance to keep from cascading failures, or you don’t have that knowledge in which case the twitter case can’t be answered without explaining first the general case of how any large software system needs regular oversight. But then, I think that even if you can’t think of an answer for Twitter, you can probably figure out a little of the answer for the general question. You can start with asking yourself another question: why would Twitter (or any very large software project) pay $Millions for maintenance that isn’t required? Is there collective delusion on the topic? Or has twitter hit upon the absolute perfectly engineered platform ever? If the answer to either of those is “no” then a twitter system failure is only a matter of time. How much time is a matter of some debate. The first time something hits a quota or something of that sort that requires a human input to ensure automatic allocations don’t accidentally get insane could be a pebble. Or maybe the first pebble is a minor unpatched security flaw, but twitter is a very sweet high profile target… whatever the first pebble is, it makes a slightly worse issue or multiple additional pebbles much easier, and you enter cascading failure. They generally happen slowly relative to the proportion of a total collapse, and then very very quickly. So another question to ask is how much of twitter’s infra teams were pure bloat, how much were excess capacity necessary to cover average amounts of employee vacation/sick/leave that goes on at any time, etc, and how elastic the remaining staff can be (putting in extra hours, working in areas that are secondary skills because no one else is available) and for how long they’re willing to do it before burnout or better job offers with less stress and overtime, whether staff losses have already hit a critical mass to make this inevitable… But I think this gives a general sense of why it can’t tick over forever on its own, or at least some inroads into thinking about the answer.