20 ms·
Twitter has an internal root CA problem
- manv1 4y agoIt's pretty amazing that a formerly public company like Twitter had such shitty documentation/processes/infrastructure. I thought SOX mandated this sort of internal controls - after all, Twitter basically seems to be full of infrastructure risks that would (and have) negatively impacted them financially in a material way. No key access? Why didn't they print it out and stick it in a safe deposit box, which is what a couple of startups I've been with have done...along with a couple of other key pieces of paper. Physical backup.
- castillar76 4y agoThe really interesting part of this is what else is tied to that CA. If it’s just Puppet, it’s bad enough; internal PKIs have a habit of metastasizing into lots of other places, though, precisely because everything internal trusts them. Worst-case here is that some piece of the internals of the Twitter app relies on things from that CA—-for instance, it relies on packages to do app config changes or updates and the packages have to be signed from that chain or served from something with a cert from it. In that case they’d be hosed: you’d have to replace every copy of the Twitter app. Fairly unlikely, but wouldn’t be the first time I’ve seen it happen. Beyond that, though: Internal build systems? Data encryption? User client auth to critical services? Internal app mTLS for data exchanges? The list of possibilities goes on and on…
- walrus01 4y agovery possibly bullshit but huge if true: https://twitter.com/davidgerard/status/1634633886712954881 https://twitter.com/davidgerard/status/1634633886712954881
- nunez 4y agowouldn't put it past him since he only wanted "builders" and I bet he doesn't consider platform ops "builders" (even though they build tons of stuff; twitter's platform was basically a product in and of itself)
- DrScientist 4y agoIf this is true - who knows - then it reflects rather badly on the people who were fired - as they didn't implement safeguard for a 'run over by a bus' scenario when they were in charge.
- dragonwriter 4y agoThere’s “someone got run over by a bus outside of our control” and “The people in charge direct a bus to run over everyone covering a key function”. You don’t really plan for the latter scenario when you are in charge, instead, you just don’t direct a bus to do that. If your successor decides to do that, that’s…on them.
- zer0tonin 4y agoThere's "run over by a bus" and "90% of the company got ran over by a bus" scenarios. The second one is rarely worth implementing.
- walrus01 4y agoThere's also possibility of: "If this person hadn't been fired, they could use some other form of credentials within twitter's internal systems plus a passphrase they have memorized to login to the private-key-repository system where the credentials for the root CA are stored and retrieve them. But as they were fired abruptly they are not inclined to help Musk. And nobody has asked them".
- concordDance 4y agoArent abrupt firings the norm in the USA? My company had layoffs last year and the US people were gone the same day.
- dragonwriter 4y ago> Arent abrupt firings the norm in the USA? Abrupt firings of everyone with critical access, primaries and backups, is not, because its suicide. (Also why critical access roles are vetted carefully, because you want to make sure there is a lower-than-normal chance you will need to fire any of them, since that’s how you minimize the chance of a situation where you’d want to fire enough of them to cause a critical situation.) If you do need decide there’s a problem that requires you to fire those people, you find every way possible to delay firing some of them while you expand the set of people with that access (which may be only momentary, by compelling them to hand over credentials as part of the exit process, if you have confidence that you can do that successfully.)
- jtvjan 4y agoThe certificate for Twitter's hidden service expired a full week ago and they still haven't exchanged it.
- 8organicbits 4y agoI'll take the rumor with a grain of salt, but can anyone unpack what the recovery plan would be for something like this? It would obviously be a big problem, but where would you even start?
- jon-wood 4y agoAssuming they’ve still got access to the servers themselves via SSH, you’d start by issuing a new root CA cert for the Puppetmaster and putting that in place, then you’ve got to issue a new cert for every client and distributing those. It’s not impossible, but it’s also going to be a pain in the backside to do.
- justsomeadvice0 4y agoBeen there before, we did exactly this; except over OOB+reboot-into-single-user (because SELinux). Took us a few days (~5k servers) but managed to get out of it with no public-facing downtime. The other way would have just been to rekick the world one box at a time. A number of integration tests were added after that disaster :)
- threeseed 4y agoAccording to this [1] Twitter has 500,000+ servers spread across DCs, GCP and AWS. Which if we assume only a team of your size remains then it would take 300+ days. That would mean no OS patches etc which would put them firmly in the crosshairs of the FTC. [1] https://twitter.com/d_feldman/status/1562265193249390593 https://twitter.com/d_feldman/status/1562265193249390593
- justsomeadvice0 4y agoIf you are split-cloud under a homogenous puppet master without homogeneous break-glass SSH access (which would be crazy) then probably your best bet is to just re-kick the world. But the scaling factor for this sort of thing is most certainly not team size; it's "how many X servers can be down at the same time", which will increase with your number of servers. In any case I think the FTC is the least of twitter's concerns right now.
- donohoe 4y agoTaking it with a pinch of salt, but this stuff does happen. I've received calls from past employers, usually when they migrate a site I worked on to a new CMS or platform. There is some critical service (AWS, CDN credentials, domain related) etc. that no one knows who has access... Happily those appear to get resolved... but this... yikes (if true)
- ilyt 4y agoFunnily enough putting it in configuration management (like Puppet) can make it nice and automatic. But, well, if you fuck up your CM...
- plorg 4y agoIn a possibly more pedestrian example, my organization needed a re-mailer service set up and found out that the IT worker previously tasked with administration for that service had the MFA set up for his personal phone. I think they eventually got a hold of him to coordinate transfer of credentials, but knowing him, there was a 50% chance he could have left the company on bad terms and would have made things quite a bit more difficult.
- tedivm 4y agoI had something similar happen when I left a company, only I'm fairly consistent on deleting credentials to systems I'm not supposed to have access to. Fortunately it was for an internal service and nothing customer facing, so they were able to wipe and redeploy.
- Volundr 4y agoOne of the first things I do when leaving a company is remove all credentials from my password manager. Sure they should disable my accounts, but on the off chance they don't I still want it clear I don't have access. It doesn't have to be a departure on bad terms, if they needed my TOTP codes I can't help them. That secret is already gone.
- deleted 4y ago[deleted]
- ilyt 4y agoOh, I know that problem, we did change Puppet root CA due to mishap of one of the admins during updating to sha256 certs. But IIRC (it was long time ago) Puppet CA cert by default are issued for like 10-20 years, would be a bit weird if true. Also, old versions didn't had trust chain "just" root CA so puppet master would have to have key for that on disk anyway, proper "root CA + leaf CA for puppetmasters" have been a thing for just few years in Puppet. It would only be really problematic if they also lost SSH access to those machines using Puppet. If you have root access the fix is not exactly hard. But then they fired people that did had access so that might also be a problem We made sure all of our machines can be accesses both by Puppet and by SSH kinda for that reason; we had both accidents of someone fucking up Puppet, and someone fucking up SSH config rendering machines un-loggable (the lessons were learned and etched in stone). So really, depending on who has access to what, it can be anything from "just pipe list of hosts to few ssh commands fixing it" to "get access manually to the server and change stuff, or redeploy machine from scratch". Again, assuming muski boy didn't fire wrong people
- brazzy 4y ago> It would only be really problematic if they also lost SSH access to those machines using Puppet. If you have root access the fix is not exactly hard. > But then they fired people that did had access so that might also be a problem Oh my, wouldn't that be delicious... Gotta wonder how you'd go about fixing that, though. Assuming that those people's access was also tied to their employment and irrevocably voided when they were fired: I guess it would depend on how well those machines are secured against attackers with access to the hardware.
- gorjusborg 4y agohttps://archive.is/vCoD6 https://archive.is/vCoD6
- hayst4ck 4y agoI think people who work in reliability see this type of thing as the real existential threat to twitter. It's unrealistic that a large infrastructure would fall over overnight, but what is very realistic is small problems being neglected until they become big problems, or multiple problems happening at the same time. This alone is probably manageable, it might even be simple but painful to handle for 2-15 of twitters employees (pre-firing) with specialized knowledge. If 3 people knew the disaster recovery plan and they all got fired because they were so busy maintaining things and fighting fires that they failed to get good reviews by building things, well I wouldn't be surprised. Likewise the employees trusted with extreme disaster recovery mechanisms are not the poor souls on H1Bs who don't have the option of leaving easily, so the people trusted with access might have already jumped ship since they aren't being coerced into staying on board with a mad man. The real existential threat is another problem compounding on top of this or a disastrous recovery effort. Auto-remediation systems could do something awful. A master database could fall over and a replica be promoted, but if that happens twice, 4 times? Without puppet to configure replacement machines appropriately, there could be a very real problem very quickly. Similarly, extremely powerful tools, like a root ssh key, might be taken out, but those keys do not have seat-belts and one command typed wrong could be catastrophic. Sometimes bigger disasters are made trying to fix smaller ones. Puppet can be in the critical path of both recovery (via config change) and capacity.
- karmakaze 4y agoWhenever I hear specifics about likely ways things could fail, I always see a plan. "Hey this all makes sense, lets focus on having these areas covered before they come to pass." Same goes when someone lists all the reasons why a proposal isn't viable. "Great, so we'll address those and be golden then?" Often they list them as fact without considering (or the ability to imagine) that they could be made viable with additional effort.
- ep103 4y agoThat's okay, Musk tweeted that Twitter needs a complete, green-field rewrite 5 days ago, I'm sure that will solve the problem.
- yuppie_scum 4y agoThat Mastodon server has a load time problem.. took a solid 30 seconds for me to load
- dpkirchner 4y agoWonder if their servers all share a common NTP server/pool (that they control).
- not_enoch_wise 4y ago[flagged]
- yeahsure22 4y agoThe good part of your insightful analysis is you can keep rolling the dates so you never have to revise it.
- dx034 4y agoSo far, there seem to be surprisingly few issues. Some glitches here and there, but overall stability looks still quite good. I would've expected major issues much sooner, especially as they did push out new features in the meantime.
- giraffe_lady 4y agoDo you use it actively? There have been minor problems on a near daily basis and moderate ones a handful of times. It was only ever average stable before but it was at least always that, it's far less consistent recently.
- lambo4bkfast 4y agoDo you have data on this? Its not like other apps don't have issues. Sometimes I open netflix and it takes 30 seconds to show my profiles; that doesn't mean the app is garbage.
- giraffe_lady 4y agoDo I have data on a website I use sometimes? No.
- dx034 4y agoI use it actively and feel like recommendations have improved, at least for me. Ads are a bit more annoying but I've faced zero problems with stability. But I'm only reading, not actively tweeting.
- parasense 4y ago> Musk fired everyone with access to the private key to their internal root CA, The way forward is to generate a new CA root certificate. > and they can no longer run puppet because the puppet master's CA cert expired They can reconfigure internal tools to use the new CA root certificate, or rather one of the signed intermediate certificates. > and they can't get a new one because no one has access. They can simply generate new CA root certificates, and sign or create new intermediate certificates. > They no longer can mint certs. Yes, they, can... > My limited understanding in this area is that this is...very bad No, it, is, not... There are two immediate issues that come to mind. * Twitter was so awful before, that it relied on people to safeguard the keys to the kingdom. This is very bad practice, and one of the many things Musk will no doubt be fixing. For any mission critical assets, and especially certificates, but also passwords... current modern day corporate practice is to have a secure ledger of these that can be accessed by the board of directors, the executive managers, and designated maintainers. At no point ever should the password be entrusted to anybody, but rather a "role" that functions as the one who has access. Say for example, the CIO/CTO and their subordinates. * The Second issue is the one everyone is fixating upon, and that's firing important people who put the company at risk. This is a big issue, and certainly Musk could have done a better job of scoping out who represents a single-point-of- failure at twitter, eliminate that risk, and then proceed with the culling. In a modern enterprise no single person should be capable of putting the entire operation at risk. It's just that simple. So in a way, Musk accelerated what was probably inevitable at Twitter already. They were probably precariously close to destruction already, and now they can learn the hard way of not repeating these mistakes.
- ryan_lane 4y ago> For any mission critical assets, and especially certificates, but also passwords... current modern day corporate practice is to have a secure ledger of these that can be accessed by the board of directors, the executive managers, and designated maintainers. At no point ever should the password be entrusted to anybody, but rather a "role" that functions as the one who has access. Say for example, the CIO/CTO and their subordinates. Maybe in hacker movies. In real life, you try your best to avoid anyone having access to keys or passwords, and rely on HSMs, cloud KMS, secret services, etc. Access to those things is controlled by your security team, with multi-factor authentication, often stored in safes, with alerts being fired when they are used (because they should never be used). The audit logs that trigger these alerts should be written in WORM storage, so you can track access back down to individuals, and so that you know when you need to rotate secrets accessed by humans. Ideally your CA infrastructure automatically rotates and distributes. There's absolutely no way in hell you should allow your board to have access to these things. Most companies slowly work their way towards full automation, and until that happens, your security team usually owns manual rotations of critical systems like this. Only a fucking moron would fire all of these people.
- mariusor 4y agoSince about 3-4 days I had issues with using tweetdeck, which doesn't load in Firefox with a key pinning failure. I'm not sure it's related, but it seems like to large of a coincidence not to.