17 ms·
Results of technical investigations for Storm-0558 key acquisition
- ratg13 3y agoI feel like there is a lot missing from this writeup, but I can't put my finger on exactly what. Also it feels strange that Government doesn't have its own signing key and they just use the same as everyone else. Which they didn't address and apparently do not intend to change.
- sidewndr46 3y agoif the government had its own key, you could trace anything they signed. Governments likely want code and other stuff they sign to appear as if another actor signed it
- ratg13 3y agoThe key belongs to Microsoft. Microsoft is the one signing the auth tokens, not the end users. I'm saying that Microsoft should have a separate private key to sign government auth tokens with.
- shadowgovt 3y agoIIUC in general they do. One of the steps of this failure is that a key that had no business signing off on accessing government data was granted that scope by MS's cloud software because they changed the scope-checking API in such a way that their own developers didn't catch the change ("Developers in the mail system incorrectly assumed libraries performed complete validation and did not add the required issuer/scope validation"). So instead of failing safe, lack of new code to address additional scope features "failed open" and granted access to keys that didn't actually have the right scope.
- mrguyorama 3y agoHow banal can a software mistake be before we aren't allowed to besmirch the name of the devs involved? Is forgetting a test case a shameable offense? What about ignoring authentication? Rolling your own? Turns out when you write APIs that access security related things, you have to treat everything coming in as a threat, right? Shouldn't that be table stakes by now? We need a professional gatekeeping organization because the vast majority of us suck at our jobs and refuse to do anything about it.
- deleted 3y ago[deleted]
- phillipcarter 3y agoThe issue was caused by a race condition in extremely complicated software. Good luck setting up a gatekeeping organization that can track that level of detail (and understand every dimension of a possible fault like this).
- lcnPylGDnU4H9OF 3y agoI wouldn't expect the gatekeeper to track these issues but rather to sign the credentials that developers have. Then the individual developers (ostensibly) have a base level of training that set them up to more likely avoid these issues.
- phillipcarter 3y agoI don't know if you appreciate the level of complexity here. We're talking about a core diagnostics system (extremely complex software), that already has guards in place to protect against this stuff (complex again), but there was a race condition (complex again), and this one instance in likely billions and billions of transactions is what led to an issue. What training do you think could have prevented this? Microsoft deals with complex software at enormous scale. Bugs happen. And in this case, it was a severe one that's already been dealt with.
- ano-ther 3y ago
- The28thDuck 3y agoDoesn’t seem like an operational issue, seems more like 99.9% design coverage of preventing these issues from arising. I will say, I’m not very satisfied with just “improved security tooling.” My gut is telling me that there is a better solution out there to guarding against credential leakage, but I feel wrangling memory dumps to have “expected” data is a fools errand.
- gordian-not 3y agoWeird they don’t have logs saved for something that happened two years ago due to ‘retention policies’. That’s something I would fix
- paxys 3y agoIt works this way by design. Most companies will retain logs for exactly as much time as legally required (and/or operationally necessary), then purge them so they don't show up in discovery for some lawsuit years down the line.
- eli 3y agoIt's also a GDPR requirement to minimize the collection of personal data and to purge it as soon as it is no longer needed.
- wglb 3y agoThere is a way to keep arbitrarily large logs and be fully compliant with GDPR with a little engineering.
- gordian-not 3y agosounds interesting, can you elaborate a bit?
- wglb 3y agoFor each piece of PI/PII data, generate a mapping in a table of that piece to a secure random number, and store the generated random number in place of the personal data, and use that in the log. Then, if deletion is required, simply erase the row that holds the mapping. And finally, be sure to not store that mapping table in the same place as your backups or your logs.
- eli 3y agoIn a way that lets you go back and identify behavior of an individual person? I doubt that.
- munificent 3y agoI feel this everytime one of these articles comes out, but it seems totally bizarre to me that we rely on private enterprises to deal with state-level attacks simply because they are digital and not physical. If a Chinese fighter jet shot down a FedEx plane flying over the Pacific, that would be considered an attack on US sovereignty and the government would respond appropriately. Certainly we wouldn't expect FedEx to have to own their own private fleet of fighter jets to protect their transport planes. No one would be like, "Well it's FedEx's fault for not having the right anti-aircraft defenses." But somehow, once it hits the digital domain we're just supposed to accept that Microsoft is required to defend themselves against China and Russia.
- eli 3y agoIsn't that the idea behind CISA?
- tptacek 3y agoI don't think so? What makes you think it is?
- dboreham 3y agoNSA
- c0pium 3y agoThey’re explicitly forbidden from doing things like that. As they should be; do you really want the government to have access to the kind of private corporation data they would need in order to defend them?
- c0pium 3y agoVery much no. CISA defends the federal executive branch and advises critical infrastructure. They don’t and shouldn’t have a proactive role in defending private companies.
- tschwimmer 3y agoThe digital domain is fundamentally lower stakes and harder to protect than the physical one. It is good that we do not respond to cyber attacks like we do physical ones because we would have escalated to nuclear war over a decade ago. The scope and volume of cyberattacks is very high but my understanding is that the US has a correspondingly high volume of outbound attacks as well.
- cobertos 3y agoSome things were not plainly spelled out: * July 11 2023 this was caught, April 2021 it was suspected to have happened. So, 2+ years they had this credential, and 2 months from detection until disclosure. * How many tokens were forged, how much did they access? I'm assuming bad if they didn't disclose. * No timetable from once detected to fix implemented. Just "this issue has been corrected". Hope they implemented that quickly... * They've fixed 4 direct problems, but obviously there's some systemic issues. What are they doing about those?
- ranting-moth 3y agoAdverse inference is very valid. https://en.m.wikipedia.org/wiki/Adverse_inference https://en.m.wikipedia.org/wiki/Adverse_inference
- jeremyjh 3y agoIts valid in a civil court where discovery processes exist. It doesn't really apply to public relations, where information could be withheld for a number of unknowable reasons. Of course everyone is free to speculate, but its not supported by a link to theories of common law in civil torts.
- tadzikpk 3y agoIf this credential is still valid 2 years later, what is their credential rotation policy?
- natas 3y agoI would fire their entire security team on the spot.
- shadowgovt 3y agoRisky. Those are the only people on the planet you can trust to never make this mistake again.
- bagels 3y agoIs it possible to build a security team in a way that you can guarantee to never have any vulnerability ever?
- nazgulsenpai 3y agohttps://github.com/kelseyhightower/nocode https://github.com/kelseyhightower/nocode
- sfink 3y agoYeah, why do companies even hire security people who allow security problems to exist? Or coders who write bugs?
- tptacek 3y agoIt feels like there are some missing dots and connections here: I see how a concurrency or memory safety bug can accidentally exfil a private key into a debugging artifact, easily, but presumably the attacker here had to know about the crash, and the layout of the crash dump, and also have been ready and waiting in Microsoft's corporate network? Those seem like big questions. "Assume breach" is a good network defense strategy, but you don't literally just accept the notion that you're breached.
- shadowgovt 3y ago> but presumably the attacker here had to know about the crash, and the layout of the crash dump If I were an advanced persistent threat attacker working for China who had compromised Microsoft's internal network via employee credentials (and I'm not), the first thing I'd do is figure out where they keep the crash logs and quietly exfil them, alongside the debugging symbols. Often, these are not stored securely enough relative to their actual value. Having spent some time at a FAANG, every single new hire, with the exception of those who have worked in finance or corporate regulation, assumes you can just glue crash data onto the bugtracker (that's what bugtrackers are for, tracking bugs, which includes reproducing them, right?). You have to detrain them of that and you have to have a vault for things like crashdumps that is so easy to use that people don't get lazy and start circumventing your protections because their job is to fix bugs and you've made their job harder. With a compromised engineer's account, we can assume the attacker at least has access to the bugtracker and probably the ability to acquire or generate debug symbols for a binary. All that's left then is to wait for one engineer to get sloppy and paste a crashdump as an attachment on a bug, then slurp it before someone notices and deletes it (assuming they do; even at my big scary "We really care about user privacy" corp, individual engineers were loathe to make a bug harder to understand by stripping crashlogs off of it unless someone in security came in and whipped them. Proper internal opsec can really slow down development here).
- Eduard 3y ago>... you have to have a vault for things like crashdumps that is so easy to use that people don't get lazy... Let's assume a crash dump can be megabytes up to gigabytes big. How could a vault handle this securely? the moment it is copied from the vault to the developer's computer, you introduce data remanence (undelete from file system). keeping such coredump purely in RAM makes it accessible on a compromised developer machine (GNU Debugger), and if the developer machine crashes, its coredump contains/wraps the sensitive coredump. A vault that doesn't allow direct/full coredump download, but allows queries (think "SQL queries against a vault REST API") could still be queried for e.g. "select * from coredump where string like '%secret_key%'". So without more insight, a coredump vault sounds like security theater which tremendously makes it more difficult for intended purposes.
- _tk_ 3y agoI am very curious why Microsoft is insisting that the key itself was „acquired“ without having anything to show for it. The wording seems a little odd to me, the constant repetition even more so.
- ynniv 3y agoAnd that's why you should keep your key material in an HSM, kids
- baz00 3y agoSo if we remove the careful wording, someone downloaded a minidump onto a dev workstation from production and then it was probably left rotting in corporate OneDrive until that developer's account was compromised. Someone took the dump, found a key in it and hit the jackpot.
- trifurcate 3y agoAnd, crucial to this exploit actually working to the extent it did, Microsoft's own developers failed to implement a secure authentication check on top of their own libraries and infrastructure.
- runeks 3y agoHow so?
- baz00 3y agoLeaving credentials and keys in memory.
- drodgers 3y agoAlso completely failing to check the scope of the request before validating it! > Microsoft provided an API to help validate the signatures cryptographically but did not update these libraries to perform this scope validation automatically
- remram 3y agoAnd there was a redaction system, which did not redact the key ("race condition"). Then a detection system, which didn't detect the key. And then the key was used to access an entirely different system with an entirely different access level and it just worked anyway. The phrasing as "some obscure bugs were carefully exploited" seems a bit off, it looks more like a comedy of errors where none of the security systems served its purpose at all.
- 1970-01-01 3y ago>Our investigation found that a consumer signing system crash in April of 2021 resulted in a snapshot of the crashed process (“crash dump”). The crash dumps, which redact sensitive information, should not include the signing key. In this case, a race condition allowed the key to be present in the crash dump (this issue has been corrected). Correction is good, but why can't they go one more step and allow everyone to scan their server minidumps for crash-landed keys?
- dgudkov 3y agoA breach like that requires a very good understanding of Microsoft's internal infrastructure. It's safe to assume that the breach was a coordinated effort of a team of hackers. This is not a cheap effort, but the payback is enormous. Hyper-centralization leads to a situation when hackers concentrate their efforts on a few high-value targets because once they are successful, the catch is enormous. I'm pretty much sure that there are teams of (state-sponsored) hackers that are already doing deep research and analysis of the internal infrastructure of Google, Microsoft, Amazon, etc. The breach gives an idea of how well already the hackers understand it. I would argue, it's time to decentralize inside a wider security perimeter.
- splitstud 3y ago[dead]
- deleted 3y ago[deleted]
- sargun 3y agoYou have to assume that you have nation state actors working at your organization at sufficient size. Unfortunately, it’s difficult to work around this assumption, because anyone can be compromised at any time.
- transcriptase 3y agoI find it somewhat amusing that companies like Microsoft and Google that have pivoted a large portion of their business model to collecting, keylogging, recording, scanning, exfiltrating, telemetrizing, collating, inferring, and analyzing every last iota of data they can about as many people as possible under the guise of improving their products or personalizing ads... ... can't identify nation state actors within their own company. I suppose that would be illegal. Whereas using it to improve AdSense CTR or selling it to brokers is perfectly acceptable.
- bostik 3y agoIf you are up against an adversary with an unlimited budget and organisational event horizon measured in years, your quarter-to-quarter thinking will always kneecap you.
- rdtsc 3y agoWonder if the actor caused the crash of the system in the first place? Or it was crashing so often they didn’t have to. Race condition to scrub the crashdump sounds fishy. When the system is crashing it’s hard to make assumptions or have any guarantees any cleanup and scrubbing is going to happen.
- natch 3y ago> Due to log retention policies, we don’t have logs with specific evidence No “this issue has been corrected” for this one. Are we still budgeting storage like it’s the 1990s for logs?
- cesarb 3y ago> > Due to log retention policies, [...] > Are we still budgeting storage like it’s the 1990s for logs? Retention policies are not necessarily about storage space; sometimes, they are there to avoid being required to provide that old data during lawsuits.
- c0pium 3y agoRetention policies at cloud providers are 100% about storage space (and accompanying cost). At companies like Microsoft saying “reduced cogs” is a very reliable way to get bonuses.
- lazyasciiart 3y agoIf storage space was 100% of the decision about logs they would not retain logs. The other factors include technical utility and legal considerations.
- c0pium 3y agoHuh? Logs are how you run the service. Cost is what keeps you from retaining more of them. Since you seem not to be familiar with the subject, retention policy is something of a misnomer; it would be more accurate to call it a deletion/destruction policy. The default is retain everything forever. There are no legal considerations here. That’s what lawyers are for, and big tech has a ton of lawyers. If you work somewhere that people are deleting things to keep them from being discoverable, run away as fast as you can.
- lazyasciiart 3y ago
- mh8h 3y agoWhy don't they use HSMs instead? The whole point of those hardwares is to prevent leaking the key materials.
- olliej 3y agoThey way they discribed I assumed that the crashlog they got was from an HSM?
- agrajag 3y agoIt would be a gross failure of an HSM to allow private key material to leak in any way.
- SgtBastard 3y agoAs a sibling commenter mentioned - if a HSM dumps its memory where it contains private key material, that’s a spectacularly bad HSM, which MS wouldn’t have been able to fix the race condition of. Reading that MS were able to fix the crashing system’s race condition that included the key, it’s likely to have been a long-lived intermediate key for which the private key was held in memory (with a HSM backed root key for chain of trust validation, assuming MS aren’t completely stupid). The challenge is the sheer scale these servers operate in terms of crypto-OPS… it would melt most dedicated HSMs.
- toast0 3y agoThese guys [1] claim to have "the fastest payment HSM in the world, capable of processing over 20,000 transactions per second." I imagine the peak load for signing authtokens for Microsoft accounts is way higher than that. [1] https://www.futurex.com/download/excrypt-ssp-enterprise-v-2-overview/ https://www.futurex.com/download/excrypt-ssp-enterprise-v-2-...
- timmclean 3y agoIs there a reason why they couldn't split the load across multiple HSM? For something so sensitive I would've expected a design where one or more root/master keys (held in HSM) are periodically used to sign certificates for temporary keys (which are also held in HSM). The HSMs with the temporary keys would handle the production traffic. As long as the verification process can validate a certificate chain, then this design should allow them to scale to as many HSMs as are needed to handle the load...
- bananapub 3y agoas far as I can tell, the only non-bug mistake here was allowing coredumps to leave production ever. if this is your attacker, you are pretty fucked no matter how good you are.
- Eduard 3y ago> The key material’s presence in the crash dump was not detected by our systems (this issue has been corrected). Now hackers have it even easier to find valuable keys from otherwise opaque core dumps: Microsoft's corrected detection software will tell them as soon as it finds one.
- Arrath 3y agoWhile true that it is easier for malicious actors to find this kind of thing with a tool that goes DONG! after a quick scan, its not as if the previous security through obscurity of "key hidden in megs or gigs of crashdump" was much of a blocker for a suitably motivated adversary.
- time4tea 3y agoWhat this means is that the keys are not stored in non-recoverable hardware, they are available to a regular server process, just some compiled code, running in an elevated-priv environment. There is no mention that the systems that had access to this key were in any other than the normal production environment, so we may extrapolate that any production machine could get access to it and therefore anyone with access to that environment could potentially exfil the key material.
- planetjones 3y agoNone of the reports mention if two stage authentication or any other extra factor authentication that enterprise accounts would be secured with were bypassed too. Am I right to assume that because the attacker had the signing key all of the extra authentication mechanisms that would have been enabled on accounts were bypassed by the attacker (because the attacker could create a token that bypassed all the extra authentication methods)? And I presume there has been no known dump of e-mails exfiltrated during this attack?
- ShadowRegent 3y ago> Am I right to assume that because the attacker had the signing key all of the extra authentication mechanisms that would have been enabled on accounts were bypassed by the attacker...? That's my understanding.
- SgtBastard 3y agoBecause it was a signing key that was stolen, the attackers could move straight to the post-authentication phase and forge authorization tokens. Those email accounts could have had multiple authentication factors enabled, other conditional access policies applied (geo-location, device trust, time of day etc)… all of which were skipped over.
- bobalob_wtf 3y agoWith the signing key they could mint the same type of token you get once you pass all of the authentication steps.
- 65a 3y ago1. Use HSMs, especially for big important keys with longevity 2. Minimize workflows that cross security boundaries Are my takeaways
- drodgers 3y agoThis is not confidence inspiring. These fixes are only surface-level, rather than looking at the underlying systemic failures: - Not storing key data in an HSM. - Exporting crashdumps outside of the production account. Redacting private keys and other sensitive data from these will always be failure-prone. Keys are also only just one problem, there will be personal data, passwords etc. in those dumps depending on what process it was. - Corp environment infiltration not detected at the time (presumably, since this part is pure guesswork) - Not enough log retention in the corp environment to track a 2 year old infiltration. - Not assuring that key validation correctly denied requests with the wrong scope. Fails-with-a-valid-key-with-wrong-(scope/date/subject etc.) are the kind of cases that always deserve test coverage, especially for a dedicated key-validation endpoint. Also concerning that this wasn't found by manual-testing/red-teaming/pen-testing in the time since it shipped. - Slow detection, slow response, poor communication
- deleted 3y ago[deleted]
- nijave 3y ago- Lack of timely key rotation The lack of scope checking seems especially egregious. It sounds like any number of keys would have been incorrectly trusted. If it was an RSA key that signed JWTs which it sounds like or similar, Microsoft has an issuer endpoint for all customers and it's critical to check the issuer/scope for those since any number of things can create a token with a valid signature.
- yyyk 3y ago>crashdumps.. Redacting private keys and other sensitive data from these will always be failure-prone >Not enough log retention in the corp environment to track a 2 year old infiltration. The two are in conflict. Redaction is a problem for logs as well. A longer retention period implies a more damaging vulnerability.
- drodgers 3y agoI was more talking about the problems of moving data across trust zones rather than retaining it at all. Retaining logs and crashdumps for several years is good, but moving them from a locked-down production environment to a less secure corp account (where they were presumably easier to work with because of the lower security requirements) is why this leak happened.
- noodlesUK 3y agoI feel like an issue that really got them was that the keys weren’t rotated. It sounds like quite some time passed between when the key was moved where it didn’t belong and when it got snatched. If keys were rotated frequently, it would not have been possible to use it to forge a token.
- ranting-moth 3y agoMicrosoft investigating a ginormous breach at Microsoft. Common guys. This requires at least one 3rd party investigation to be credible.
- est 3y agoThen a bunch of Chinese guys show up as the 3rd party.
- aftbit 3y agoI do not fully understand the nature of the credential used here. Why was it still accepted several years after it was generated?
- AtNightWeCode 3y agoIs the key a certificate? Then it might be installed on all Service Fabric clusters. There is overall something fishy with X.509 handling in Azure.
- PublicCoconut 3y agoRegarding: “requires detailed knowledge of internal infrastructure” MSFT decided to move their APAC support center for Office365 to China. If there is an issue in APAC engineers usually respond from there. It makes sense regarding time zone and cost. However, people change jobs etc. Hence this setup will result in a number of good engineers with very detailed knowledge of those systems who live in China. On the more tinfoil end of ideas: if the office for this work is in a particular country it brings corresponding risks for physical opsec of the infra.
- acdha 3y agoLooking at the validation section of https://learn.microsoft.com/en-us/azure/active-directory/develop/access-tokens https://learn.microsoft.com/en-us/azure/active-directory/dev... - did I miss something or does that still lack any mention of the importance of checking dates or revocation for the issuer? Since the pseudo code doesn’t I’d bet there are more implementations which trust any key Microsoft has ever published (modulo some kind of cache purge).
- trifurcate 3y ago> Since the pseudo code doesn’t I’d bet there are more implementations which trust any key Microsoft has ever published Bingo, this is the most worrying bit to me. Pasting from a comment I posted earlier: > Microsoft's own developers failed to implement a secure authentication check on top of their own libraries and infrastructure. If Microsoft can't use its own identity platform correctly in a flagship Microsoft product (Outlook), what chance does anyone else stand?
- alphager 3y agoThis is the core problem. Everybody is discussing the crash dump and the exfil, but the core problem is that Microsoft neither validated the validity of keys (the leaked key was already invalid) nor the context of the key usage (the key wasn't allowed to generate admin tokens). They just checked if the key was signed by the Microsoft CA. This is something that's incredibly obvious in a code review.
- nicknow 3y agoThere are entire systems engineering courses focused on failure resulting from a series of small problems that eventually in the right succession result in catastrophic failure. And I think we can say this was a catastrophic failure. Think about it, first you need a race condition, and that race condition has to result in the unexpected result. That right there, assuming this code has been tested and is frequently used, is probably a less than 10% chance (if it was frequently happening someone would have noticed.) Then you need an engineer to decide they need this particular crash dump. Then you need your credential scanning software (which again, presumably usually catches stuff) to not be able to detect this particular credential. Now you need an account compromised to get network access and that user has access to this crash dump and the hacker happens to get to it and grabs it. But even then, you should be safe because the key is old and is only good to get into consumer email accounts...except you have a bug that accepts the old key AND a bug that didn't reject this signing key for a token accessing corporate email accounts. This is a really good system engineering lesson. Try all you want eventually enough small things will add up to cause a catastrophic result. The lesson is, to the extent you can, engineer things so when they blow-up the blast radius is limited.
- xwolfi 3y agoRace condition is the reason we all use to explain to management why we wrote a stupid bug. Everything is a race condition: "the masker is asynchronous so the writer starts writing dumps before the masker is setup" sounds like a completely moronic thing to do. Say there is a race condition, and people say "a less than 10% chance from happening", but what do we know, maybe it happens each big crash, and it just doesn't crash that often. Why isn't it masking before writing to disk ? God only knows.
- michaelt 3y ago> Why isn't it masking before writing to disk ? Crash handlers don't know what state the system will be in when they're called. Will we be completely out of memory, so even malloc calls have started failing and no library is safe to call? Are we out of disk space, so we maybe can't write our logs out anyway? Is storage impaired, so we can write but only incredibly slowly? Is there something like a garbage collector that's trying to use 100% of every CPU? Are we crashing because of a fault in our logging system, which we're about to log to, giving us a crash in a crash? Does the system have an alarm or automated restart that won't fire until we exit, which our crash handler delays? It's pretty common to keep it simple in the crash handler.
- jokoon 3y agoHow did those developers had access to that key, and how did it land in crash dump? I smell like gross negligence here.
- yalok 3y agoI always hesitate to send the crash dump for this very reason...
- deleted 3y ago[deleted]
- lucasRW 3y agoI like how the fact that Microsoft's normal corporate environment (ie. the one with Internet connection) being compromised is casually considered as a secondary issue here.
- mall0c23 3y ago[flagged]
- markus_zhang 3y agoThe crash dumps, which redact sensitive information, should not include the signing key. In this case, a race condition allowed the key to be present in the crash dump (this issue has been corrected). Can someone explain a bit about what kind of race condition this could be and how it exposes the signing key (in this case seems to be exposed without encryption)?
- oger 3y agoOccam's razor tells me that this exploit route as described in the press release sounds way too complicated and that too much luck must have been involved. Therefore I am sceptical if we will ever hear the real story.
- freedude 3y agoThe list of failures is so long how much security engineering took place at Microsoft? It is possible the crash that contained the race condition was planned. But how would anyone know the race condition existed? We know now so it is likely, based upon the successful attack, it was known prior to the key being exfiltrated. How many of our organizations have the same race condition and our keys are/have been vulnerable? No one else has keys that are valid for more than a year, right? Let's say the race condition was unknown to the attacker. How can the attacker find the key in the dump? Did they have a scanning tool looking for key data? If so, how long was that running on MS network before it found the keys in the dump file? How many of our organizations' keys are in dump files? Are they all expired?
- _8j50 3y agoI was expecting "this was corrected" next to the log retention policy lol. Log collection is liking building an ark in the desert.