16 ms·
A deeper dive into our May 2019 security incident
- CountHackulus 6y agoI found it interesting that the attacker looked for help on the attackee's own site. I guess it truly proves how good of a repository of information StackOverflow is.
- ballenf 6y agoThere's a new service SO could offer: help a company under attack or recently attacked correlate the methods with suspicious users on SO, based on IP addresses and the presumption that attackers would use the same system to get help as used in the attack.
- bentcorner 6y agoI feel like there's privacy implications around that. You'd need to be careful that it couldn't be abused.
- anticristi 6y agoShould probably be classified as a meta-breach. :)
- fredley 6y agoAlthough the article was written in an extremely straightforward and dry technical manner, this was comedy gold.
- brazzy 6y agoYeah, I'm pretty sure the reason why the article keeps repeating that is not a desire to provide the most detailed information about the breach...
- sradman 6y agoThe report describes a security breech in 2019; the report was held back until now for legal reasons: > Sunday May 5th > ...a login request is crafted to our dev tier that is able to bypass the access controls limiting login to those users with an access key. The attacker is able to successfully log in to the development tier. > Our dev tier was configured to allow impersonation of all users for testing purposes, and the attacker eventually finds a URL that allows them to elevate their privilege level to that of a Community Manager (CM). This level of access is a superset of the access available to site moderators. EDIT: clarified that the report was held back
- kmontrose 6y agoThe breach itself was announced shortly after it was discovered: https://stackoverflow.blog/2019/05/16/security-update/ https://stackoverflow.blog/2019/05/16/security-update/ And affected users were notified once identified, which was shortly after the announcement: https://stackoverflow.blog/2019/05/17/update-to-security-incident-may-17-2019/ https://stackoverflow.blog/2019/05/17/update-to-security-inc... This is an update with more details, which was held back for legal reasons.
- sradman 6y agoYes, thank you. My wording was ambiguous; my bad.
- blakesterz 6y agoThat was an interesting read. I'm left wondering "why" though. Anyone care to take a wild guess what they were after? That seems like quite a bit of work to be just doing it for no particular reason.
- plibither8 6y agoExactly what I thought. Probably more beneficial for attacker would be to report the security vulnerabilities and receive a bounty in turn.
- ballenf 6y agoGiven the focus on enterprise systems and teams, really looks like it was a Solarwinds type (but lower sophistication) attack where SO wasn't really the target. The targets were users of SO Enterprise or teams products.
- lima 6y agoIf that's the case, why would they elevate privileges on the main SO site and draw attention to their successful intrusion?
- lbriner 6y agoBecause they are working blind. They are trying to find something that they don't necessarily know exists. They also don't know what tripwires exist and after trawling around for so long, they might even have assumed that SE didn't have any monitoring systems.
- glenstein 6y agoOf the two responses here, yours strikes me as the more plausible, non-cartoonish one. I think it's good to come to these things with an understanding of how behaviors can be happenstance and come from an attacker negotiating with limited information or their own limited understanding.
- deleted 6y ago[deleted]
- dsr_ 6y ago"However, there is a route on dev that can show email content to CMs and they use this to obtain the magic link used to reset credentials." Zawinski's Law: "Every program attempts to expand until it can read mail. Those programs which cannot so expand are replaced by ones which can."
- Shog9 6y agoGonna clarify here, because that description is a bit misleading: this wasn't a route that allowed viewing sent emails, it was a route that allowed viewing what would be sent if a password reset was requested. The story behind that route might be interesting... See, originally Stack Overflow didn't have passwords - all logins were done via OpenID, so any credential management you'd need to do was done through your provider (Google, LiveJournal, myOpenID, etc) This made account recovery assistance pretty simple: given a verified email address, the system would just send that address an email that reminded the owner of any and all OpenID providers that they'd associated with their account. From there, it was up to the account owner to work with a provider to do things like reset passwords. Skip forward a few years, and Stack Overflow had its own OpenID provider - now you could sign up with an email and password just like a normal site, except really you were creating an account on https://openid.stackexchange.com/ https://openid.stackexchange.com/ - so the recovery process remained pretty much the same, just with a new provider that happened to be run by the same company. So far so good... Except, this was awkward to explain to folks. Really, that was what ended up killing OpenID: folks wanted a "Google" or "Facebook" button, not a whitepaper on fancy new authentication systems. At this point Stack Overflow decided to try to streamline the login process, making signing up and logging in with their own provider seamless: no need to know anything about OpenID. Now recovery emails started including password reset links, and also reduced or removed information on other OpenID providers that were associated with the account in an effort to reduce confusion. The decision tree for generating those emails got complex. And the decision tree for supporting users got complex as well. Support staff got frustrated; they'd been used to knowing what would and wouldn't be in a recovery email, and had a pile of templates ready to help folks navigate login issues based on that. But now they were getting replies back from folks who were confused and upset because their recovery email didn't contain information that the support person had asserted it would! This was the genesis of the vulnerable route: a way for support staff to ensure that they were providing accurate information to users about how they could recover their accounts. By the time of this attack, it was already obsolete; the login system had been redesigned twice since the confusing and complex system that first required it. It was vestigial and forgotten... The ideal breeding ground for vulnerabilities. (source: I worked at Stack Overflow through the time period described in this post, and was involved in support during the period when the relevant route was useful)
- lima 6y ago> A significant period of time is spent investigating TeamCity—the attacker is clearly not overly familiar with the product so they spend time looking up Q&A on Stack Overflow on how to use and configure it. This act of looking up things (visiting questions) across the Stack Exchange Network becomes a frequent occurrence and allows us to anticipate and understand the attacker’s methodology over the coming days. Awesome writeup - this gave me a good laugh :-)
- TacticalCoder 6y ago> However, there is a route on dev that can show email content to CMs and they use this to obtain the magic link used to reset credentials So many sites do this: allowing major changes to be effective immediately (like resetting credentials/password) by simply opening a "magic link" sent by email. I think that this "immediately" is a major security antipattern. I prefer it when such changes have a "cooldown" period of, say, 72 hours, during which the change is "ongoing" but not effective yet and during which the user can veto the change (say by either login on the site, where they'd then get a warning that a major configuration change is ongoing, and denying the change on the site or by opening another "magic link", sent by email, which allows to deny the change). It's not a perfect solution but it stops so many of these oh-so-common attacks dead in their tracks. Because there's a big difference between being able to read an email meant to someone (as happened here, on the server side) and being able to prevent a legit user from receiving emails while also being able to prevent that legit user from login onto a website with its correct credentials.
- NieDzejkob 6y agoI don't think it's a good idea to make a password reset take 3 entire days. In this case, I'd say the costs outweigh the benefits.
- curiousllama 6y agoYea, 10 minutes & a text message would suffice, IMO...
- LeifCarrotson 6y agoI think the most important part would be to give someone time to vet that it's legitimate. Stack Exchange has on the order of 100 developers, it wouldn't be hard to CC account creation or password reset notices to the manager of a new hire, and in that case, 10 minutes would often be enough to say "Uh, I haven't hired anyone named Curious Llama, who are they and why are they requesting developer access to an obsolete resource?" and put the brakes on.
- CapriciousCptl 6y agotldr; 1. Attacker found a stackoverflow dev environment requiring a login/password and access key to get in. 2. Attacker was able to login to the dev environment with their credentials from prod (stackoverflow.com) by a replay attack based on logging in to prod. 3. The dev environments allows viewing outgoing emails, including password reset magic links. The attacker triggered a reset password on a dev account, and changed the credentials. This gives them access to "site settings." 4. Settings listed TeamCity credentials. The attacker logged into TeamCity. 5. Attacker spends a day or so getting up to speed with TeamCity, in part by reading StackOverflow questions. 6. Attacker browses the build server file system, which includes a plaintext SSH key for GitHub. 7. Attacker clones all the repos 8. Attacker alters build system to execute an SQL migration that escalates him to a super-moderator on production (Saturday May 11th). 9. Community members make security report on Sunday May 12th, stackoverflow response found the TeamCity account was compromised and moved it offline. 10. Stackoverflow determines the full extent of the attack over the next few days.
- itsdrewmiller 6y agoKudos to the team over there for being as transparent about what happened and where they were not following best practices - I am pretty sure most companies would not publicly admit this: we had secrets sprinkled in source control, in plain text in build systems and available through settings screens in the application.
- akersten 6y agoInteresting that most of the mitigations are "move resource behind firewall." Kind of an indictment of the whole BeyondCorp idea - unless we really trust our 2FA to never have any access bypass issues like the initial access to the dev environment here. Speaking of that, I didn't see "fix bug allowing unauthenticated access to dev environment" listed as one of the mitigations, but maybe I glossed over it.
- invokestatic 6y agoI agree. I’ve employed the BeyondCorp philosophy behind a VPN as an extra measure of security, which is to say that all services are authenticated and encrypted inside the VPN perimeter. As shown in this article, service accounts are a major concern for attacker lateral movement which can’t be effectively protected with just 2FA.
- vntok 6y ago> Hardening code paths that allow access into our dev tier. We cannot take our dev tier off of the internet because we have to be able to test integrations with third-party systems that send inbound webhooks, etc. Instead, we made sure that access can only be gained with access keys obtained by employees and that features such as impersonation only allow de-escalation—i.e. it only allows lower or equal privilege users to the currently authenticated user. We also removed functionality that allowed viewing emails, in particular account recovery emails.
- deanward81 6y agoIt's in the remediations section, but maybe the wording isn't clear: *> Hardening code paths that allow access into our dev tier. We cannot take our dev tier off of the internet because we have to be able to test integrations with third-party systems that send inbound webhooks, etc. Instead, we made sure that access can only be gained with access keys obtained by employees and that features such as impersonation only allow de-escalation—i.e. it only allows lower or equal privilege users to the currently authenticated user. We also removed functionality that allowed viewing emails, in particular account recovery emails.* There was no "unauthenticated" access into dev - the access key here is what allows login at all to our dev environment, but the attacker was able to bypass that protection.
- 120bits 6y agoThank you SO for being open and listing the best practices. It seems like even few security best practices makes it harder for hackers to get in to your system. I have database connecting strings and password as ENV variables. But I still don't know what is the best practice. Lets say someone gets access to the server, they can still read the ENV vars, right? It definitely prevents from accidently checking in your code git repo. But still . Does anyone has good recommendation for storing credentials like database passwords in a way secured way.
- cbg0 6y agoI don't think there's a magic way to do this, if your app can connect to the database and someone has access to your app server - they have access to your database as well.
- dividuum 6y ago> Lets say someone gets access to the server, they can still read the ENV vars, right? Correct. Easiest way is to look at `/proc/$pid/environ`. It contains the \0 separated values for that process.
- cbg0 6y ago> Our dev tier was configured to allow impersonation of all users for testing purposes, and the attacker eventually finds a URL that allows them to elevate their privilege level to that of a Community Manager (CM) > After attempting to access some URLs, to which this level of access does not allow, they use account recovery to attempt to recover access to a developer’s account (a higher privilege level again) but are unable to intercept the email that was sent. However, there is a route on dev that can show email content to CMs and they use this to obtain the magic link used to reset credentials. Many of these debugging tools are great for devs to test things quickly but I've always felt very weary of having these exist in an app without some strict access control with 2FA. Ideally you'd not have them in the app at all, maybe just on local dev.
- Eduard 6y ago> Fortunately, we have a database containing a log of all traffic to our publicly accessible properties https://stackoverflow.com/legal/privacy-policy https://stackoverflow.com/legal/privacy-policy GDPR anyone?
- TrickyRick 6y ago> When you visit the Network or use our Apps, Stack Overflow automatically receives and records information from your browser or mobile device, such as your Internet Protocol (IP) address or unique device identifier. Cookies and data about which pages you visit on our Network allow us to operate and optimize the Products and Services we provide to you. This information is stored in secure logs and is collected automatically. Clear as day that they're doing exactly that. You agree to this when you use the site.
- eitland 6y agoCollecting data might well be acceptable with the GDPR. However what makes it legal isn't if it is written in the TOS or in a cookie banner. AFAIK what matters is either: - if you have a specific, valid (according to the GDPR) reason, - or if you have the users free and informed consent. ... and yes, I think a number of the things I still see on the web is not OK: - dark pattern where if you click manage settings everything is opted out, but there's a big green "Accept everything" and a small bland "Confirm my choices"? Doesn't fly because the rule that it should be equally easy to opt out. - Cookie banners with no real opt out? No way. - Cookie banners where you have to deselect 927 "partners"? Also no way. The only ones that seems legal are those who either uses a pure minimum of cookies for preserving state and those allow one to opt out directly but inform you that ads might become less relevant.
- whimsicalism 6y agoThe chronology has some issues in dating/day starting on "Tuesday May 15th" (Tuesday was the 14th) and continuing on.
- deanward81 6y agoOuch, good catch, fixing now
- mark-r 6y agoSomething to do with the difference between UTC and New York time zones, perhaps?
- Robelius 6y agoAs someone who doesn't work much with software teams, can someone fill in my gaps for understanding timeline. I'm imagining after a security issue is identified, the steps taken are roughly in the below order and close-ish for the date. I guess my question is, why does it take 20 months from start to blog post? -Contain the issue (1wk) -Remove the threat (1wk) -Build up remedies (a few months) -Check and recheck what happened to make sure you're accurate when submitting final reports (a few months) -Release a blog post (1month) The timeline is a cool day by day instance, but I just don't understand the larger timeline.
- tclancy 6y agoI assume it's related directly to "It’s been quite some time since our last update but, after consultation with law enforcement, we’re now in a position to give more detail".
- kmontrose 6y agoIt's this. Discovery, immediate mitigation, deeper mitigation, general notice, notifying effected users - all these can happen pretty quickly once the ball is rolling. Once you're dealing with "the law" in any capacity you are constrained in what you details you can share broadly, and when. I'm happy we were finally able to share this level of detail.
- riston 6y agoDid they figure out who were behind these attacks? These seems to be quite sophisticated and quite long taking attacks to dig so deeply into SO system.