7 ms·
All HTTP-based services unresponsive
- itake 9y agoits amazon having problems. https://status.aws.amazon.com/ https://status.aws.amazon.com/
- kkirsche 9y agoNorthern VA is having a ton of power outages due to high winds. NOVEC electric company is reporting power issues for almost 10% of customers so far
- cimmanom 9y agoYou would think Amazon would have sufficient backup power to make that a non-issue, as would any significant third party data centers involved in routing data. However, a downed data line might make sense as a root cause.
- thegeomaster 9y agoAnd it's only one AZ, so it sounds like Atlassian services aren't spread out over AWS properly. The recent S3 incident really highlighted the importance of this.
- bklyn11201 9y agoAmazon's description of S3: "Data is automatically distributed across a minimum of three physical facilities that are geographically separated within an AWS Region". What are people doing wrong with how they use S3? Until AWS provides a cross-region S3 that is master-master or self-healing, the suggestion that people are using AWS improperly seems incorrect.
- acdha 9y agoThat’s currently showing an issue with Direct Connect in northern Virginia. That seems like a bit of a stretch and it certainly wouldn’t say anything good about their DR planning if one region can take the whole thing down.
- colinbartlett 9y agoThere are actually quite a lot of services down across the web now. Maybe they are unrelated... but it could all be related to AWS. https://statusgator.com/services/amazon-web-services https://statusgator.com/services/amazon-web-services My side project, StatusGator, monitors something like 250 status pages and there's quite a spike in warn or down notices at the moment that I can see.
- edaemon 9y agoStatusGator is neat, thanks for linking that. Do you have graphs anywhere to track the number of outages/problems over time? It would be nice to see if there's been an uptick in problems generally, across certain services, etc.
- colinbartlett 9y agoNo, but that's a great idea! I have 3 years of data now from hundreds of services including severity of reported outage and text about why it went down. So I could show graphs over time for sure.
- shoover 9y agoAh. I wonder if that's why Capital One's login is down. I'm a little surprised their app isn't more resilient, but this is the second multi-hour outage I've noticed in the past couple months.
- filchermcurr 9y agoGitHub DDoS, Sourceforge DDoS, BitBucket 'routing issues'... somebody hates version control.
- rambojazz 9y agoThe target could be the companies, not VCSes per se. As long as they don't ddos notabug I'm fine with it :)
- smaili 9y agoAppears to be due to an upstream dependency: > Some component services are currently unreachable due to an upstream incident on a cloud provider. We're attempting to route as much traffic as possible away from the affected components, and are working with our vendor now.
- edwinksl 9y agoLooks like access over SSH is still working. Not the worst.
- simlevesque 9y agoThis is the last straw for me. I'm gonna stop using them to host my code. They have been down way too many times in the last year. It's been six failures from them in the last two weeks alone. I'm gonna self host Gitea to fix my issues. I cannot believe that they fail so hard. Why does a failure mean that I cannot read AND write from BitBucket ? Why are those two things even related ?
- bg4 9y agoDDoS?
- simlevesque 9y agoMy self hosted repos won't get DDoS. BitBucket is a large target and that's a problem for me.
- mtgx 9y agoIt's likely a DDoS attack against Akamai (again). Github also saw a record-breaking 1.3 Tbps attack recently: https://www.wired.com/story/github-ddos-memcached https://www.wired.com/story/github-ddos-memcached
- simlevesque 9y agoThat's something that is bound to happen more often. I cannot let that affect me. I cannot let their problems become mine.
- zbentley 9y agoYou use the internet. Their problems are already yours. The only situation in which self-hosting or ditching BitBucket/similarly large providers will help protect you from the fallout of historically-large, catastrophic attacks is if you do all of your development on the same local network as your hosted server (and don't rely on any internet services to access it). And even if you diligently self-host every part of your own services (not as easy as just plop a gitlab/gitea install on a host you own and start it up), you have to deal with the fallout from internet-breaking DDoS attacks and other malicious activity if you want to use the internet to run or use your code: from congestion caused by compromised devices in your network "neighborhood" (same/similar ISPs or last-mile providers) to DNS outages to BGP hacks, we have seen time and time again that, if not exactly centralized, the systems that comprise the usable internet are certainly highly interdependent. Large-scale attacks of many kinds compromise them. Instead of tantrums, it might behoove users to understand what kind of SLAs they can promise in order to operate their self-hosted services in such an interdependent environment. Some examples: Do you need local power? A local ISP to be up? More than one? If you have more than one internet link, how do you pair the connections--if it's via BGP, what happens if the central authority on that has issues? If local power is down, does your local ISP's connection stay up? How long does it stay up (is there a node/amp somewhere on the line that cut to battery)? Do you need to access internet services by hostname? If so, do you do local DNS caching? If so, how stale can it get in the event of a loss of external DNS? Most importantly: how much (it's a nonzero number unless you're developing for yourself, by yourself, on your LAN) dependence on external services are you comfortable with, and how much time are you willing to spend eliminating the long tail of such dependencies?
- thegeomaster 9y agoAnother ongoing discussion: https://news.ycombinator.com/item?id=16501731 https://news.ycombinator.com/item?id=16501731 Hosted JIRA is down too (at least for me). Interestingly, I can seem to be able to find it only via search, it doesn't show up on the frontpage at all.
- arkad 9y agoXaaS has many benefits, but uptime is not one of them anymore. I self-host my repos, had a few downtimes but thanks to this DDoS my local services have better uptime. ( Disclaimer: I know it's not apple to apple comparison as scale is massively different)
- convolvatron 9y agodistributed source code management theoretically doing this in a robust and replicated manner quite a bit easier. if you ignore partitions, it seems pretty straightforward to make a git push-all, and a recovery process for stale nodes coming back.
- david-giesberg 9y agoDavid from the Atlassian SRE team here. AWS Direct Connect is experiencing an outage in their US East Region: https://status.aws.amazon.com https://status.aws.amazon.com, which is causing connectivity issues for most Atlassian products and services. We're working hard to get everything back up and running. Please check http://status.atlassian.com http://status.atlassian.com for the latest updates. We're posting regularly and will continue to provide updates there.
- gtsteve 9y agoI'm not really a networking guy, so perhaps this is an obvious question, but why don't you have a failover configuration to send traffic over VPN or the public Internet? I would expect the latency to increase but otherwise still work. Is it a cost concern, is DC reliable enough that it's just an accepted risk, or is there some other reason?
- irenna 9y agoHello, I'm Irena from the Networking Engineering team at Atlassian. I have been directly involved with this incident and wanted to provide some answers to the questions. We’ve built our architectures based on the AWS Direct Connect service because it’s the most reliable and scalable solution based on our customer and network needs. The AWS Direct Connect service we use in the US East Region has multiple redundant links (4x 10Gbps) optimized for data throughput requirements and availability, and to our knowledge the AWS Direct Connect transit facilities have power backups that would help contribute to its reliability. But, as we saw from today’s event, something still failed. I should note that we have both publicly and privately reachable resources in AWS. The publicly reachable resources have fail-overs built in for situations like these (it happens automatically), but the private reachable resources with our architecture depend solely on AWS Direct Connect. For example, our Bitbucket failure today was due to the fact that we rely on AWS Direct Connect to link between the Bitbucket Cloud components that we host in our data centers and others that we host on AWS. Bitbucket could continue connecting to services in our own data centers and the public Internet/AWS, but could not talk to the privately reachable resources in the Atlassian infrastructure hosted on AWS. We understand the importance and the impact for our customers, and dedicated several teams to this issue as soon as it was reported. AWS has resolved the issue, but we will look into ways to help prevent and better mitigate these types of issues in the future as part of our incident review and improvement processes.
- krallja 9y agoStop using AWS US East!
- igammarays 9y agoCurious, why? Is us-east-1 known to be problematic? What about us-east-2 (Ohio)?
- insomniacity 9y agoBecause it was the first AWS region, it is known to have the oldest hardware, the most dirty hacks, and the most outages.
- tedmiston 9y agous-east-1 is the oldest region with the oldest hardware which is probably why it has more issues than others
- deleted 9y ago[deleted]
- murph-almighty 9y agoWhy? It's literally an AZ out of many that AWS provides.
- krallja 9y agous-east-1 feels like the region where Amazon rolls things out first. This has two primary effects: * sometimes new things break in unexpected ways * sometimes things get changed for the Rev. B and those revisions don’t get done in Virginia because they’ve already completed the Rev. A rollout. Also, there’s a secondary effect: because it’s the “default” region, it has a LOT more tenants, which means it probably has scaling and HA problems that none of the other regions do.
- bklyn11201 9y agoAWS has an easy nudge available to get people to begin using us-east-2 over us-east-1: introduce new features, new EC2 instance types, etc in us-east-2 first.
- deleted 9y ago[deleted]