4 ms·
Our sincere apologies for tonight's downtime. We're back up now after 30 incredibly frustrating minutes, but we're making changes to ensure this incident can't
by seldo 12y ago
Our sincere apologies for tonight's downtime. We're back up now after 30 incredibly frustrating minutes, but we're making changes to ensure this incident can't be repeated.
The root cause was a network failure at our CDN, Fastly. The incident was limited to a single Point of Presence (POP) in San Jose, so if you were in Europe or Asia you didn't see anything wrong, but obviously at this time of day most traffic is from the west coast.
While our uptime over the last few months has been pretty great, in the last week we've had two non-trivial incidents. That's unacceptable to our users, and to us, and we're not just sitting around hoping it doesn't happen again, but will be making fundamental architectural changes to eliminate the sources of failure we've seen.
- thedaniel 12y agoHow fundamental are we talking here? Is npm still going to be a couchapp?
- seldo 12y ago99.9% of requests to npm over the last 3 months have been hitting things other than couch. Binaries are being served directly from disk by nginx, and 99% of JSON GETs are served from our CDN's cache. Couch is still the source of truth, but we treat it much more like a database than an app these days, and it's been much more reliable that way. The only time you're really hitting couch directly these days is when you publish. Almost none of our downtime since February (when we sorted out some bugs in our app that were affecting couch performance) is attributable to couch. Mostly it has been network-related problems at the caching layer, and the fixes we need to make involve making failover faster and more reliable, as failures are inevitable in any large distributed system.
- bjornstar 12y agoI'm in Tokyo and it was definitely down for me and now it's down again.
- peterbraden 12y agoI'm in Zurich and it was also down here.
- seldo 12y agoYes, we had another 20 minutes of downtime 4 hours later, caused by our CDN accidentally re-instating their broken datacenter. The second outage is documented here: http://status.npmjs.org/incidents/jc65gc8tzk5v http://status.npmjs.org/incidents/jc65gc8tzk5v