8 ms·
Fog Creek is about to go down
- redler 14y agoWhat's a little surprising to me is that as of now www.fogcreek.com returns nothing but an error message. Presumably they have control of their DNS, and could quickly throw something -- even a simple web page with a status message -- on a server at some other datacenter.
- deleted 14y ago[deleted]
- ph0rque 14y agoWhat about Trello, is that implied as well?
- deleted 14y ago[deleted]
- seiji 14y agoThere's pretty much two ways to deal with this. Either admit this is a low probability failure scenario and it isn't cost effective to have global redundancies. The outage will be resolved as soon as possible. Or, admit you failed to build a georedundant HA infrastructure and apologize with a tentative plan to build out a redundant infrastructure in a different catastrophe zone. move the servers? On what planet is server infrastructure movable on a whim?
- deleted 14y ago[deleted]
- gecko 14y agoIt is movable. It would take a few days. We have another DC. We wouldn't say that if we knew it wasn't doable. EDIT: Just because it's feasible doesn't mean we will actually do it; just wanted to clarify that we weren't firing from the hip.
- tinco 14y agoWow you expect fog creek to be down for that long when it does? Why can't you just scp everything on there to your other DC? or at least just move the harddrives? Seems like it would cost less time.
- gecko 14y agoKiln alone has >4 TB of data; you want to SCP that with 90 minutes heads-up? Having power and/or rack space is not the same as having servers, switches, etc. anyway. Hopefully it will not be down that long. We'll let you know more when we know more.
- hellmans 14y agoThat 90 minutes claim is a bit disingenuous. We're at the same data center. Internap told us that the fuel pumps were flooded and asked us to shut down at ~9:30pm last night. The generators lasted until 10:52am. So 12 hours+ of warning. To be fair, at 9:30 they did warn that the generators would only last for 4-5 hours, but customers like us who proactively shut everything down extended that significantly.
- ynniv 14y agoYou could have rsync'd it with a few days heads up, and freshened that in the last 90min.
- modoc 14y agoDo you guys do off-site backups?
- unwind 14y agoToo bad that the backup generator refuelling pumps have been submerged (while the generators themselves are running). That sounds like some really ... unfortunate planning of the positioning of these machines, made me think of the Daiichi incident when backup assets failed to come online because parts of the backup infrastructure were destroyed. Not as serious, of course. Fog Creek's hosting isn't a nuclear power plant. :) Makes me glad I don't work with things that are as hard to test in the real world as these kinds of backup solutions must be. I hope they manage the refuelling, somehow.
- mikeash 14y agoThat's the first thing I thought of too. Seems like there should be a new rule that your backup generator infrastructure should be on at least the second floor.
- smackfu 14y agoMost basements are already waterproof (since you don't want groundwater getting in). It's generally the first floor that is the weak point, when the water gets over the basement walls. So if you think about it, it's not that hard to make the basement waterproof to a flood... just make the walls higher.
- emeraldd 14y agoThat means that the fuel tanks will have to be there as well or you're going to need some specialized lines/equipment to prime the pumps. How much would a 3" cylinder of diesel fuel about ten foot tall weigh? (What ever it is that would be a crap load of vacuum to produce/maintain). Then you have the weight of the tanks themselves and refilling logistics. All fun engineering problems. :)
- barrkel 14y agoWhat if you pumped it up by having a hydrostatically balanced system: instead of sucking up the fuel, pump water down to drive a piston that pushes the fuel up.
- andyjohnson0 14y agoSubmerged fuel pumps are the same problem that Internap are experiencing. https://news.ycombinator.com/item?id=4715889 https://news.ycombinator.com/item?id=4715889
- teuobk 14y agoFog Creek's services are hosted at Internap's LGA11 (75 Broad St) data center.
- samarudge 14y agoI'm assuming their servers are in LGA11, according to Internap the basement has flooded and damaged the fuel pumps for the generators. They also said in an earlier email they have no staff on site and are urging all customers to shut down their servers. Edit: here are the emails I've received from Internap regarding LGA11 https://gist.github.com/3980482 https://gist.github.com/3980482
- richardwhiuk 14y agoCut long lines to make it more readable. https://gist.github.com/3980570 https://gist.github.com/3980570
- sparkinson 14y ago> For our cloud customers, we will also being shutting down the infrastructure at this time. That bit made me chuckle.
- RobAley 14y ago"Given the preparation work that's gone into this, we are confident that all of our services will remain available to our customers throughout the weather." - yesterdays update. Try not to let your fingers type cheques your datacenters can't cash...!
- jrockway 14y agoNobody was really expecting this much of Manhattan to lose power.
- smackfu 14y agoWell, it seems like their confidence was based on their single datacenter not going down. Which seems misplaced.
- jpdoctor 14y ago... especially since there is more than one historical example for all of Manhattan losing power. (One of those times involved looting and civil strife.) Combined with the verbiage about "once-in-a-generation storm", it is fortunate that any part of Manhattan had power.
- count 14y agoNobody except the whole world watching the Weather channel. I'm frankly amazed that Manhattan survived!
- jusben1369 14y agoI would tend to disagree with you jrock. Nearly all the models talked about this storm wrecking this type of havoc.
- jrockway 14y agoI meant ordinary citizens, not emergency planners. Despite being told that they could be without power, many of my friends did not believe it. "How could Manhattan lose power?"
- kevingessner 14y agoWe are beginning to bring down all Fog Creek services (FogBugz, Kiln, Trello, etc.) as our datacenter is shutting down. fogcreekstatus.typepad.com and @fogcreekstatus on twitter will continue to have updates. edit: All Fog Creek services have been shut down ahead of power failure. We'll update the status blog as we know more from our DC.
- df07 14y agoStack Exchange (Stack Overflow) barely made it out. We are in the same datacenter but we just finished building out and testing a secondary datacenter in Oregon literally last weekend. We did an emergency failover last night after the datacenter went to generators. Read more at http://blog.serverfault.com http://blog.serverfault.com
- ChuckMcM 14y agoWow, nice timing there.
- cookiecaper 14y agoJust curious, wouldn't it have been wiser to put the failover servers somewhere in the Midwest? It's pretty much as far away from the ocean as one can get, making tsunamis/hurricanes/etc. irrelevant, low earthquake risk, and a shorter flight from NYC. Seems a little inadvisable to place the infrastructure in two coastal areas; I guess it's probably about the local talent pool.
- logn 14y agoThe Midwest isn't immune to the effects of hurricanes. We've frequently lost power, sometimes up to a week, when bad hurricanes come through. They just become super storms over the Midwest and take down trees which cut power lines. And snow/ice can often take out power for up to a day. So, I'd think that rather than place something in Ohio and New York, both which can be affected by the same weather pattern within a day or two, different coasts offers better protection.
- cookiecaper 14y agoIn raw geographical terms, Ohio only barely counts as "the Midwest"; I think it is included in that region primarily for cultural/economic reasons. What about Kansas City or Omaha? I lived in KC for some years and we'd occasionally get remnants of gulf hurricanes, but they'd just be severe storm systems that would pass through. We rarely lost power, but if we did, it was never for protracted time periods, and generators in data centers should easily be able to deal with power loss from severe storms. Tornadoes are pretty much the least threatening natural disaster out there, as their area of effect is usually small and their duration is usually short, so I think "tornado alley" is actually a fine place for a data center meteorlogically.
- Supreme 14y agoNo redundancy what-so-ever? What an amateur operation. I still say Joel is a fraud. EDIT: this site is amazing - divergent opinions seem to be actively discouraged given how many "points" I've lost thanks to stating mine. Is the point of this site for all of the members to think in the same way?
- tinco 14y agoOfcourse they have redundancy, just not cross-datacenter redundancy. And if you knew anything about cross-datacenter redundancy you'd know that cross-datacenter redundancy is something you do not decide upon lightly. Then again, having cross-datacenter backups that can easily be taken online would be a bit more professional than 'we want to physically move the servers'.
- Supreme 14y agoAre you kidding me? If you run big sites like FogBugz then ofcourse you have cross-datacenter redundancy. It's not complicated to host your staging site in another physical location and point the DNS records to it when things go pear-shaped.
- almost 14y agoSo which of your big sites have cross-datacenter redundancy? Why don't you talk about the decision process that lead to that and costs associated? Unless you're just talking out of your arse of course and you have no experience with that sort of thing at all.
- michaelhoffman 14y agoThe relationship between willingness to opine on a topic and knowledge of that topic: http://www.smbc-comics.com/?id=2475 http://www.smbc-comics.com/?id=2475
- tinco 14y agoYes, so this staging site of you has exactly the same databases as your production site? Without customer data Fogbugz and Trello are useless. This means that this simple staging site of yours needs to have all data replicated to it, which means it also needs the same hardware provisioned for it, effectively doubling your physical costs, your maintenance cost and reducing the simplicity of your architecture. Ofcourse, if you're big enough you can afford to do this, and one could argue fogcreek is big enough. I'm just saying it's not a simple no-brainer. What is a simple no-brainer how ever is to have offline offsite backups that can easily brought online. A best practice is to have your deployment automated in such a way that deployment to a new datacenter that already has your data should be a trivial thing. But yeah, if you're running a tight ship something things like that go overboard without anyone noticing. Remember the story of the 100% uptime banking software, that ran for years without ever going down, always applying the patches at runtime. Then one day a patch finally came in that required a reboot, and it was discovered that in all the years of runtime patches without reboots, it was never tested if the machine could actually still boot, and ofcourse it couldn't :)
- panda_person 14y agoI initially read that as Fog Creek the company was soon to be going out of business.
- ewams 14y agoReminds me of NASDAQ's post from over a year ago: http://news.ycombinator.com/item?id=2928519 http://news.ycombinator.com/item?id=2928519 Good luck FogCreek.
- tbourdon 14y agoI guess "The Cloud" got rained on in this case.
- mdc 14y agoWe use Fogbugz for all our internal project tracking. The consensus among our engineers is that this downtime is understandable and we'd rather deal with it, even in a mission-important web app, than pay more every month to insure redundancy was available. Frankly this is just making us appreciate Fogbugz all the more since tracking our time without it will be a real PITA.
- j45 14y agoMy first thought is that everyone affected is doing OK. I'm sure that everyone just wants their stuff working too, like electricity. :) I do hope at a better time, Fogbugz can consider redundancy/failover that Stack Overflow enjoyed.
- kranner 14y agoI don't know what Fogbugz costs, but they really should start charging for Trello now, esp. if they can use some of that revenue to add geo-redundancy. Trello is fantastic, but now I'm worried that I'm too dependent on it and I should arrange an offline alternative. Take my money, Fog Creek.
- stonnyfrogs 14y agoHi guys, Joel from Fog Creek here. We have absolutely no idea what failover means, or what the hell these people are talking about with this "two datacenters in different places" bullshit. Thanks for your money.
- deleted 14y ago[deleted]
- deleted 14y ago[deleted]
- calinet6 14y agoThis is a good reminder that no system is immune to failure, cloud or otherwise. Georedundancy is expensive and difficult, so it's a delicate trade-off, but engineering good physical backup systems is also difficult. Our servers are in a state far away from hurricanes, but in a state with many other natural disasters, including tornadoes, so it's hard to say if it's a good trade or not. Interesting question: why aren't there more DCs in Utah, Wyoming, Idaho, or New Mexico? And is physical location a huge determinant in where you colo your servers?
- Pwntastic 14y agoI know that Arizona is a huge place for DCs just because there aren't any natural disasters there. There are a few of them out here in Utah that I know of, but none at the scale that they really could be. It would make a lot of sense to put some out here, I would think
- meaydinli 14y agoI guess one of the most famous DCs in Utah is the NSA one: http://www.wired.com/threatlevel/2012/03/ff_nsadatacenter/ http://www.wired.com/threatlevel/2012/03/ff_nsadatacenter/
- drone 14y agoWe once had space in a datacenter in Arizona, but everyone else had the same idea. We had to move out and find a new datacenter when they were at capacity and we couldn't expand. While they were building a new facility, it was over a year away and our expansion needed to happen sooner. As a final point, the added space was already being pre-reserved at a premium, and we couldn't afford the new rates vs. other areas of the U.S. At the time, physical location wasn't a big deal, but as the company grew, and the data center overhead did as well - it became cheaper to have the core data centers closer to our operations, where our staff could be utilized. Georedunancy ultimately ended up being used for DR and minimum required service availability during major issues.
- lmm 14y agoMost places I've worked have had the first DC close to the office, then the second one in another country. When you're just starting up it's important to have easy physical access, and it's seldom worth migrating away from an existing DC rather than just opening up a new one.
- erre 14y agoWell, good luck to them, both personally and in bringing it back soon. I only wish Trello hadn't tried to reload on its own, so I could still see the screen before the shutdown. Now all I have is a blank page :(
- william_uk 14y agoI've managed to successfully failover to our own emergency backup version: https://twitter.com/williamlannen/status/263294924382937090 https://twitter.com/williamlannen/status/263294924382937090
- medinismo 14y agoCan someone explain to me how someone like Fog Creek would let an app like Trello go dark. Dont their carefully selected and perfectly screened engineers get paid gobs of money to prevent exactly this scenario from happening by having data centers in other locations replicate the one you have in your own house.
- danmaz74 14y agoLast time I checked, Trello was a free app...
- kalininalex 14y agoBecause, per Spolsky, Fog Creek wants to achieve a widespread adoption first before they start charging for it (with probably a free tier remaining). Other than that, Trello is very much a commercial app. Fogbugz IS a commercial app, and at $25/mo (if I remember right) it's in the same category as most other commercial apps, i.e. it's not particularly cheap. It's still down.
- cdmoyer 14y agoThey decide that the cost is not worth the benefit. Excluding back-seat systems engineers on sites like this, I suspect that most of their customers will be a bit upset, but give them the benefit of the doubt and be glad to pay slightly less monthly (or nothing for Trello) and suffer a short outage.
- blorenz 14y agoSure, the situation is not ideal. Trello is a boon to my productivity and has been a gift at being free. Spolsky has given so much to the community that I, personally, can tolerate this inconvenience to my workflow.
- trimbo 14y agoAs a newish customer of Fogbugz, I'm disappointed. This storm had days, if not a week, of head's up. While I agree "no one expected Manhattan to lose power" (another comment), as someone who heads up development of a SaaS product, I would have spent most of that week planning for worst cases and recovery. I constantly think about the worst case scenarios and how long they'll take to recover, even without huge storms bearing down on the data center. So, I'm disappointed. I really love the Fogbugz and Trello products. Now I'm in a position where I have to question whether we should depend on them.
- tghw 14y agoSo, you agree that no one expected Manhattan to lose power. No one expected that in 2003 during the blackouts, either. And in that case, Peer1 was able to keep the servers running without disruption. This was a monster of a storm, with unprecedented water levels. Buoys around New York reported waves 5 times higher than anything on record. So consider it this way: would you rather the services be significantly more expensive (remember, doubling the hardware is the cheap part of it) or have the possibility of a few hours of downtime in a once in 100 years event?
- abarringer 14y agoSeems their previous post[0] from last night showed a bit of overconfidence? "Consider this the "Everything is Perfectly Fine Alarm." Having run a few HA Datacenters I don't think that level of confidence is ever warranted. [0] http://status.fogcreek.com/2012/10/feelin-fine-no-expected-downtime-due-to-hurricane-sandy.html http://status.fogcreek.com/2012/10/feelin-fine-no-expected-d...
- shill 14y agoLatest tweet from PEER1, Fog Creek's datacenter. Storm #Sandy highlights value of #cloud storage http://www.peer1hosting.co.uk/industry-news/us-storm-highlights-value-cloud- http://www.peer1hosting.co.uk/industry-news/us-storm-highlig... https://twitter.com/PEER1/status/263197209959493632 https://twitter.com/PEER1/status/263197209959493632
- jwr 14y agoWe depend on FogBugz (hosted) to answer our support E-mails. If the downtime is on the order of several hours, I'm fine with it, these things happen. But if (as it looks like) it is on the order of days, I'll be looking for another solution. When you offer hosted services (not cheap, mind you), you take on responsibilities. Among them are disaster recovery scenarios. We do have ours and I'm expecting any company for a cloud-hosted solution to have theirs.
- meaty 14y agoThis is precisely why we have a self host requirement for all of our software. We did have a ton of stuff in salesforce but due to a number of problems with salesforce availability and the inevitable problem of relying on British Telecom's infrastructure monkeys, it got moved to a locally hosted dynamics CRM solution with off site transaction log shipping should the office catch fire. Cost a small fortune but there is nothing more expensive than not being there for your paying customers.
- kami8845 14y agoIf you don't mind my asking, who is `we` in this?
- meaty 14y agoWe have a comms NDA which prevents me revealing the company name but we're in the financial sector and are an old fashioned "enterprise company".
- kranner 14y agoJudging by current Twitter traffic for @trello, there is a clear need for a self-hosted version.
- tghw 14y agoThat's not going to happen.
- smackfu 14y agoCurious: What kind architecture shows a 503 error when your servers are dead (like they are currently doing), but can't show an error status page? Presumably that server is not in the dead datacenter. Or is it just that something at the datacenter level is redundant?
- lmm 14y agoAny kind of front-end reverse proxy could be doing that. At a complete guess, might they be using cloudflare?
- kevingessner 14y agoWe've shut down all of our servers to protect data, except some of the outermost infrastructure and gateways. That 503 is coming from HAProxy, our load balancer -- it's unable to send your traffic to any of the (powered-down) servers.
- patrickgzill 14y ago"When the tide goes out, you find out who is swimming naked." Warren Buffett said it in reference to economic problems, but somehow it applies...
- rdl 14y agoWow, that kind of sucks for them. I hope they consider prioritizing some kind of geographic replication after the storm is done. It adds cost and complexity (which slows down development, too), but seems like a good tradeoff when you have customers depending on it. The geographic load balancing side is basically a solved problem (although you don't want to use only DNS-based load balancing like Route53 in most cases), but the hard part is wide area replication of databases for hot failover. It's pretty easy to do failover if you'll accept a 5-10 minute outage, though.
- eric5544 14y agoLet me join in the chorus of people that are slightly miffed that I don't have access to my Trello boards. In a way it highlights again the importance of local storage and off-line accessibility for web-apps. I just checked and in the trello app on my nexus 7 I can still see (and browse) all my boards for example (the only problem is that the content is a few days old as I have not used the app recently)
- spolsky 14y agoI'm kind of an optimist; I believe it'll only be a matter of hours total outage. The generator is fine, the equipment is fine, the internet connectivity is fine... the only problem is getting fuel up to the generator on the 17th floor, while the fuel pumps in the basement are submerged. Someone will carry it up 17 flights if need be.
- hashtree 14y agoI feel for the guy/gal/group that draws the short straw and has to tote a 400-ish pound 55 gallon drum of diesel up 17 flights. "Whelp that lasted for 5 minutes, again!" ;)
- emmettnicholas 14y agoThey are carrying half-full 50-gallon barrels of diesel (source: http://forums.peer1.com/viewtopic.php?f=37&t=7532&sid=d17175713731c19a4c9a4ce1e2b7bafc&start=10#p9475 http://forums.peer1.com/viewtopic.php?f=37&t=7532&si...). Still, carrying 200 lb barrels up 17 flights of seawater/diesel-slicked stairs sounds ...unpleasant.
- bonaldi 14y agoGiven an average requirement of 500 gallons an hour? That's nearly 2 tonnes. Nobody's carrying that up 17 flights.
- bonaldi 14y agoOh, you magnificent eejits http://status.fogcreek.com/2012/10/diesel-bucket-brigade-maintains-services.html http://status.fogcreek.com/2012/10/diesel-bucket-brigade-mai...
- alkaramba 14y agoGood luck, guys!
- theycallmemorty 14y agoThey've since completely evacuated the building, no?
- whatgoodisaroad 14y agoDoes anyone else find it funny to think of the internet running on diesel?
- hashtree 14y agoYou might be surprised to know how much is normally running on petroleum: http://en.wikipedia.org/wiki/List_of_power_stations_in_New_York#Petroleum http://en.wikipedia.org/wiki/List_of_power_stations_in_New_Y...
- dbecker 14y agoI keep hearing that this is the worst natural disaster that people (who are currently alive) have ever seen in NYC. It's a bummer that their sites are down... but I think I can go a day without my to-do list when they have a once-in-a-lifetime natural disaster. For perspective, it's not like they have down-time once every few months.
- joevandyk 14y agoWhy would loss of power cause "unrecoverable data corruption"? Don't modern databases work hard prevent this sort of thing?
- atesti 14y agoThey use SQL server for StackOverflow, FogBugz and Kiln. But for Trello they use only MongoDB.
- drewcrawford 14y agoSo a couple of things: 1) Like any prepared person, I got all my data out of the east coast before this whole hurricane thing. If someone from Fog Creek hooked me up with some emergency licenses while they got their stuff sorted out, we'd be fine. Actually, this would be a good time to switch to self-hosted. 2) Second, while I was doing our hurricane prep, I ran into this blog post from Joel: > Copies of the database backups are maintained in both cities, and each city serves as a warm backup for the other. If the New York data center goes completely south, we’ll wait a while to make sure it’s not coming back up, and then we’ll start changing the DNS records and start bringing up our customers on the warm backup in Los Angeles. It’s not an instantaneous failover, since customers will have to wait for two things: we’ll have to decide that a data center is really gone, not just temporarily offline, and they’ll have to wait up to 15 minutes for the DNS changes to propagate. Still, this is for the once-in-a-lifetime case of an entire data center blowing up http://webcache.googleusercontent.com/search?q=cache:lHEK939AKiEJ:www.joelonsoftware.com/items/2007/07/09.html+&cd=5&hl=en&ct=clnk&gl=us&client=safari http://webcache.googleusercontent.com/search?q=cache:lHEK939... Obviously this was written in 2007, but they claim to be geographically redundant and have geographic backups that are "never more than 15 minutes behind". Presumably things haven't deteriorated since then.
- larrys 14y ago"and they’ll have to wait up to 15 minutes for the DNS changes to propagate" Not sure I see the need for a propagation delay if customers can be pointed to the new site by simply using domainbackup.com instead of domain.com (in other words completely separate dns as well as a completely different domain (even through a completely separate registrar) to a site hosted elsewhere. They can know this in advance of course.
- atesti 14y agoThat text says "To implement this warm backup feature, I wrote a SQL mirroring application that implements transaction log shipping: ..... Right now, we’re log shipping twice a day, so you might lose a day of work if an entire city blew up, but in a couple of weeks, we’ll implement a system that does more continuous backups, and we expect that the warm backups will never get more than 15 minutes behind." What happened to that? Was it turned off? How long did it last? I also wonder how long FogCreek will still maintain the for-your-server version of FogBugz? Will it still be available next year?
- kevingessner 14y agoAnd we're back! All Fog Creek services are back on line. Our datacenter has enough fuel for several hours and is working on getting a delivery of more. We are hoping that Kiln, FogBugz, Trello, and all our services will remain up, though things are still a bit dicey. The details are all here: http://status.fogcreek.com/2012/10/fog-creek-services-updates.html http://status.fogcreek.com/2012/10/fog-creek-services-update... Thanks everyone for your patience!
- mfrankel 14y agoUnfortunately smaller Peer 1 companies like mine were told that power was going to be shut and we therefore brought down our servers. We're still down. I'm happy for FogCreek, and I'm generally happy with Peer 1, but I wish they would have been honest with everyone in this situation.
- mhp 14y agoThe issue is communication between the boots on the ground and the NOC. Earlier today, we brought our servers down too. People that stayed up (squarespace) have been up the whole time. Until I understood the whole situation, we made the same choice you did. Now that I know more and I spent the day at the DC, I realized we should get 1-2hr warning before power goes out. Peer1 should be running even after all fuel in the header (shared tank) on 17 runs dry... at least for a little bit. I can't guarantee it, but that's my take on the facts I have at hand (I am President of Fog Creek and spent the day at Peer1 and our office. They are currently using my aquarium pumps to try to pump diesel to the tank on the 17th floor).
- mfrankel 14y agoThanks for your assessment. I actually chose Peer 1 because of the recommendation from FogCreek and I'll still thank you even after this incident. I think it's ironic, that Peer 1 is getting accolades from Business Insider's squarespace article, while their misinforming email has hurt my firm and the small companies that use our software. I won't hold my breath for an apology email. That's life in the big city.