4 ms·
For example, if we had notifications in place to alert us 12 hours earlier that we needed more capacity, we could have added a third shard, migrated data, and t
by jranck 16y ago
For example, if we had notifications in place to alert us 12 hours earlier that we needed more capacity, we could have added a third shard, migrated data, and then compacted the slaves.
Where did Foursquare find their engineers? I hope no one lost their job here but this is pretty elementary stuff.
- driverdan 16y agoA bit harsh but a good point. There are a few red flags here. Resource monitoring on the servers, like you mentioned, seems pretty obvious. Especially considering that the system was almost guaranteed to crash when it ran out of RAM. Sharding needs to be managed in a way that evenly distributes the data. I would never do it by user ID unless a proper analysis of the data showed it to be a fairly even distribution.
- slantyyz 16y agoIt's probably a technical debt thing. FourSquare is immensely popular, and the team probably meant to have those checks and balances in place but got 'too busy' to implement them. That doesn't excuse anything, but these oversights can happen even when you've got primo talent on board.
- deleted 16y ago[deleted]
- harryh 16y agoIt's true that this is elementary in and of itself, but looking at things with a bit of a wider lens shows the complexity. We're a small engineering team (10 people) working on a product that is growing extremely fast both in terms of usage and feature set. Meanwhile we're also pretty much constantly re-architecting things to keep up with growth and also doing the immense work of growing the company up from 3 people to 33 and beyond (this has turned out to be WAY HARDER than I would have guessed going in). Further, there are lots and lots of different things that we need to be monitoring at any different time to make sure that everything is going ok and we aren't about to run into a wall. Automated tools can help a lot with this, but these tools still need to be properly set up and maintained. I'm not saying we didn't screw up. We had 17 hours of downtime over two days. We screwed up bad, and we feel horrible about it, and are doing a lot to make sure that we don't screw up the same way again. But it's not because we're morons that never thought about the fact that we should be monitoring memory usage. We just got overwhelmed with the complexity of all that we're doing at once. -harryh, foursquare eng lead
- chewbranca 16y agoI agree completely about the complexity of monitoring solutions for small teams with fluctuating applications. However, something as simple as htop running on an extra monitor would have alerted you of this issue long before downtime resulted.
- cullenking 16y agoAwww come on, it's easy to say from an outside perspective. Regardless, the problem was handled well in the end, and we all get the benefit of understanding these limitations better. I think us tech people got the good side (information) out of this ordeal :)
- chewbranca 16y agoI work on a team of 2 where I'm responsible for a handful of servers. I chimed in because I'm in a similar position, I've been looking for a monitoring solution for a while now. Things like nagios and zenoss are over kill, but lack of time has prevented me from finding an ideal solution. That said, I keep htop open and running at all times, and its saved my ass on more than one occasion. I say htop because of the color coding it provides, if things start going red it attracts my attention.
- sjs 16y agoNagios is pretty nice. It's dead easy to write custom monitors and clients are everywhere, there's even a Firefox extension. It requires a bit of learning to get going with but it's not so bad and the pay-off is big. That said I'm looking at monit too. I hear it's quite nice and has less of a learning curve.
- hartror 16y agoI'll put a vote in for Zabbix as a good option, it has saved us more times than I can count.
- jasonwatkinspdx 16y agoWhen you're a small company growing fast you have an endless list of things you really should have done a long time ago. That's no excuse for not preventing a predictable and painful failure, but do try to remember it's a lot easier to criticize as an observer with no time commitments.
- tripngroove 16y agoDisclaimer: I work at Cloudkick. We can help you all with these problems. Here's how fast/easy it is: 1. Create an account (~30 sec) 2. Add your cloud credentials (~45 sec) 3. Install the monitoring agent (~90 sec/node) 4. Create a CPU/Memory/Disk monitor, query-targeted all your servers (~60 sec) (example: "provider:EC2") 5. You get an email whenever the monitors you created reach the thresholds you set There are a lot of other cool features - check 'em out on the site: www.cloudkick.com All our plans are free for 30 days.
- daleharvey 16y agoWhile I usually frown on such blatant self promotion, I couldnt help but upvote this for being such a good advertisement, I want to use you now.
- al_james 16y agoOr you could install Ganglia for free.
- dedward 16y agoPerhaps it's just my own background that pushes me this way - but as soon as my deployment was growing ot eat up anything over about 60% of my available memory, I'd be looking at deploying more. Your high water mark for spinning up more nodes and/or upgrading shouldn't be your maximum physical ram capacity - that's just nuts.