9 ms·
Post mortem of a failed HackerNews launch
- antirez 14y agoWhatever was the real cause for your issues, Linode's default small swap space is a plague. A system starts to misbehave much gently if there is enough swap.
- zorlem 14y agoFor a production server I think that the opposite is a better _general_ advice - reduce the available swap, because if your server gets to a point to need it, the performance will suffer so much that your server will become completely unresponsive. Having less swap will allow the OOM to kill the run-away process and allow you to login and fix the problem instead of rebooting the server or waiting in vain for it to recover by itself. edit: typo
- X-Istence 14y agoDrop your swap on a separate drive from your main drive that is serving your data. Solves the problem nicely.
- deleted 14y ago[deleted]
- mp3tricord 14y agoOnce memory goes to swap you already lost. Personally I rarely configure swap on servers, save the DB. I would reconfigure your services to not grow past physical free memory. After that you are going to have to scale servers horizontally.
- elchief 14y agolamesauce. 1. HN should let you pay them $10 and let them hammer your server(s) before your story goes live. good for you. good for them. 2. there's a deal at lowendbox right now for a 2GB VPS for $30 a YEAR. you could have a healthy server farm for pretty cheap.
- achacha 14y agoThings you should do before going live: load test obviously and monitor how memory gets used (if you hit swap you are too low of memory or mem consumption is too high and needs to be separated/load-balanced). CPU usage should be monitored and if you are going over 90% for a long duration you may be putting too much in one instance. Things like that, but from what I gathered your biggest flaw was not doing any load testing (jmeter, floor, etc, lots of tools to help you). I've been doing performance and optimization most of my web life (at least since 1992), you are not alone, people throw un-loadtested sites into the wild all the time and fail every time when slashdot/reddit/hn -effect occurs. But you can always do better now that you have learned from this failure.
- fox91 14y agoPlease, fix your CSS for mobile use. It's impossible to read because, if I zoom, the sidebar gets bigger too
- runarb 14y agoI have been having some of the same issues on a site I run ( http://www.opentestsearch.com/ http://www.opentestsearch.com/ ). Under heavy load solr will grind to a halt if you don't have enough ram available. Putting a dedicated Varnish server in front of the search servers helped a lot. Using a cdn may also be a viable option, but haven't tried it myself.
- druiid 14y agoWell, and this is why I recommend running solr on a standalone instance. It (and java/jetty/tomcat/etc.) are very memory hungry in general, so it is worth your while and money to spend a bit more and spin up a separate instance or whatever type of services you are using to run solr. It'll also run faster. One last thing you can do if none of that is possible is use a better VM like Jrockit (http://www.oracle.com/technetwork/middleware/jrockit/overview/index.html http://www.oracle.com/technetwork/middleware/jrockit/overvie...). Jrockit with the right GC in my experience is much better about running in lower memory type situations.
- bdcravens 14y agoFirst off, best of luck with your project. Secondly, kudos on writing the post-mortem, as I know it takes some guts to own a "failure". I think, however, the need to write something like this speaks to an incorrection assumption: you need a "launch". Of course, TC and HN can give you a nice bump in traffic and even signups. However, in the long run, this really doesn't accomplish much for you. It gives you the kind of traffic that will likely leave and move on to the next article, skewing your metrics. There's certainly qualified prospects in there, but it's hard to decipher with all the noise. Again, the concept of a "launch" speaks to poor business models. It really benefits businesses where the word "traction" is more important than "revenue". Build a business that provides a service that others will pay for and grow as fast as the business can bear, bringing in those visitors that are truly valuable to you.
- gtd 14y agoI know as well as anyone the relative futility of relying on HN, Reddit, or TC coverage for building a successful tech product. Feedback and traffic from social news is merely a blip that says next to nothing one way or another about your long-term prospects. However, if your site goes down for any reason a postmortem of this sort is definitely warranted. The word "launch" is not signifying much more than a point in time in this case, and I think you're jumping to a lot of conclusions about what hopes they were pinning on this event.
- dklounge 14y agoRight on. I second these sentiments. First, keep up the good work and best of luck moving forward. Very good that you're also reflecting on your successes and failures - always be learning. The most challenging piece of a new business is, well, new business. And it's about growing your value proposition organically, one customer at a time, and refining the business. Analyzing bump in media attention won't really help you on that piece of the search. Once you've nailed down the search, and you're simply focused on getting more publicity as you scale, then perhaps that sort of analysis will be of more use. But, I doubt it.
- gingerjoos 14y agoI think the author wasn't just looking at a launch per se. He was looking for feedback and HN is arguably the best place for feedback w.r.t startups > HN community’s remarks and constructive criticism are pearls of wisdom
- maxent 14y agoYour project, Cucumbertown, is a cooking/recipe site/platform/network. Hacker News is not your audience/customer. Any "launch" on Hacker News is a fail, regardless of downtime.
- tzaman 14y agoCouldn't disagree more. HN community is the most (brutally) honest and it's feedback can prove very valuable - regardless whether HNers are the target market or not.
- envex 14y agoJust because you like hacking/coding/engineering/hn-job-title-here, you can't enjoy cooking?
- bdcravens 14y agoI've never thought of HN as a marketing platform. It's a place for Hackers (perhaps the "news" fails to encompass the full site) to discuss PG's concept of hacking entrepreneurs. Far too many Show HNs target developers.
- ErrantX 14y agoCooking often requires hacking. Engineers love to cook. Also, startups often look for cheap/simple food to sustain them. Cooking/food articles consistently do well here.
- lmm 14y agoRunning with swap enabled is a terrible idea. The authors mention how it was only once solr crashed that they were able to actually log in and start fixing problems; having swap means that rather than the OOM killer terminating processes, instead your whole system just grinds to a halt. (it's strange that they recommend enabling swap when they also recommend enabling reboot-on-oom, which is pretty much the complete opposite philosophy)
- richardwhiuk 14y agoIn my experience, the linux kernel handles no swap at all very badly, so you need a small amount. Increasing the swap, which is the suggested solution, is however, a terrible idea. As soon as you hit high memory usage, your IO load will go through the roof, and everything will grind to a halt. The solution here is separation of services - i.e. put Solr on a different box, so that if it spirals it doesn't take out other services. The OOM killer is your friend for recovering from horrible conditions, but as soon as you hit it or swap, somethings gone wrong.
- haberman 14y ago> In my experience, the linux kernel handles no swap at all very badly, so you need a small amount. Why? I'm pretty sure we disable swap at Google. Maybe swap was necessary back in the days when memory was really tight, but it seems like a terrible idea now. Especially since the scheduling is completely oblivious to swap AFAIK, which means that a heavily swapped system will spend most of its timeslices just swapping program code back into memory. It's the worst kind of thrashing.
- Cherian 14y agoYou are right. The best solution is separation of services. But for a startup than runs 7-2 services like this – it’s a close call. You’ll often have to run 2-3 services together, else $100 * 7 machines is too much burn
- zorlem 14y agoIt's not a problem running several services on the same box as soon as each of them is sized appropriately. What I suggest at least roughly calculate how much eg. RAM could each service use at peak time. This usage should be limited so that the sum of memory used by all services at peak time is less than the amount of RAM you've got on your server.
- buro9 14y agoIf anyone owns a blog or site that they suspect may appear on HackerNews (especially if you're posting it), then please take the small amount of time to put an instance of Varnish in front of the site. Then, ensure that Varnish is actually caching every element of the page, and that you are seeing the cache being hit consistently. You should expect over 10,000 unique visitors within 24 hours, with most coming in the 30 minutes to 2 hours after you've hit the front page on HN. You need not do your whole site... but definitely ensure that the key landing page can take the strain. Unless you've put something like Varnish in front of your web servers, there's a good chance your web server is going down, especially if your pages are quite dynamic and require any processing.
- deleted 14y ago[deleted]
- Cherian 14y agoCucumbertown co-founder here. Nginx was serving the cache and our sense was it was caching. But then the day before we put in csrf validation to the login form and it was bypassing the caching. So in theory we were positioned to serve from Nginx cache.
- gingerjoos 14y agoThe blog mentions that they did have caching on with nginx (which is what Varnish does, isn't it?). The problems were because of nginx not caching the frontpage (configuration issue) and because there was an unexpected hit on solr.
- boundlessdreamz 14y ago1. So nginx didn't cache because of cookie? 2. Isn't swapping bad? I don't think I've ever had a situation in which swap more than say 100MB was helpful. Once the machine starts swapping, a bigger swap just prolongs the agony. 3. If you couldn't ssh, why didn't you just reboot the machine? Edit: 1. What did you use for the graphs? 2. What is the stack?
- Cherian 14y agoWe use DataDog - http://www.datadoghq.com/ http://www.datadoghq.com/ . It’s a statsd, graphite manifestation but with much more capabilities. Stack is Python, Django, PostgreSql, Redis, Memcache etc.
- debacle 14y agoI clicked on the link to Cucumbertown and was immediately greeted with a picture of Italian seasoned chicken thighs. I think I really like your website. I really like the simplicity of the presentation to the user.
- nemesisj 14y agoThis is a great way to make lemonade out of the lemon of getting hosed by a lot of traffic. Write an informative post-mortem and resubmit! I know I missed the original submission and clicked through to the site, and there you have it. I'd say being humble and trying again is never a bad idea.
- jsaxton86 14y agoThis post mortem has me thinking about the best way to handle the situation in which you can't SSH into your server. The OP decided to trigger a kernel panic/restart on OOM errors, but I have a couple of concerns about this approach: * If memory serves correctly, if your system runs out of memory, shouldn't the scheduler kill processes that are using too much memory? If this is the case, the system should recover from the OOM error and no restart should be needed. * OOM errors aren't the only way to get a system into a state where you cannot SSH into a system. It would be great to have a more general solution. * Even if you do restart, unless you had some kind of performance monitoring enabled, the system is no longer in the high-memory state so it will take a bit of digging to determine the root cause. If OOM errors are logged to syslog or something, I guess this isn't a big deal. I suppose the best fail-safe solution is to ensure you always have one of the following: * physical access to the system * a way to access the console indirectly (something like VSphere comes to mind) * Services like linode allow you to restart your system remotely, which would have been useful in this scenario
- adrianpike 14y ago* In linux-land, there's an OOM killer (http://linux-mm.org/OOM_Killer http://linux-mm.org/OOM_Killer) that would have started taking processes out. You have to exhaust swap for it to really take effect, and once you hit swap, your entire machine suddenly becomes hugely IO bound - in shared or virtual hosting environments, this usually makes the machine totally unresponsive. * I've never seen any sort of virtual hosting service without either a remote console or a remote reboot. Usually both.
- lnanek2 14y agoWow, nice post with real data, graphs, and helpful tips. As the Germans would say, I have nothing to complain about.
- TeeWEE 14y agoStress test, load test before launch! It doesn take more then an hour, and you quickly know what your upper limits are, and where the bottlenecks are. I use gatling in favor of JMeter: https://github.com/excilys/gatling https://github.com/excilys/gatling
- pothibo 14y agoThat's why I like to use Heroku/EC-2 for launching new webservice. If shits hit the fan, you can jack up the processing power/database/RAM/whatever to scale to your demand. Once you have a good idea of the traffic it generates, you can then move it to a cheaper service. Obviously, it's easy to say that when you're on the bench. Congratulations on the launch by the way.
- Cherian 14y agoCucumbertown co-founder here. Actually I dislike this idea though we should have been better prepared. At my previous firm we had this culture that whenever traffic peaks we spin up new instances. And tools like RightScale & Chef make it ridiculously simple. So our style was to do that than to optimize strains in code paths. Because this is so so convenient. And before you know it, this notion of hardware is cheap becomes a culture. Soon enough if you grow you’ll be serving 100K users with 250 machines.
- amikazmi 14y agoSounds like a good thing to me- If it was better for your previous firm to pay more than to optimize, it's actually preferring developers time (which cost money, as you know) over servers cost. This can be cost effective until some point. I don't think it could get to the level of "100k users on 250 servers", and if it does. If the other side of the coin is that you waste dev time AND that your site is down for a few hours.. Is it really worth the "culture fear"?
- pothibo 14y agoI understand what you are saying. I do agree that Chef [and RightScale? Never used it) makes it easy to spawn new instances and through load balancing average your load. I was talking in term of tradeoff in the first few weeks of a new service with a MVP. Obviously, you re-assess your need before you get to 100k user, and probably uses something else than EC-2. In any case, both ideas are equally good. I don't claim better knowledge in any way.
- dkokelley 14y agoI think throwing more resources at the problem is a quick and dirty solution for when things go downhill quickly (like what happened here), and having that option is incredibly nice. Still, it should take second priority to proper configuration tuning in the long term. Also, in some instances of runaway memory, there will always be a point where all the memory in the world isn't enough.
- saltcod 14y ago[ not that you asked for it here, but I've got some frontpage UI feedback: ] I think you should put a description up front to describe what Cucumber town is. I think that main image should be a slider with multiple feature images, and I think the Latest Recipes should be the first section after this. Just my 2c! Screen: http://cl.ly/image/3R2Y131Z433L http://cl.ly/image/3R2Y131Z433L
- Cherian 14y agoThis is wonderful feedback. We’ll definetly look into this.
- alexbrand09 14y agoI am currently building a site and this is definitely an experience that I can learn from. I am wondering, why was the homepage not being cached?
- zorlem 14y agoThere was a cookie set for CSRF protection and the headers specify that content should not be cached if there is a cookie (or more precisely - the cached content includes the cookie as a cache key, so each request with a different cookie gets a cache-miss).
- alexbrand09 14y agoHow would you circumvent this? I'm thinking that disabling CSRF is probably a bad idea. Maybe use AJAX to get the CSRF token after page load?
- digitalpbk 14y agoDeveloper here, since we had a click to open form at the time, we loaded the CSRF via AJAX. However that does not seem to be a good idea if we need it to work asap (and without javascript). I would look at something like SSI to put in the CSRF token to a cached page.
- Cherian 14y agoSlightly different. "Note that in 0.8.44 behaviour was changed to something considered more natural. As of 0.8.44 nginx no longer caches responses with Set-Cookie header and doesn't strip this header with cache turned on (unless you instruct it to do so with proxy_hide_header and proxy_ignore_headers). " http://forum.nginx.org/read.php?2,126312,126316#msg-126316 http://forum.nginx.org/read.php?2,126312,126316#msg-126316
- zorlem 14y agoThanks for the clarification.
- mekoka 14y agoThanks for this post, there were some nice tips in there. Although, I do have some nitpicking about your writing style. Maybe it's just me, but I found that your use of "+ve" instead of just saying "positive" and of "&" instead of "and" did not have the intended effect of speeding up reading, quite the reverse actually.
- ajross 14y agoAgree with the former, but it bears saying: you're about 2 millenia too late to be complaining about the use of the ampersand. :)
- mekoka 14y agoI'm not complaining on the use of the ampersand, but unlike many people seem to believe, it's not semantically equivalent to "and". Beyond just joining two items in a phrase, the ampersand marks an association between them and emphasizes it as a single definite idea, a "thing". Ampersands are often used to mark brands, names and cultural items made up of multiple components: Johnson & Johnson, Dungeons & Dragons, bread & butter, fish & chips, Gold, Smith & Associates. If I say "I had some fish & coleslaw" that would make a few people wonder if this is some popular recipe they should google.
- politician 14y agoSeconded. Initially, my brain told me that +ve was the name of the site, so I was confused when I looked for a product page link only saw "Cucumbertown". Granted, that's mostly laziness -- apparently I've got a rule that matches "strange words near the top of the post" to "probably the name of the product".
- chacham15 14y agoI dont know about anyone else, but when I saw "+ve", I just thought to myself, "what is that?" for about a half a second before giving up and moving on.
- 14y ago
- deleted 14y ago[deleted]
- chacham15 14y agoIs there a way to run simulated traffic to determine how your server will react based upon heavier load to try and determine how many people it can serve?
- gingerjoos 14y agoYes, there are. A previous comment by TeeWEE [1] mentions 2 tools. There are others as well - Siege [2] and Apache Bench [3] come to mind. [1] http://news.ycombinator.com/item?id=4847949 http://news.ycombinator.com/item?id=4847949 [2] http://www.joedog.org/siege-home/ http://www.joedog.org/siege-home/ [3] http://httpd.apache.org/docs/2.2/programs/ab.html http://httpd.apache.org/docs/2.2/programs/ab.html
- pearkes 14y agoBlitz is also worth a mention. http://www.blitz.io/ http://www.blitz.io/
- druiid 14y agoBlitz is good, but I've found the way to build test-cases fairly limiting. I've used http://www.loadimpact.com http://www.loadimpact.com to good success. Addendum: Oh, load impact is pricey though.
- ohashi 14y agoCan also connect newrelic with it and get stack traces logging how much time each component of your app is taking including DB queries.
- ck2 14y ago1. Reduce keepalive, even with nginx 60 is too much (unless it's an "expensive" ssl connection). 2. set vm.swappiness = 0 to make sure crippling hard drive swap doesn't start until it absolutely has to 3. Use IPTABLES xt_connlimit to make sure people aren't abusing connections, even by accident - no client should have more than 20 connections to port 80, maybe even as low as 5 if your server is under a "friendly" ddos. If you are reverse proxying to apache, connlimit is a MUST.
- Cherian 14y agoThis is excellent advice! Thank you.
- jacquesm 14y agoDon't bother reducing keepalive, just disable it altogether. Unless you have a very specific use case it is more trouble than it is worth.
- ck2 14y agoA small keepalive helps prevent browsers from trying to open too many connections and reuse existing ones more efficiently from what I have seen. Nginx handles connections more efficiently than apache so it doesn't hurt. But even Apache can benefit from a couple seconds of keepalive to get less thrashing. You can use the optional second part the keepalive_timeout setting in nginx to send timeout hints to modern browsers, ie. keepalive_timeout 10 10; Some servers like litespeed have the easy ability to do keepalive for static content (ie. a series of images) and then connection close for dynamic. This behavior can be emulated by nginx with the right configuration.
- zorlem 14y ago> Don't bother reducing keepalive, just disable it altogether. Unless you have a very specific use case it is more trouble than it is worth. Bad idea. This way you're actively increasing the latency of your site. This way, for each asset that has to be fetched you're forcing the client to open a new connection, which can add more than 150 ms of delay per item (thanks to the three way TCP handshake). What I would suggest is setting the KeepAlive timeout to a value that could handle each individual page-load. This way all the page elements will have a chance to use the connections that has been already opened.
- James_Henry2 14y agoThis is really interesting, thanks for sharing. I love this kind of transparency.
- driverdan 14y agoI'd argue the opposite of your headline, that this was a very successful launch. Since HN isn't your target audience having your site fail from the traffic was far better than having it fail from a launch in your market. You shook out some important bugs before you lost real users. Plus you got to do this followup which will bring even more traffic.
- nasalgoat 14y agoI find it very difficult to believe that this person worked on any sort of performance team, given that what they discovered is pretty much "Handling Load 101". Running everything on one box? Using swap? No caching? It's like a laundry list of junior admin mistakes.
- druiid 14y agoGenerally that's kind of how these smaller 'startups' work... and then I get a call and charge my standard rate per hour ;)
- zorlem 14y agoA few things have caught my attention in your post. Your biggest problem was that the configuration of your services was not sized/tuned properly for the hardware resources you've got. As a result of this your servers have become unresponsive and instead of fixing the problem, you've had to wait 30+ minutes until the servers recovered. In your case you should have limited Solr's JVM memory size to the amount of RAM that your server can actually allocate to it (check your heap settings and possibly the PermGen space allocation). If all services are sized properly, under no circumstance should your server become completely unresponsive, only the overloaded services would be affected. This would allow you or your System Administrator to login and fix the root-cause, instead of having to wait 30+ minutes for the server to recover or be rebooted. In the end it will allow you to react and interact with the systems. The basic principle is that your production servers should never swap (that's why setting vm.swappines=0 sysctl is very important). The moment your services start swapping your performance will suffer so much that your server will not be able to handle any of the requests and they will keep piling up until a total meltdown. In your case OOM killing the java process actually saved you by allowing you to login to the server. I wouldn't consider setting the OOM reaction to "panic" a good approach - if there is a similar problem and you reboot the server, you will have no idea what caused the memory usage to grow in the first place.
- ajross 14y agoI'd say that the biggest problem is that they tried to launch their product on what appears[1] to be a 4G host, representing maybe $3-400 of hardware cost (maybe more if you buy premium, I doubt Linode does). I mean, careful configuration and capacity planning is important. But what happened to straightforward conservative hardware purchasing where you get a much bigger system than you think you need? It's not like bigger hosts are that expensive: splurge for an EC2 2XL ($30/day I think) or three for the week you launch and have a simple plan in place for bringing up more to handle bursty load. [1] The OOM killer picked a 2.7G Java process to kill. It usually picks the biggest thing available, so I'm guessing at 4G total.
- foxylad 14y agoAppengine. You're a development shop, not scalable system builders. Deciding to build your own systems has already potentially cost you the success of this product - I doubt you'll get a second chance on HN now. If you were on appengine, you'd be popping champagne corks instead of blood vessels, and capitalising on the momentum instead of writing a sad post-mortem. I'd recommend you put away all the Solr, Apache, Nginx an varnish manuals you were planning to study for the next month, and check out appengine. Get Google's finest to run your platform for you, and concentrate on what you do best.
- srameshc 14y agowell, you got it right this time :)