6 ms·
Fighting Back After Hacker News Took Down My Site
- js2 15y agohttp://www.google.com/search?q=site:news.ycombinator.com+patio11+keepalive http://www.google.com/search?q=site:news.ycombinator.com+pat... I will also quote myself: "I first learned to disable Apache KeepAlives in 1998. Yes, 1998. It's disheartening that Apache still ships with it enabled by default. It has always allowed a relatively small number of lingering clients to completely DoS your server."
- tluyben2 15y agoYeah I hear tons of people saying 'check if KeepAlive if on, your site will be much faster if it is'. Not the smartest general remark about configuration of web servers to make :)
- spydum 15y agoThere is nothing wrong with keepalives if you give them proper timeout and request limit values. They definitely cut down tcp/ip setup overhead, and improve page render times. Just don't let the client be too greedy about how long they hold that connection.
- pornel 15y agoIt's not problem of web servers in general, but specifically Apache. HTTP keep-alive connections are faster (saves 3-way handshake on subsequent requests and allows pipelining). The problem is that Apache implements them in a horribly inefficient manner (thread/process and all its memory kept in use just to hold on to a socket). If you use nginx, lighttpd or shield Apache with haproxy/varnish, then you can easily have keep-alive enabled and clients will see better performance.
- wheels 15y agoThis is the most critical bit. We run our website (totally disconnected from our webservices, naturally) on a 256 MB VPS and frequently handle top HN stories on our wordpress blog (which runs on a host with several of our other PHP / static subsites). Key points: • Make sure that wordpress supercaching is on. You can verify this by looking at the last line of the HTML that comes back which has a timestamp for when it was generated. • Turn off KeepAlives. • Set MaxClients to 8. • Use monit to check to make sure it can connect and restart apache with a kill -9 if it can't. (This is optional, but helps if you have some random thing that ends up taking a very long time to execute and eats up connections.) With that in place even a much smaller server can easily handle a HN top story without breaking a sweat.
- rellik 15y agoDo you use something automated to ensure things are cached? I use a different page caching plugin (W3 Total Cache), but it also puts the comment in the HTML.
- wheels 15y agoNo. Once I've seen that it's working properly at one point, I've never seen a case where SuperCache stopped working later (unless I changed the config).
- peterwwillis 15y ago> Set MaxClients to 8. LOL what? what's your mpm? did you enable this to keep the backend from blowing up from too many queries? surely you can handle more than 8 connections at a time. is this the proxy layer, and if so are you using web caching on top of wordpress caching? if monit is restarting apache every time it can't connect (i hope you have a long timeout) you're denying service to a lot of people. connections are supposed to queue so they don't get dropped.
- wheels 15y agoAs mentioned, this is on a single VPS with 256 MB of RAM. Each Apache process needs about 25 MB of (non-shared) RAM, so actually 8 is pushing it. We're using mpm_prefork. There's no additional proxy nor cache. My point, specifically, was how low you can go with a cheapo VPS. We hold up fine during an HN spike. (We're B2B and not a destination site, so our usual load is trivial.) Even during an HN spike you're getting tops of 2-3 visitors per second, which can be dished out reasonably well with 8 workers. The monit thing kicks in after a 30 second timeout. With the configuration above, we don't get that because of load, but rather when something else has gone wrong (specifically there's a wordpress plugin that our internal status blog uses that sometimes hangs). But given the original poster's issue of apache getting so out of control that it took him several minutes to get a live SSH connection and a system load of 60, having monit kill things (and restart them) is a preferable stop-gap. (Note: Our actual customer facing stuff is quite different; there we're using multiple servers behind an nginx proxy and using a combination of Rails, Sinatra and Java services. The basic web stuff is segregated off from those primarily for security reasons.)
- rs 15y agoNginx has proved extremely useful and capable especially in memory constrained virtual environment. It effectively gives an easy way to move from a synchronous web server to an asynchronous one.
- teadrinker 15y agoBETTER SOLUTION: Get on the cloud. 15 minutes to upgrade your server when load gets high and billed by the hour at the ram level you're 'using. Theres very little reason to stay on traditional hosting if you're running anything close to a serious startup.
- rellik 15y agoWhile the cloud does enable you to scale your resources on demand, you still have to address the same issues I had to deal with (how to efficiently use those resources). Plus, moving to a mutli-server setup involves additional complexity in the architecture of the system, which I'd rather not mess with unless I have to :)
- teadrinker 15y agoMulti-server? Nope, single. Same as you have now, just with added flexibility on ram with 2 clicks I can handle a huge server load with ease, then drop back when done. That way I don't waste time on low level optimization tasks and can keep to what I do best. Building things. Seriously, there's no reason to stay on a traditional style server setup and you'll never go back once you've tried it. It can be pricey once you ramp up, but give it a shot as you can literally pay by the hour while you play.
- spydum 15y agoLow-level optimization can be fun though. If you always have the convenience of "throwing hardware" at the solution, you may not be forced into learning something cool, like nginx or properly configuring caching.
- rellik 15y agoHmm, didn't know you could throw ram at it without adding servers/bouncing the box. Very cool :) Still, I think there's value in figuring out this stuff by investigating efficiency gains. If I was having to do real tricky stuff (which would probably be beyond my expertise) then the cloud would make more sense for me.
- sc68cal 15y agoWhat about isolating the popular URL and temporarily serving a static version of the page?
- rellik 15y agoThat's what the page caching plugin does. It puts a static file in nginx/apache's search path, so it never hits PHP
- sc68cal 15y agoWhich plugin did you use? WP Super Cache? I'd like to take a look through the code.
- spydum 15y agototal cache, super cache, most of them create static html versions of the page. Super Cache is a bit more clever in that it depends on htaccess instead of PHP code to check to see if the on-disk cached version exists. The benefit here is PHP module does not need to fire up -- apache can serve this very rapidly with little memory footprint.
- rellik 15y agoI used W3 Total Cache. It generates a static file and puts it in apache's search path, so it never hits PHP
- bradleyland 15y agoI cringed a little bit when I read this: "...none get enough traffic for me to have made caching or performance tuning a huge priority". I don't mean to level criticism solely against the post author for this, because he is not alone, but I find this mentality inexcusable. This is especially true if you're using WordPress, for which there are a multitude of caching plug-ins. All of which will offer orders of magnitude better performance than allowing the blogging engine to build the page from scratch every time. In other words, caching is always a priority. It's not an add-on or an after-thought, it should be part of your design. Consider a scenario where someone asks you to add up three arbitrary numbers: 123 + 456 + 789 Now, imagine you type those in to a calculator to add them up. It takes you a few seconds. The output is 1368. You can easily remember this number, so the next time someone asks you what the result of 123 + 456 + 789 is, you can just say 1368. Not caching is like keying the numbers in to a calculator every time, rather than just relying on your memory. I know this is a rudimentary example, and I know that most people running a WP blog probably know how caching works, but even if you don't "need" it today, why would you leave your blog set up so that it's constantly re-building content that can be cached by simply installing a WP plug-in? I implore you. Make caching the second thing you do (after security) when setting up or building your web app.
- rellik 15y agoDefinitely agree! For me, the counter-point is that there's tremendous value in getting the Minimal Viable Product out the door, and iteratively improving it (the post in question was only the 2nd post after moving off a hosted blog). Projects that I want to be perfect before release usually don't make it out the door. Plus, I'd grown lazy from my lack of traffic :)
- benmills 15y agoI agree with you when you're taking about web apps, but what about personal blogs? If I just want to set up a personal blog that I don't plan on promoting or spending a lot of time on I wouldn't want to spend any extra time setting up things that don't help me with my primary goal, writing blog posts.
- 15y ago
- ck2 15y agoWho run WordPress without a page cache these days? Just look at the query count even without plugins. Add a few plugins and it's a total mess.
- sciurus 15y agoAn important point that was missed is how was he running the php and rails apps? mod_php? mod_rack? cgi? fastcgi? Proxying to an application server? The answer to that is going to greatly affect how well Apache scales when you're memory-constrained.
- anigbrowl 15y agoWhat a strange headline. Who is he fighting back against - Hacker News? Traitorous Wordpress? Himself? If I needed a plumber I wouldn't want one that talked about fighting back against the water. Nor would I want an architect that talked of fighting back against gravity. The language of confrontation is inappropriate here, since the problem as described stems from a failure to spend time learning or configuring Apache. This might seem trivial, but to me it suggests fundamental flaw in the approach to the problem, which probably increased the time needed to fix it.