21 ms·
>Nginx has taken over some of Apache's market because it's lighter on resources, i.e. with the same amount of resources it can process & deliver more stuff. Th
by rehack 13y ago
>Nginx has taken over some of Apache's market because it's lighter on resources, i.e. with the same amount of resources it can process & deliver more stuff.
That's precisely why we moved to Nginx. We do mostly Java so Apache was not the problem there, as it was simply proxying over to Jettys.
But the blog is wordpress that means PHP, so we had to use modphp modules with Apache. That was Okay initially, but as we scaled, we saw some 50 instances of Apache and each taking 25 Mb memory. And note, most of the bloat in memory was due to modphp being loaded in-memory in each instance.
We moved to Nginx with an external php-fpm for PHP, and just 4 Nginx instances take up the entire load and each uses some 2.5 Mb of memory.
And at the same time Google webmaster shows a significant drop in average latency.
So that's our experience of why we moved to Nginx.
Edit: typo
- thwarted 13y agoSo you've moved the memory usage from the web server proper to the external php-fpm processes, which are most likely getting recycled more frequently than the default settings for apache prefork workers are so you're not seeing the ballooning PHP memory usage. The size of the 50 instances of Apache each taking 25 meg is most likely because of the PHP code and mod_php, not because of Apache itself.
- rehack 13y agoAgreed. In hindsight, I am sure, we could have externalized the PHP with Apache as well. But the larger context of this move was doing more with less. And Nginx just working with 4 instances and serving just as much, with a reduced average latency of page load, justified the move. And I sort of felt sad doing this. Because I have a huge respect for Apache. I hope just like what Chrome did to Firefox, Nginx does to Apache, and its a win-win for all.
- thwarted 13y agoBut the larger context of this move was doing more with less. And Nginx just working with 4 instances and serving just as much, with a reduced average latency of page load, justified the move. "More with less" here is a judgement call. You've increased the number of processes and subsystems that need to be managed (nginx + php_fpm), which is more. You've increased the context switching when dynamic requests need to be passed to another process and proxied through nginx. The communication between the frontend web server and php_fpm needs to be serialized in some fashion, so you're now incurring a greater than 1x cost in processing and parsing things like HTTP headers and the request (serializing and deserializing the data to pass that between the processes). It seems you're ignoring the actual cost of running dynamic PHP code because the cost has been moved out of the "web server", the thing that listens on port 80, and into something that isn't considered "a web server", but this is just a change in bookkeeping. Do you know where that reduced average latency is coming from? Is it because nginx is serving static assets that would tie up an apache worker that could be running/serving PHP in apache? If so, you can do the same thing by offloading all your static asset serving to another server (or a completely external CDN, which has other extremely beneficial client-side advantages too) using apache too. That is, the choice of software here isn't as big a win as presented, it's the architecture of the ecosystem that is the big influence. It's great that you found something that works better for your workload, and it's appropriate to use what works and not get bogged down in why all the time. But the way you've presented it doesn't indicate well that all the causes and effects are completely understood. And I think this ends up giving Apache a bad rap.
- rehack 13y ago>Do you know where that reduced average latency is coming from? Is it because nginx is serving static assets that would tie up an apache worker that could be running/serving PHP in apache? If so, you can do the same thing by offloading all your static asset serving to another server (or a completely external CDN, which has other extremely beneficial client-side advantages too) using apache too. That is, the choice of software here isn't as big a win as presented, it's the architecture of the ecosystem that is the big influence. Yes, I understand that. I have studied about the C10K problem[1]. And have coded another service which just uses libevent[2] directly and some C++ code to serve a feature for our site very low latency. So I understand quite well why Nginx is offering low latency. As you rightly observe later on, I don't have the luxury to get obsessed with all the Whys, so often shoot for the major architectural gain and take any side effects that come along in the stride. >.. But the way you've presented it doesn't indicate well that all the causes and effects are completely understood. And I think this ends up giving Apache a bad rap. I am surprised by your this observation. I have been 100% honest in what I wrote above, and I repeat have a huge respect and thankful for the Apache team and Apache web server software. Just that, by the page loading gains we got, were clearly very good. And also there were less errors in general reported in the web master. So we stuck with the change. Also a point regarding the switching costs and extra processing between php-fpm and Nginx processes. As I said in my first comment, we do mainly Java and PHP is just for the blog. But I very clearly remember seeing all the Apache child procs bloating to 25 Mb after the first hit to the PHP code was done. So I am actually saving by externalizing on the resources. Trust me I know what I am doing. Another tangential thing not related to Nginx, which I am doing for low resource consumption. Is moving some relevant code from Java to Go. And there also am seeing huge memory gains (i.e. savings). Perhaps will share more about it at an appropriate time. [1] http://www.kegel.com/c10k.html http://www.kegel.com/c10k.html [2] libevent.org Edit: grammar
- thwarted 13y agoI am surprised by your this observation. I have been 100% honest in what I wrote above... Trust me I know what I am doing. Sorry, I didn't mean to suggest that you weren't being honest or that you didn't know what you were doing. I'm sure we've all had to deal with "Well, this guy on HN converted to X from Y and saw <insert generic gains claims here>, so why are we on Y again?", which the somewhat hand-wavy details in your original comment can end up contributing to (obviously we're limited for space and attention here, so leaving out some details can be desirable). What I failed at was communicating that my comments were not intended as an attack on your specific methodologies or choices, but was more meant to make the reader of our comments consider the wider implications of blindly following the herd and not making informed decisions. I have had people forward me lists of links to isolated, comment-less HN comments as "support" for their position. But I very clearly remember seeing all the Apache child procs bloating to 25 Mb after the first hit to the PHP code was done. So I am actually saving by externalizing on the resources. Yes, sorry, I didn't consider the exact traffic ratios between the (proxied) Java requests and the (internally handled) PHP requests. If you've got a wide spread there favoring proxied requests, then, as you've experienced, it would be advantageous to get rid of the internal handling, proxy all requests, and outsource the bloating to other processes/systems where it is more isolated, and the upper bound on the bloat be more influenced by the lower total requests for the resources that consume more. And once that's done (the getting rid of the monolithic parts) then it's a lot easier to experiment with replacing the different parts to see if there are other gains. You may have seen similar resource consumption benefits by nginx proxying to Apache+mod_php or even to bare PHP cgi. But one shouldn't even follow that method blindly either, because it's all based on workload, and everyone's workload is different. There are so system management and monitoring benefits and drawbacks that go with these kinds of changes. We both know that there is no software selection silver bullet, but it's not necessarily the readers of these comments that know that.
- xist 13y agoI hear this a lot about Apache and there's one thing I'm not sure a lot of people understand... it's SHARED memory in most cases. While it spawns several processes (depending on which worker method), it's not as awful as it looks at first glance. A quick google search finds a nitty gritty page that might help - http://foertsch.name/ModPerl-Tricks/Measuring-memory-consumption/index.shtml http://foertsch.name/ModPerl-Tricks/Measuring-memory-consump... there's probably better examples. Yes nginx does have lower memory usage, but on most non vps/shared servers this is a lot less of a concern than most people think vs Apache. php-fm works great with Apache as well and probably accounts for the huge resource gain over mod_php that Apache uses by default.
- thwarted 13y agoI hear this a lot about Apache and there's one thing I'm not sure a lot of people understand... it's SHARED memory in most cases. While it spawns several processes (depending on which worker method), it's not as awful as it looks at first glance. Apache, and specifically mod_php, doesn't have that level of control over memory usage (mod_perl (and, incidentally, mod_python) has deeper integration with Apache to encourage, with the right configuration, more of the memory to be shared between processes, but it's still not that great). The Apache parent process, primarily exists to do process, signal, and socket management, it's not really possible to do a lot of application level (pre)processing before subprocesses are forked off, which is what would be required to share a significant portion of the address space. If you have a .php file run via mod_php that looks like this: $x=""; for ($a=0;$a<(1024*1024*100);$a++) { $x.="1"; } This will produce 100MB of non-shared memory in a child process, when the request is made and serviced by that child. And, because of the way requests are dispatched to children, at a low request rate the same child could end up serving all (or a majority of) requests. This unbalances the memory usage between processes. A quick google search finds a nitty gritty page that might help - http://foertsch.name/ModPerl-Tricks/Measuring-memory-consump.. http://foertsch.name/ModPerl-Tricks/Measuring-memory-consump.... there's probably better examples. That is a great link in general for how shared memory and COW works.