6 ms·
Lwan: Experimental, scalable, high performance HTTP server
- imaginenore 12y agoNginx can do 500K to 1 million req/s http://lowlatencyweb.wordpress.com/2012/03/26/500000-requestssec-piffle-1000000-is-better/ http://lowlatencyweb.wordpress.com/2012/03/26/500000-request... So can Google's compute engine on a $10 instance: https://news.ycombinator.com/item?id=6804897 https://news.ycombinator.com/item?id=6804897 http://googlecloudplatform.blogspot.ca/2013/11/compute-engine-load-balancing-hits-1-million-requests-per-second.html http://googlecloudplatform.blogspot.ca/2013/11/compute-engin...
- acidx 12y agoPerformance isn't the only thing that you should look for in a web server. Nginx is probably the best choice for most applications, yes, as Lwan lacks lots of important features, real world testing, and community. And I say that having written Lwan: it is, for me, nothing but a toy. :) OTOH, the beefiest machine I have access to test it is a 4 year old laptop, not a 24-core Xeon.
- bhauer 12y agoThe tests at lowlatencyweb.wordpress.com were conducted without network connectivity—the load generator (wrk) was running on the same host as the web server. The results are 500K RPS for localhost connections with standard keep-alive and 1M RPS for localhost connections with pipelining. This is using a server with 24 HT cores and it's not clear to me what the response body was. Google's Compute Engine test was using 200 virtual servers, but it does include network connectivity. The response body is a single byte. Their blog entry is a celebration of the performance of their load balancer more than a statement about the performance of each VM. In March, we were able to exceed 1M requests with network connectivity and without pipelining to a single server [1]. Our project is not testing static web servers, so we don't test with plain nginx; but I expect nginx would also exceed 1M RPS in this hardware environment. This was using a server with 40 HT cores and a single-byte response body. Similarly, a highly tuned web server such as OP's Lwan should be expected to exceed 1M RPS (network-connected) on a 40 HT core server. 1M RPS with small response payloads is fairly easy on modern hardware. Incidentally, we see 6M+ RPS with pipelining in our Round 9 plaintext results [2]. [1] http://www.techempower.com/blog/2014/03/04/one-million-http-rps-without-load-balancing-is-easy/ http://www.techempower.com/blog/2014/03/04/one-million-http-... [2] http://www.techempower.com/benchmarks/#section=data-r9&hw=peak&test=plaintext http://www.techempower.com/benchmarks/#section=data-r9&hw=pe...
- imaginenore 12y agoIf you want to go further, you need to get rid of the OS: 40 million req/sec with Lua http://highscalability.com/blog/2014/2/13/snabb-switch-skip-the-os-and-get-40-million-requests-per-sec.html http://highscalability.com/blog/2014/2/13/snabb-switch-skip-...
- justincormack 12y agoThats not a web server, its doing packet processing, which is a different problem. You could connect a web server to it, but that is not a benchmark of that.
- nwmcsween 12y agoNice, benchmarks vs the competition would be interesting as well. Also with 10K+ idle connections I would be more worried about kernelspace memory requirements (maybe recommended sysctl.conf changes).
- donavanm 12y agoNah, its a couple KB per connection. the biggest consumer would be the tcp socket control structs and associated data buffers. Ball park 1.5KB for the structs and another 4-16KB for tcp buffers on a typical internet tcp connection.
- nwmcsween 12y agobut vars controlling how long a tcp sock is held for or if they are reused is controlled my the kernel.
- donavanm 12y agoI think youre talking about timewait states et al. On linux thats dominated by the MSL, which is a compile time constant of 60 seconds. You mentioned sysctls, those are primarily tcp_tw_reuse and tcp_tw_recycle (which is the worlds worst sysctl). Regardless, its a couple KB per connection. How many hundred thousand do you want to support?
- hendzen 12y ago"Hand crafted HTTP request parser" - hard to see how this really faster (and less bug prone) than generating one with Ragel.
- meowface 12y agoI am fairly ignorant of Ragel, but to my understanding Ragel is better for assuring correctness rather than performance. I don't think Ragel makes any claims regarding performance. I could see a specially written parser outperforming it, just like hand written assembly can still sometimes outperform a compiler. I agree that it's more likely to have bugs, though.
- nly 12y agoMuch of the advantage of using a DFA generator like Ragel is lost because the HTTP header grammar is actually ambiguous in several places and can't be streamed. You can use it as a component, but it isn't entirely sufficient on its own. The HTTP 1.1 RFC requires whitespace be stripped at the end of header values, yet also permits (although deprecates) header folding, giving rise to the following ambiguity (using _'s in place of leading spaces): Foo: Hello\r\n ______\r\n ________\r\n ________ world!\r\n Bar: smeg This requires that the parser buffer all the whitespace between 'Hello' and 'world!' (and the RFC doesn't put a standard limit on header value length) just in case 'world!' never comes and the value of the Foo header has to be stripped back to just "Hello" Here's a related observation from a commit[0] by the Joyent guys, who wrote the streaming parser used by NodeJS: "For http-parser itself to confirm[sic] exactly would involve significant changes in order to synthesize replacement SP octets. Such changes are unlikely to be worth it to support what is an obscure and deprecated feature" Another example is parsing the Request and Status lines: GET <uri> HTTP/1.1 Technically <uri> can't contain spaces, but the RFC says you MAY accept them [RFC7230: 3.5. Message Parsing Robustness] ... which then gives rise to the possibility of <uri> containing the literal string " HTTP/1.1", and ultimately opens up bad user agents that send spaces to header injection. Resolving these ambiguities require implementing your own buffering, and dropping down to Ragels 'state charts' feature to avoid your semantic actions being munged... which leaves you to design the top level state machine yourself. [0] https://github.com/joyent/http-parser/commit/5d9c3821729b194eef60f41fcc5f8b4657c3d8ff https://github.com/joyent/http-parser/commit/5d9c3821729b194...
- djcapelis 12y ago> "Hand-crafted HTTP request parser" Uh oh.
- codingbeer 12y agoFor more information on some of the C magic behind this well-written piece of software, check out the authors blog[1]. It should be pretty interesting for any systems programmer. [1]: http://tia.mat.br/blog/html/index.html http://tia.mat.br/blog/html/index.html
- e12e 12y agoAny relationship to G-WAN[g]? I can't see any mention of it in the readme, or on the web page (perhaps I didn't look hard enough) -- but there appears to be some resemblance? (Use as a C web framework work-a-like, fast webserver etc)? [g]http://gwan.com/ http://gwan.com/
- acidx 12y agoApart from sharing the "wan" suffix and choice of language, there's no relationship whatsoever.
- nwmcsween 12y agoA few issues: * rawmemchr - it might be faster as it doesn't have to decrement the size_t but this only mildly relevant for lwan as many instances of rawmemchr use are simply rawmemchr(ptr, '\0') which is exactly the same as ptr + strlen(ptr) + 1 and even less optimized. * pthread_tryjoin_np - __linux__ is defined by gcc not glibc, you should check for __GLIBC__ if you want to use glibc specific functions. * underscore prefixed functions - pedantic I know but it is reserved for the implementation.
- acidx 12y agoThese are easy things to fix: feel free to issue a few pull requests. :) Regarding rawmemchr(): both are pretty well optimized. Both are implemented in glibc using the same technique (reading a byte at a time until it is aligned, then moving to multibyte reads). strlen() might be faster, yes, considering that the implementation can hardcode some magic numbers. In other words: some micro benchmarks might help decide here. Regarding __linux__ vs. __GLIBC__: Lwan works with some alternative libcs (such as uClibc), so relying on __GLIBC__ being defined for things like this doesn't seem like a good idea. In any case, since Lwan isn't portable anyway, one can just assume it is always running on Linux and get rid of these #ifdefs.
- abionic 12y agoCan't spot an OSS license. What license is aimed for it?
- acidx 12y agoGPLv2 (or later) at the moment. Might change to LGPLv2 (or later) soon, though.
- jagger27 12y agoAny roadmap for HTTP2 support?
- acidx 12y agoNot planned ATM. Still need to read more about it before giving it a go.
- zongitsrinzler 12y agoHow does this compare to Nginx?
- sauere 12y agoPretty good, i'd say about 50% faster on raw request speed. Anyway, it isn't a fair comparison given that nginx's feature set is much larger.
- mmastrac 12y agoI had to Google "rebimboca da parafuseta" out of curiosity. Apparently it's a Brazilian term roughly analogous to "reticulation of the splines".
- acidx 12y agoAuthor here. This gave me a chuckle. :) And, yes, that's pretty much it, although "rebimboca da parafuseta" is less obscure than "reticulation of the splines" (if you're Brazilian, anyway). It is used to denote a fictitious part, which name or function are unknown, of a car engine or any other machine. It was coined in the 70s TV show, and later used in some ads aired during the same period; since then, it's an expression used for humorous effect. It's not very common these days but is unlikely you'll meet someone down here that never heard it.
- kaoD 12y agoWe've got a similar idiom in Spain: "la junta de la trócola" ("trócola's joint", not really translatable, though trócola is apparently a synonym for "polea", "pulley"). It was coined for a cigar brand TV commercial back in the 90s and also alludes to a fictional part, in this case used by a car repairman to fraud a customer into paying more in said commercial.
- mutagen 12y agoAha, not unlike the retro-encabulator: https://www.youtube.com/watch?v=RXJKdh1KZ0w https://www.youtube.com/watch?v=RXJKdh1KZ0w
- techdragon 12y agoLove that video, I may have shared it with more people than I did the original rick roll video lol
- emmelaich 12y agoWhich is of course a riff on the original turbo-encabulator: https://www.youtube.com/watch?v=Ac7G7xOG2Ag https://www.youtube.com/watch?v=Ac7G7xOG2Ag