19 ms·
Asmttpd – Web server for Linux written in amd64 assembly
- voidiac 11y agoThe name is confusing, the first thing i thought of was an SMTP-server.
- pavlov 11y agoIt's been closer to 20 years since I last read a complete program in x86 assembly, so this is quite fun to look at. I'm somehow disappointed (quite unreasonably, of course) that the code uses plain old zero-terminated C strings instead of something more exotic. One of the fun things about assembly is that you get to reinvent basic language features on the fly -- calling conventions, data layout, strings, everything.
- JoshTriplett 11y agoIt needs to do so to interoperate with the OS, so using those avoids having multiple conventions and converting between them.
- eddyb 11y agoAFAIK Linux doesn't use zero-terminated strings anywhere in its syscalls, or at least not in those like write, where you pass a size alongside the buffer.
- reipahb 11y agoHowever (almost?) all syscalls dealing with filesystem paths take null-terminated strings. See for example the implementation of the open() syscall: https://github.com/torvalds/linux/blob/fb65d872d7a8dc629837a49513911d0281577bfd/fs/open.c#L999 https://github.com/torvalds/linux/blob/fb65d872d7a8dc629837a... (Hence the need for the strncpy_from_user()-function: https://github.com/torvalds/linux/blob/fb65d872d7a8dc629837a49513911d0281577bfd/lib/strncpy_from_user.c#L99 https://github.com/torvalds/linux/blob/fb65d872d7a8dc629837a...)
- eddyb 11y agoAhh, I totally missed the path arguments, my bad.
- ranit 11y agowrite() syscall writes a sequence of bytes (not string) therefore cannot use zero-terminated convention.
- JoshTriplett 11y agoopen takes a NUL-terminated filename.
- JoshTripIsWrong 11y agoNo it doesn't? Networking syscalls generally take length-specified strings. Unless this calls into libc or something...
- dang 11y agoPlease don't make usernames that attack another user. It's uncivil and distracts from the topic.
- TheLoneWolfling 11y agoI'm surprised it doesn't do length-prefixed strings with a null terminator anyways. Makes a whole lot of things easier.
- KMag 11y agoAs a general principle, bugs are reduced when you use representations that don't allow inconsistent representations. Unless you have an overriding reason, it's best to use a data representation that doesn't allow non-canonical representations (if such exists). If you null-terminate a length-prefixed string, what if there's a null in the middle of the string? (1) Allow inconsistency, and go with the length prefix in case of inconsistency. You could allow null bytes in the middle, treating it as a normal length-prefixed string, but then why do you null-terminate the string? (Is it so that you can still pass the string to functions that will choke on embedded nulls? Why would you do that?) This is just asking for kernel bugs. (2) Allow inconsistency and go with the position of the first null byte in the case of inconsistency. If the length prefix is inconsistent with the position of the first null byte, you could go with the position of the first null byte, but then why even have the length prefix? (3) Disallow inconsistency. You could disallow embedded nulls, but then the length prefix is just there as a place to cache strlen calls? If you're defining a syscall interface and requiring the length and first null to be consistent, then you need to run strlen anyway in order to sanity check what userspace gave you... why not simplify the external interface to just be either null-terminated or length-prefixed?
- TheLoneWolfling 11y agoI should have specified a little more, perhaps. Don't think of it as a null-terminated + length-prefixed string. It's effectively a purely length-prefixed string. There just happens to always be a null one byte after the end of the string. The length prefix is the only thing you use, ordinarily. The only time the null comes into play is if you've already had a bug. Think of it like a stack protector.
- KMag 11y ago> Think of it like a stack protector. ... but one that doesn't terminate execution, but instead hides your bugs. In most use cases, I'd prefer to find my bugs in the majority of cases, rather than to hide the bugs except for corner cases.
- maguirre 11y agoOut of curiosity What would you have done for strings?
- FreeFull 11y agoCould have a pointer, length pair. That's how it's done in some non-C languages.
- bkeroack 11y agoPointer, length, encoding. In 2015, giving me a bag of bytes labeled "string" is about as meaningful as saying it's "music" or "a picture".
- vardump 11y agoIt's 2015. UTF-8 won.
- kinghajj 11y agoIf only Java and NT would get the message...
- danbruc 11y agoUTF-8 is not well suited for a general purpose string implementation because it is a variable length encoding and therefore addressing a character becomes a linear time operation. UTF-16 would probably be a better choice in most cases.
- pjc50 11y agoIt needs to be security-hole compatible.
- lelf 11y agoMaybe it's just me, but I honestly don't know what this is doing near the HN top. It's more or less literal translation from C.
- kbenson 11y agoPossibly because it illustrates that using assembly isn't necessarily the insane scary idea that it first seems to many (like me), even to those that should know better because they have used it in the past (like me).
- bjourne 11y agoNo, a literal translation is what you get when you write a http server in C and inspect what assembly code it produces for x86.64. Since this assembly code is nowhere close to that output it is not a literal translation.
- rdc12 11y agoOnly if you use a completley naive compiler, any level of optimisation moves from being a literal translation
- bjourne 11y agoNo, any correct translation a c compiler produces is, by definition, a literal translation.
- rdc12 11y agoNo, that simply means that they are semantically equivilant, which is very different to a literal translation. To quote an online dictionary [0], "2. Word for word; verbatim: a literal translation.". Optimisation in compilers is certainly not word for word. Take a simple problem like the FizzBuzz problem, write it as the simple obvious branching style. Now compile it with GCC or Clang (with -O3) and you end up with lookup table (or at least I did a few months back. Semantically equivalent but not literal "word for word" translation. [0] http://www.thefreedictionary.com/literal http://www.thefreedictionary.com/literal
- markus2012 11y agobenchmarks? :-)
- edsiper2 11y agosomething like "wrk -d 10 -c 1000 -t 4 http://127.0.0.1/byte.txt http://127.0.0.1/byte.txt (one byte file) did: Requests/sec: 100.00 Transfer/sec: 11.91KB for a JPEG image (200KB) the results are similar: Requests/sec: 99.86 Transfer/sec: 19.27MB
- nXqd 11y agothis is really really fast.
- edsiper2 11y agoare u kidding ? expectation is no less than a 100k req/s.
- trentnelson 11y agoDoesn't look like it supports HTTP pipelining: [trent@ubuntu/ttypts/4(~s/wrk)%] ./wrk -c 1 -t 1 --latency -d 5 http://localhost:8080/Makefile Running 5s test @ http://localhost:8080/Makefile 1 threads and 1 connections Thread Stats Avg Stdev Max +/- Stdev Latency 35.00us 0.00us 35.00us 100.00% Req/Sec 10.00 0.00 10.00 100.00% Latency Distribution 50% 35.00us 75% 35.00us 90% 35.00us 99% 35.00us 1 requests in 5.10s, 1.57KB read Requests/sec: 0.20 Transfer/sec: 314.55B Note the 1 request. For 1000 clients, it's only doing 1000 requests: [trent@ubuntu/ttypts/4(~s/wrk)%] ./wrk -c 1000 -t 1 --latency -d 5 http://localhost:8080/Makefile Running 5s test @ http://localhost:8080/Makefile 1 threads and 1000 connections Thread Stats Avg Stdev Max +/- Stdev Latency 307.93ms 552.88ms 1.63s 87.10% Req/Sec 414.29 439.07 1.34k 85.71% Latency Distribution 50% 4.11ms 75% 407.63ms 90% 1.63s 99% 1.63s 1000 requests in 5.01s, 1.53MB read Socket errors: connect 0, read 41, write 0, timeout 0 Requests/sec: 199.79 Transfer/sec: 312.96KB
- kryptiskt 11y agoI propose that a web framework be called Assembly on Ambulator.
- lsiebert 11y agoGiven that it's so small, 6k, I'd called the framework based on this Assembly on Alleys. It could totally implement a DSL... call it "C" for convenience, that generates the required assembly code :-).
- wyc 11y agoShameless plug for a companion IRC bot in ARM assembly: https://github.com/wyc/armbot https://github.com/wyc/armbot
- mariopt 11y agoWhy would someone write a web server in assembly? just for fun?
- octo_t 11y agoExactly. Because why not?
- fleitz 11y agoProbably to lower overhead associated with C language features, similar to the reason why many people write things in C instead of a higher level language.
- stass 11y agoTo avoid overhead people implement specialized compilers suited for the task at hand. If anything, going down to C, and especially assembly will hurt performance as low-level code is much harder to optimize for obvious reasons. Above all of that real-world performance comes from proper system-level design, not micro-optimizations, and using a low-level language (be it C, C++ or assembly) will prevent one from quickly iterating over different ideas.
- kibibu 11y ago> low-level code is much harder to optimize for obvious reasons What are the obvious reasons I'm missing?
- cgabios 11y agoLow-level code often omits high-level intended behavior (description, pseduocode, documentation, test cases, etc.) and semantic meaning like variable names. In such styled codebases, the absence of these makes it harder to refactor, reuse and/or modify than say concisely & precisely documented codebases in higher-level languages (Python, Ruby, Go) or quality asm.
- 11y ago
- ndesaulniers 11y agoI'll try and get this building for OSX. For the uninitiated, might I recommend my: http://nickdesaulniers.github.io/blog/2014/04/18/lets-write-some-x86-64/ http://nickdesaulniers.github.io/blog/2014/04/18/lets-write-... Though, this is written in yasm syntax, which is slightly different. Also, keep an eye out for a blog post on Interpreters, Compilers, and JITs I'm working on (cleaning it up and getting it peer reviewed this or next week)! update 1 Actually, would the syscall's be different between Linux and OSX? Let's find out, once this builds! hammers away update 2 Got it building and linking. bus error when run, debugging with gdb. update3 Can't generate dwarf2 debug symbols for OSX? $ yasm -g dwarf2 update 4 Careful, this tries to listen on port 80 [0] (0x5000 (LE) == 5*16^1 == 80), I would never run any assembly program off the web with elevated privileges. I recommend 0xB8B0 (LE, port 3000). update 5 > Actually, would the syscall's be different between Linux and OSX? Looks like yes: http://unix.stackexchange.com/a/3350 http://unix.stackexchange.com/a/3350 These might be close to shim out (OSX and Linux at least share a calling convention, unlink Windows). I'll upstream what I have. [0] https://github.com/nemasu/asmttpd/blob/master/main.asm#L24 https://github.com/nemasu/asmttpd/blob/master/main.asm#L24
- j42 11y agoFreudian slip? unlink(Windows) indeed...
- yellowapple 11y agoRelevant: http://i.imgur.com/INBvStO.png http://i.imgur.com/INBvStO.png
- e12e 11y agoA little strange to see: "Sendfile can hang if GET is cancelled." in the readme and no corresponding issue. Not even one closed as "wontfix". Sounds like DOS?
- barosl 11y agoDays ago I saw a lightweight httpd written in C here, yesterday a C++ header-only httpd library caught my mind, and now an httpd in assembly. I'm curious what would come next...
- vardump 11y agoOh, I bet someone writes httpd in vhdl or verilog... Unbeatable header parsing time I'm sure. Or maybe in CSS. https://news.ycombinator.com/item?id=9567183 https://news.ycombinator.com/item?id=9567183
- nathan_f77 11y ago> Oh, I bet someone writes httpd in vhdl or verilog I would love to see that.
- deleted 11y ago[deleted]
- taftster 11y agoProbably one written in JavaScript, I'm guessing.
- userbinator 11y agoThat already exists, Node.js.
- mhd 11y agoWhere the web-serving part is written in C.
- SteveBerta 11y agoAmazing!
- kragen 11y agoI wrote httpdito, a web server for Linux in 386 assembly, a couple of years ago (mostly outdated discussion is at https://news.ycombinator.com/item?id=6908064; https://news.ycombinator.com/item?id=6908064; a README is at http://canonical.org/~kragen/sw/dev3/httpdito-readme http://canonical.org/~kragen/sw/dev3/httpdito-readme) and I was happy to get the executable under 2000 bytes. I actually used it the other day to test a SPA, although for some things its built-in set of MIME-types leaves something to be desired. But it doesn't have default documents, different kinds of error responses, TCP_CORK, sendfile() usage, content-range handling, or even request logging. So asmttpd is way more full-featured than httpdito, and it's still under 6K. (...httpdito possibly doesn't have any bugs, either, though ☺)
- sagargv 11y agoWhat's the performance like?
- polarbaer 11y agoNo dependencies at all, runs on Docker 'FROM scratch', Nice! - https://registry.hub.docker.com/u/0xff/asmttpd/ https://registry.hub.docker.com/u/0xff/asmttpd/