6 ms·
That's the beauty of original HTTP - simplicity. Same as parsing HTML(in the 90s). With HTTPS(S as Secure) it's a whole different story and most programmers use
by milansuk 5y ago
That's the beauty of original HTTP - simplicity. Same as parsing HTML(in the 90s). With HTTPS(S as Secure) it's a whole different story and most programmers use some library.
- nly 5y agoNo, parsing HTTP/1.x is a nightmare and definitely not simple. It wasn't even particularly well defined until 2014 when the original RFCs were modernized, and even now there are bugs reported in HTTP parsers all the time. Node.js came out in 2009, a full ten years after HTTP/1.1 (RFC 2068) and its original http-parser is rather hard to follow, doesn't conform to the RFCs for performance reasons, and is considered unmaintainable by the author of it's replacement[0] As for parsing HTML, well go look at how Cloudflare have stumbled[1] [0] https://github.com/nodejs/llhttp https://github.com/nodejs/llhttp [1] https://blog.cloudflare.com/incident-report-on-memory-leak-caused-by-cloudflare-parser-bug/ https://blog.cloudflare.com/incident-report-on-memory-leak-c...
- ibraheemdev 5y ago> Node.js came out in 2009, a full ten years after HTTP/1.1 (RFC 2068) and it's original http-parser is full-on spaghetti code, doesn't conform to the RFCs for performance reasons, and is considered unmaintainable by the author of it's replacement That's because of the way the parser is written. There are other simpler parsers that are much more readable.
- hdjjhhvvhga 5y agoThe fact that someone wrote a parser that's hard to follow doesn't mean that parsing HTTP/1.x is extremely difficult. What is really hard is to construct a parser that is at the same time (1) fast, (2) complete, (3) secure. It is much easier to choose just two, compare e.g. the one based on Nginx[0] vs picohttpparser [1]. [0] https://github.com/Samsung/http-parser/blob/master/http_parser.c https://github.com/Samsung/http-parser/blob/master/http_pars... [1] https://github.com/h2o/picohttpparser/blob/master/picohttpparser.c https://github.com/h2o/picohttpparser/blob/master/picohttppa...
- a-dub 5y agofor the basics, however http/1.x is pretty simple. you can test webserver health by literally typing in the request. i suspect the complexity you speak of is similar to MIME. where SMTP/POP/IMAP are pretty simple, things got pretty hairy with the introduction of MIME, SASL and friends. i think, though, that most of the complicated stuff in http is optional, is it not? like if you don't send a header that compression is supported, the server won't compress... or am I misremembering? either way, simpler to understand from a packet capture than a grpc stream or spdy/http2 stream.
- jsjohnst 5y agoI once heard that it’s impossible to build a “spec compliant” IMAP4 library as the spec itself is contradictory. Don’t have a reference to prove it, so I could be wrong.
- jart 5y agoPretty much everything is optional if you stick to http/1.0. If you implement http/1.1 then you're required to do a lot of non-essential stuff like chunk encoding, pipelining, and provisionals which themselves are reasonably trivial too but they make the server code less elegant. If you want a protocol that's actually hard, implement SIP.
- giancarlostoro 5y agoSomewhat related but in the Python space of things: I love that Python has a standard for web frameworks so much so that you can build your own web framework that targets said standard and it can be deployed anywhere without getting lost in the weeds of parsing HTTP. For example FastAPI is directly a ASGI compliant framework, and it is known as one of the fastest Python web frameworks out there. Bottle I think is also a raw WSGI framework and its all in one file. (ASGI is what became the natural progression for WSGI, think of it like the http package Rust wants to standardize).
- jart 5y agoI'm the author of the fastest open source HTTP server. Parsing HTTP 0.9, 1.0, and 1.1 is trivial. It's a walk in the park. It only takes about a hundred lines of code to create a proper O(n) parser. https://github.com/jart/cosmopolitan/blob/0b317523a0875d83d650ce8a7b288e3b3500fbba/net/http/parsehttpmessage.c#L54 https://github.com/jart/cosmopolitan/blob/0b317523a0875d83d6... The Joyent HTTP parser used by Node is very good but it's implemented in a way that makes the problem much more complicated than it needs to be. The biggest obstacle with high-performance HTTP message parsing is the case-insensitive string comparison of header field names. Some servers like thttpd do the naive thing and just use a long sequence of strcasecmp() statements. Joyent goes "fast" because it uses callbacks, which effectively punts the problem to the caller, and, for a few select headers which it handles itself, like Content-Length, it uses this really complicated internal "h_matching" thing for doing painstakingly written out hardcoded character compares. Redbean solves the problem by using better computer science: perfect hash tables. Thanks to gperf command. That makes the API itself much more elegant since the parser can not only go faster but return a hash-table like structure where individual headers can be indexed without performing string comparisons.
- mariusor 5y agoI think that implementing a proper state machine for the header parsing with ragel would give a more comprehensive result than using gperf or even the handmade one from your code. I think there are already some versions of the ragel code online, but they might be for other target programming languages.
- jart 5y agoI'm one of the authors of Ragel and I disagree with you. HTTP is trivial enough that you'd be better served writing the state machine yourself using a switch statement. See my GitHub link above for an example. The code easily ports to other languages, like Java. Lastly when it comes to Ragel and gperf, they both do two completely different things. Ragel would generate a prefix trie search in generated code which would have enormous code size compared to what gperf is doing, which is much faster. With gperf, you only need to consider exactly O(3) octets total to tell which header it is. After that, it does a single quick string compare to confirm it's one of the predetermined headers rather than some unknowable value.
- strictfp 5y agoThe whole idea behind Node.js was to write a super-efficient completely nonblocking http server in C, while keeping all the business logic in a simple scripting language. You should not expect the Node.js parser to be simple.
- na85 5y agoSeems like it's yet another example of the node ecosystem being amateur hour, rather than a problem with HTTP.
- secondcoming 5y agoThere's nothing simple about HTTP. It looks like it should be simple, but it isn't.
- Arnavion 5y agoBut HTTPS just adds TLS. You can use "some library" to do the TLS handshake and subsequent encryption, and end up with a readable-writable stream that you can then parse HTTP from yourself. Your code is the same as when it was dealing with a TCP stream directly.
- darnir 5y agoHTTP/1.x is anything but simple. They were under defined and overly complex in many ways. The original RFC was so complex that when reworked, they split it into 6 documents. I've worked heavily on some HTTP implementations and its ridiculously hard to get them right. Not to mention, this "server" only responds to a simple well formed GET request. Without handling about 90% of what the HTTP specifications talk about. Its a nice project, but it doesn't speak to the simplicity of HTTP
- kevinoid 5y agoI agree. As with many things, it's only simple as long as you ignore the complexities. As they say, the devil's in the details. > this "server" only responds to a simple well formed GET request. And not even that. The Request-URI in a Simple-Request line (inherited from HTTP/0.9) may contain escape characters. (e.g. `GET /my%20file.txt` to get `my file.txt`) HTTP/1.0 states "The origin server must decode the Request-URI in order to properly interpret the request."[1] This server does not. Which is not to say that this server isn't interesting. Just that it's not a demonstration of how easy HTTP/1 is to parse. [1]: https://www.w3.org/Protocols/HTTP/1.0/spec.html#Request-URI https://www.w3.org/Protocols/HTTP/1.0/spec.html#Request-URI