7 ms·
That kind of stuff is necessary for binary protocols to evolve in a compatible way. When a TLS 1.4 is defined, we need a window where clients and servers can st
by codeflo 3y ago
That kind of stuff is necessary for binary protocols to evolve in a compatible way. When a TLS 1.4 is defined, we need a window where clients and servers can still negotiate 1.3 until both have been upgraded. And 1.3 had to find ways to be compatible with 1.2, and so forth. Decades of that kind of evolution are guaranteed to leave some marks in the protocol.
But let's keep a sense of proportion. I find it hard to worry about "wasting" maybe 0.01% of the total internet bandwidth for a few extra bytes here or there, when that's necessary to keep the internet working at all, when on the other hand, for no end-user benefit, we don't hesitate to waste maybe 15% (after compression) by insisting on using text-based formats for the payload in all our web standards.
In fact, I'd love to see a back-of-the-envelope calculation of how many tons of CO2 would have been saved in total if HTML was a well-engineered binary format. (Including bandwidth, storage, parsing on the client etc.) The number must be insane.
- maccard 3y agoThe thing about html is that it's a verbose text format on the surface, but it compresses incredibly easily, and support for gzip is widespread. If you're concerned about size, brotli is better again.
- codeflo 3y ago> The thing about html is that it's a verbose text format on the surface, but it compresses incredibly easily, and support for gzip is widespread. If you're concerned about size, brotli is better again. That's why I said it wastes 15%, not 100%. Whenever text-based formats and binary formats are compared, the results after compression (of both) are usually in that range. You can debate those precise numbers, they might be lower for the brotli/HTML combination. That won't really change the point I was making in context: That a few extra bytes for backwards compatibility in the TLS handshake pale in comparison to the amount of waste we accept for encoding the payload.
- saiya-jin 3y agoI dare to say that human readability that accounts for that extra size allowed many people to learn basics about how internet works, namely HTML and its compatriots. Mankind would be poorer for this enriching experience, even if it resides mostly in the past, and for me personally 15% (or 100% extra) could easily justify that. Plus add easiness of debugging this (before you bring some javascript monstrosity that hacks around it and will stop working probably in weeks after paid support of it ends)
- toomim 3y agoA better example is HTTP, which is text, but read by humans much less frequently.
- tialaramex 3y agoOnly HTTP/0.9 HTTP/1.0 and HTTP/1.1 are probably text in the sense you mean HTTP/2 and HTTP/3 are binary formats, they are semantically very similar to the older formats in some sense, but you would benefit from more tools to examine them properly because human readability was not the priority.
- philsnow 3y agoYou’re saying that compressed HTML, because it is text, wastes 15% (could be however much, the number doesn’t really matter) over some binary format that would express the same content? What binary format would/could that be? I’m just not seeing how (if it expresses the same content) it could be smaller than compressed HTML.
- foobazgt 3y agoThey're comparing compressed text vs compressed binary, apples to apples. While text compresses amazingly well, binaries aren't 100% entropic themselves. They usually also benefit from compression. For example, the defacto "executable" format for the Java runtime is a compressed archive (jar file).
- foobiekr 3y agoThe compression doesn’t help memory usage or cache locality on the receiving end. Http is terrible and JSON and other web technologies are terrible. The whole concept of how to write efficient code and the resulting performance left on the table by terrible coders is enormous. A terrible waste.
- Szpadel 3y agothe thing is that you need additional CPU power to compress and decompress that data, maybe that is not much by today standards, but when accumulated it could be a significant number
- szundi 3y agoThe alternatives are not free either and would accumulate different costs probably as well
- ChoHag 3y agoIf HTTP or HTML were binary formats 1000s of embryonic engineers would not have been able to learn them by hacking codes into telnet or notepad and the web would now be even more concentrated in the hands of the few players who are busily engaged in fucking it to death.
- tialaramex 3y ago> That kind of stuff is necessary for binary protocols to evolve in a compatible way. When a TLS 1.4 is defined, we need a window where clients and servers can still negotiate 1.3 until both have been upgraded. And 1.3 had to find ways to be compatible with 1.2, and so forth. Decades of that kind of evolution are guaranteed to leave some marks in the protocol. It's necessary because people are incompetent, and because overall the market rewards them for incompetence. From the outset TLS provided a trivial version negotiation mechanism but it was easier to ignore it and write incompatible garbage especially for so-called "Middle boxes" often sold as a drop-in security "solution" for businesses. So when it came time to ship TLS 1.1, it was soon discovered that in practice you can't just say "Hi, I speak TLS 1.1" which would be a couple of bytes - that won't work, you need to find some other way to quietly signal to competent people that you know the newer protocol. So they did, slightly weakening the security in the process, and this continued into TLS 1.2 where browsers began doing "Fallback" which was a risky but sadly necessary process where you give up attempting the new protocol altogether sometimes, thus opening yourself up to supposedly obsolete attacks. By TLS 1.3 things had become so bad that TLS 1.3 essentially begins, as you can see if you inspect the data shown on that page, by pretending we're speaking TLS 1.2 and then saying we want to negotiate an optional "extension" to TLS 1.2 which is where we confess we actually speak TLS 1.3. Every single packet of TLS 1.3 encrypted data is also wrapped in a TLS 1.2 layer saying "Don't mind me, I'm just application data". Why? Because we can't confess we're not speaking TLS 1.2, ever, and if we said we were doing TLS 1.2 crypto system stuff the same incompetent garbage software would try to get involved because it "understands" (badly) how to speak TLS 1.2, so we just pretend it missed the negotiation phase, this is just application data, nothing to see. And it works. That's crucial. It's why we did all this, and yet it reveals that because the products people bought were developed incompetently they wouldn't even have detected serious attacks anyway, let alone prevented them. Need to sneak 40GB of stolen financial data over a network "protected" by this Genuine Marketing Leading Brand Next Generation Firewall? Don't worry, just label it "application data" with no explanation and it'll be completely ignored. If you've ever watched a Lock Picking Lawyer video on Youtube, it was like one of the ones where it's several minutes so you expect it'll be hard to pick, but then you find out he's actually so disgusted by the lacklustre security of this $150 "Pick Proof High Security Lock" that although he rakes it open in 2 seconds with a cheap tool, and then shims it open with a discarded Redbull can, and then knocks it open with a hammer, and then uses a purpose built bypass tool to open it instantly in a single flowing motion, he also takes time to disassemble it and show you that the manufacturer fucked up, wasting material solving a non-existent problem and in the process making the lock much worse, which is why the video was so long. Learning from their experience with these "Security" products for TLS 1.3, the QUIC people designed QUIC specifically with the intent that you can't even tell what version it is unless you're the client or the server, and then they shipped a new QUIC version to check that works even though they don't really need one yet, so that they don't have to do this whole dance again every few years.
- MrGilbert 3y ago> In fact, I'd love to see a back-of-the-envelope calculation of how many tons of CO2 would have been saved in total if HTML was a well-engineered binary format. I wonder if it is feasible to create something like that. Because, a binary format requires specialized tooling, which needs to be created and maintained. But maybe, with further adoption of WebAssembly, HTML will be less important in 10 - 15 years. :)
- codeflo 3y agoYou wonder if its feasible to create what, a binary format?
- MrGilbert 3y agoI wonder if there is a significant CO2 reduction from binary over plaintext, if you figure in the additional work that a binary format requires. As such I wonder if it is feasible to create a report, as it will never be able to factor everything in. You assume that there will be a significant reduction. I assume that we don’t know for sure.
- wizzwizz4 3y ago> In fact, I'd love to see a back-of-the-envelope calculation of how many tons of CO2 would have been saved in total if HTML was a well-engineered binary format. Far, far less than would be saved if HTML was still expected to be human-readable. Things like React (server-side and client-side: think of all those poor <div>s!) waste far more resources than decompressing and parsing HTML.
- chx 3y agoDon't forget to add a few billion dollars saved because of how easy it is to debug HTML. How easy to is for people to start. You press ctrl+u and see the source. There were not many platforms where you could just view the underlying instructions... when I grew up, some 8 bit machines had cartridges to hack but let's face it 6502 or Z80 assembly is nowhere near as friendly as HTML.
- stusmall 3y agoctrl+u could just make the displayed representation as easy to read. There are many binary protocols/formats that are a breeze to work with if you have the right tooling. You never see the 0xa 0x8 0xf. Your tools will either show you what the bytes represent or that there was a parsing error. Those parsing errors would be rare in the average case, just like debugging strange unicode issues in HTML today.
- philsnow 3y agoExactly; when you click the padlock icon, the browser shows you the parsed representation of the x.509 certificate chain, not the ASN.1 bytes.
- foobiekr 3y agoMost binary formats are easier, not harder, to parse. It’s sad. An entire generation or two that doesn’t know the basics.
- easrng 3y agoMost text formats are easier to parse but in incorrect ways.
- chx 3y agoAt that point, how big is the win in a binary protocol versus compression? Wouldn't a binary version be just a mapping between the HTML tags and a bit representation which a Huffman compression directory can just recreate on the fly but better?
- dylan604 3y ago>In fact, I'd love to see a back-of-the-envelope calculation of how many tons of CO2 would have been saved in total if •we didn't include massive JS libraries that are not truly necessary •we didn't track user's every move and report that data back