19 ms·
The obvious risk with plain-text protocols is that you don't write a rigorous spec, and don't write a strict parser, but some hack up least-effort thing with a
by Pinus 6y ago
The obvious risk with plain-text protocols is that you don't write a rigorous spec, and don't write a strict parser, but some hack up least-effort thing with a few string.split() and whatever. This means there is a lot of slack in what is actually accepted, and unless you are in full control of both ends of the protocol, that slack will be taken advantage of, and unless you are more powerful than whoever is at the other end (which you aren't, if they're your clients and you are not Google or Facebook), you have to support it forever. So write plain-text protocols if you like, but make sure to have a rigorous spec, and a parser with the persicketyness of... I don't know.
- rini17 6y agoThat works until next time someone considers the strictness of the existing protocol too overwhelming or complicated and invents a new "simpler" one.
- GoblinSlayer 6y agoTCP is strict. If you want something rich, you just run another protocol on top of the strict one. There can be many layers: TCP -> TLS -> HTTP -> JSON -> Base64 -> AES -> PNG -> Deflate -> Picture.
- mumblemumble 6y agoI'm thinking here of "standards" like csv and mbox that are almost impossible to handle with 100% reliability if you don't control all the programs that are producing them. It can get even worse with some niche products. I used to work with a piece of legal software that defined its own text format, and had a nasty habit of exporting files that it couldn't import. There was a defined spec, but it was riddled with ambiguities. I'm coming to think that, when it comes to text formats, it's LL(k) | GTFO.
- deleted 6y ago[deleted]
- wtetzner 6y agoDon't you have the same problem with ill-specified binary protocols?
- Pinus 6y agoTo a certain extent, yes. However, with a binary protocol, at least the syntactic level tends to be fairly solidly locked down, simply because you have to. Two bytes bing-endian unsigned integer of that, four bytes big-endian signed integer of that, etc. The confusion happens at the semantic level. With text protocols, the confusion starts in the parser (see other comments about continuation lines in HTTP, for example). And I forgot to say in my first comment: If you design a text-based protocol that can carry textual data, make sure it can handle any text you throw at it (and preferrably any byte sequence), i.e. decide and implement a good quoting convention from the start. I have seen a text format, initially used to store config data, but later extended to other things in the system because it was there, that used <<< and >>> as string delimiters. Then some client named something <<<foo>>>. (This was early 90:s, so JSON wasn't invented yet!). And yesterday I spent a large chunk of my working day sorting out problems originally caused by a system exporting semicolon-separated almost-CSV, and an enterprising tester putting a semicolon in a text field.