3 ms·
(Author here.) I agree the only correct way C1 controls can work is encoded within UTF-8 data, else nothing works. The context is escaping C0 control characte
by dgl 3y ago
(Author here.)
I agree the only correct way C1 controls can work is encoded within UTF-8 data, else nothing works.
The context is escaping C0 control characters is simple, you look for a single byte and filter it as you need. C1 controls when not encoded as UTF-8 are also single byte characters and therefore easy to filter out. For multi-byte encodings, you need to correctly decode the encoding, then filter, this has synchronisation problems and is tricky[1] (particularly as a Unix byte stream doesn't tell you it's encoding, although you can assume if someone is trying to display it as text it is probably UTF-8 these days).
I should probably expand on that recommendation more, there's a lot in the paper, but the crux of the issue is C1 controls are a legacy thing and serve no useful purpose anymore and as I've shown there are enough issues just dealing with C0 controls.
If C1 controls are encoded as UTF-8 as is the only way they can work on a modern system, then they take up as many bytes as C0 controls (e.g. CSI is "\e[" or U+009B, which encoded as UTF-8 is 0xC2 0x9B) so they don't even save bytes on the wire.
[1]: see https://www.openwall.com/lists/oss-security/2015/09/20/1 https://www.openwall.com/lists/oss-security/2015/09/20/1 and some replies to that.