3 ms·
In case anyone sees this: I initially agreed with you, thinking that pure UTF-8 validation is a bit more of a niche than might be expected. But if I'm not mista
by zwegner 7y ago
In case anyone sees this: I initially agreed with you, thinking that pure UTF-8 validation is a bit more of a niche than might be expected. But if I'm not mistaken, I think at least for JSON validation, the UTF-8 validation can happen completely independently of the parsing. I don't think any control characters (braces, commas, quotes, escapes, etc) will change the validity of the UTF, or vice versa. So it might be beneficial (both for speed and security, as the sibling comment notes) to validate chunks of input as UTF-8, then pass them off to the parser, which now doesn't need to deal with validation.
Maybe this could be applied to XML, but that's quite a behemoth of a standard--I really wouldn't be surprised if there was a way to switch encodings mid-stream. I have no idea though...
- nitwit005 7y agoSure, you can validate and then parse the JSON, but parsing it is also effectively validating it as UTF-8, except for the string values in the JSON, so it's heavily going to be duplicated effort. And unfortunately, the strings can contain things like invalid surrogate pairs by using unicode escapes, so you may want to validate after escaping. You can't normally blindly parse XML as UTF-8. The encoding has rules for detecting the character set.