4 ms·
There are sequences in ISO-8859-1 that will cause UTF8 application to balk if they are read without conversion - sometimes in a part that is well away from the
by vaduz 6y ago
There are sequences in ISO-8859-1 that will cause UTF8 application to balk if they are read without conversion - sometimes in a part that is well away from the input itself.
A common example "órny", part of a number of Polish place and personal names. ISO 8859-1 (and -2, and others) form would be F3 72 6E 79. UTF-8 treats F3 as a start of a 4 byte sequence where each subsequent character is in 80 - BF range (as specified in RFC 3629):
UTF8-4 = %xF0 %x90-BF 2( UTF8-tail ) / %xF1-F3 3( UTF8-tail ) /
%xF4 %x80-8F 2( UTF8-tail )
UTF8-tail = %x80-BF
This nicely causes an exception in both Java and Python, for instance, if they were to do what Go does here - and there is nothing anyone from adding an X- header to the request or response to try just that.