4 ms·
I think what you are trying to say is: "because UTF-8 has invalid character sequences, we could potentially use one of them to represent end-of-string, which w
by bobbydavid 14y ago
I think what you are trying to say is:
"because UTF-8 has invalid character sequences, we could potentially use one of them to represent end-of-string, which would allow us the flexibility of a null-terminated string (not keeping track of the length) without the restriction of no-nulls-allowed."
You're right! Great. But you are not revealing a "strange thing" about Unicode. You are instead making a general comment about null-terminated strings. So why use such inflammatory and misleading language like "If you claim to support Unicode, you have to support NULL characters"?
Update: I don't object to your idea at all, it's a neat trick! It's just that the way it's phrased, it sounds like Unicode's design contributed to this NULL-terminal problem, when in fact even NULL-terminated ASCII strings cannot 'handle' a null character in this sense.
To augment your idea, though, how about you use '0xFF 0x00' as a terminator? This way, backward-compatibility is preserved in all cases except UTF-8 => ASCII with NULLs, and in this case the string will be truncated rather than a buffer overflow (i.e. "fail closed").