3 ms·
So...what do we do about it? How about a standard algorithm/mapping table for grouping unicode chars to known languages and flagging when a unicode string pull
by jbert 16y ago
So...what do we do about it?
How about a standard algorithm/mapping table for grouping unicode chars to known languages and flagging when a unicode string pulls chars from more than one language? (possibly only flagging if a change occurs within a 'word').
A standard way of rendering unicode strings could then be too highlight flagged chars (perhaps with a different colour). More sensitive areas (such as the browser url bar) could fire additional alerts if any chars were flagged.
Feels like a nice minor RFC to me - anyone see a problem with it?
- woodall 16y agoI like your idea of changing the color of the letters if they are in a different language. I might actually work on that.
- jbert 16y agoCool. I think a well-defined algorithm for defining a flagged character is important, to get cross-platform/application support and to validate the algo. (Any bug in the algo is a possible security hole, moreso if people become to trust that any dodgy characters are highlighted). An exposed API to handle UTF8 strings to get count and/or indices of flagged chars would also be useful. Perhaps it could work as an extension to ICU? http://site.icu-project.org/ http://site.icu-project.org/ If nothing else, the "change language within a word" should use a unicode-sane definition of 'word', which ICU would give you. Once you have the two pieces above, adding the colouring etc should be reasonably straightforward. But it'd be a shame if the two pieces above weren't factored out - since I don't think the widespread adoption would follow otherwise.