3 ms·
String Prism
- nneonneo 3y agoGreat! A few notes: - It'd be nice to have additional transformations: NFKC, NFKD, and maybe some others like CLDR transliteration, confusable mapping, ... - Characters outside the BMP get broken down as surrogate pairs, which is not what you want. Non-BMP characters they should be treated as a single codepoint; surrogate pairs are an implementation artifact of JavaScript's internal UTF-16 representation.
- tialaramex 3y agoWell, of UTF-16 itself, not specific to Javascript. What they're exposing are the 16-bit Code Units of UTF-16, which happen to be (reserved) Unicode Code Points, but are not Unicode Scalar Values. I guess I don't know what this software is for? But the references to the normalisation forms makes me assume it's about Unicode and so it ought to work in the actual Code Points, not these surrogates.
- yuchi 3y agoThe UI breaks pretty easily with a “zalgo” amount of diacritics. That said, very interesting. A cool tool to have in your belt.
- speps 3y ago[flagged]
- benj111 3y agoIt might be useful to say what it is. Especially when you're on a mobile when you're presented with 2 exactly the same breakdowns and what looks like the bottom of the page.
- savolai 3y agoLove it. Broken on ios tho, for anything longer start doesn’t show.
- kamray23 3y agoCharacter rendering and recognition breaks with characters outside plane 0 (BMP), however, normalization seems to still work correctly. For example, [◌𑄮] u+11131 u+11127 (two characters) is interpreted as five characters but normalized correctly to [◌𑄮] u+11123 (one character) which is nevertheless interpreted as three characters for some reason.