3 ms·
Symbolically same but semantically different characters are most of the time represented differently in the byte level so that searching in large text can still
by diegoperini 6y ago
Symbolically same but semantically different characters are most of the time represented differently in the byte level so that searching in large text can still be implemented as simple byte comparison, if Unicode can be called simple.
- lonelappde 6y ago"simply byte comparison" is not a term associated with Unicode. Unicode is usually in strings, and reverse engineering a string is -- very hard, and surprisingly I haven't found a paper proving whether it is NP complete.
- ken 6y agoThat's the opposite of the rule for CJK characters, right? For those, symbolically-same characters get the same codepoint regardless of semantics (language) or shape (glyph).