4 ms·
OP converted a Unicode file to an ascii file. Isn’t the loss of precision what is “compressing” the file? The Shakespeare corpus probably doesn’t use any Unic
by tarr11 3y ago
OP converted a Unicode file to an ascii file.
Isn’t the loss of precision what is “compressing” the file? The Shakespeare corpus probably doesn’t use any Unicode characters.
- mbb70 3y agoA file cannot be 'Unicode'. Unicode is just a catalog of symbols one might want to express in writing. An encoding (like ASCII or UTF-8) describes how bytes are mapped to such symbols. ASCII only describes 128 characters whereas UTF-8 can represent any Unicode symbol in bytes. Notably though, UTF-8 is a strict superset of ASCII. That is, every character ASCII represents is represented by the same bytes in UTF-8. So encoding Romeo and Juliet in ASCII or UTF-8 results in the same exact file.