21 ms·
By the same reasoning, the 7-eyed O has now been used more than once, so it deserves a glyph! So the right way to do this is to introduce a new character for th
by etamponi 4y ago
By the same reasoning, the 7-eyed O has now been used more than once, so it deserves a glyph! So the right way to do this is to introduce a new character for the correct glyph, and also leave the current one (perhaps changing the title). Otherwise these tweets won't make when read by someone that updated to Unicode 15.0
- jotato 4y agoMy thought as well
- baybal2 4y agoUnicode basic rule is that character definitions never ever change, even when enumerated erroneously.
- Arnt 4y agoYes, but this is a change either way, because that codepoint's definition referred to that character. Either the reference or the description of the appearance has to change.
- echelon 4y agoꙮ ꙮ ꙮ ꙮ ꙮ ꙮ ꙮ ꙮ ꙮ ꙮ ꙮ ꙮ ꙮ ꙮ ꙮ ꙮ ꙮ ꙮ ꙮ ꙮ ꙮ ꙮ ꙮ ꙮ Make a new character. Updating the existing character ruins the meaning of all previous usages. It's like trying to change an API. Don't disrespect your existing users. Make a new version. (ꙮ ͜ʖꙮ) Think of all the ASCII art this botches. That has to have some historical importance to the Unicode standards body. (⌐ꙮ_ꙮ) For scholarly digital (unprinted) documents where the correct character rendering matters, erroneous past usages can be trivially found with grep, a date search, and easily corrected. The domain experts will familiarize themselves with this issue and fix the problem. Don't take a shotgun to it! This message wꙮn't have the ꙮriginally intended meaning if the characters are updated from underneath. ꙮ ꙮ ꙮ ꙮ ꙮ ꙮ ꙮ ꙮ ꙮ ꙮ ꙮ ꙮ ꙮ ꙮ ꙮ ꙮ ꙮ
- koboll 4y agoHonestly it probably deserves the Pluto treatment: decertification as a character. One historical use in the 1400s doesn't merit a character and never did.
- colejohnson66 4y agoUnicode's mission is to make every document "roundtrip-able". Even if a character is only used once, it should be possible to save a plaintext version of the containing document without losing any information. Roughly, I should be able to put a transcription of that one translation from the 1400s on Wikisource without using images. You may disagree with me, and that's fine, but it doesn't change Unicode's mission. Besides, there's room for 1,112,064 codepoints[a], and only 149,146 are in use. It's predicted we'll never use it up, so what harm is there in one codepoint no one will ever need? [a]: U+10'FFFF max; it used to be U+FFFF'FFFF, but UTF-16 and surrogates ruined that
- modzu 4y agowhy isnt the artist formerly known as prince in unicode?
- djur 4y agoUnicode doesn't have a character for every illuminated initial, nor should it. I'm not clear on why this character should be considered any differently.
- j-bos 4y agoBecause it's already been added to unicode. Now it's not a question of whether or not to add, rather to remove, and unicode almost by definition does not remove.
- thayne 4y agoUnicode does have deprecated code points though. Not that I necessarily think making this character deprecated makes sense.
- martin_a 4y agoUff. I'm not sure we have space for another glyph in Unicode. Looks pretty packed in here...
- BlueTemplar 4y agoUTF-8 is still more than 80% empty, and can be potentially extended...
- colejohnson66 4y agoTheoretically, UTF-8 can encode up to 31 bits (U+7FFF'FFFF)[0], but for compatibility with UTF-16's surrogates, it's officially capped to 21 bits with the max being U+10'FFFF[1]. That decision was made November 2003, so there's two decades of software written with hard caps of U+10'FFFF. [0]: https://www.rfc-editor.org/rfc/rfc2279 https://www.rfc-editor.org/rfc/rfc2279 [1]: https://www.rfc-editor.org/rfc/rfc3629#section-3 https://www.rfc-editor.org/rfc/rfc3629#section-3
- echelon 4y agoThis thread on HN won't make sense in the future if the Unicode body replaces ꙮ Make a new character!
- nerfhammer 4y agowhy not make an additional eye a diacritic mark so you can just add an arbitrary number of eyes