5 ms·
UNICODE and ASCII apostrophes are a bit absurd. For KeenQuotes[1], my library to automatically curl straight quotes, there's an Apostrophe type that defines var
by thangalin 2y ago
UNICODE and ASCII apostrophes are a bit absurd. For KeenQuotes[1], my library to automatically curl straight quotes, there's an Apostrophe type that defines variations on how to convert a straight apostrophe to a curled one. The main issue is that most suggestions are to use ’, which isn't semantically correct[2], and at one point Michael Everson noted, "the alphabetic property should be restored to U+02BC"[3]. I've bucked the x27, U+2019, and rsquo trend with:
/** No conversion is performed. */
CONVERT_REGULAR( "'", "regular" ),
/** Apostrophes become MODIFIER LETTER APOSTROPHE ({@code ʼ}). */
CONVERT_MODIFIER( "ʼ", "modifier" ),
/** Apostrophes become APOSTROPHE ({@code '}). */
CONVERT_APOS_HEX( "'", "hex" ),
/** Apostrophes become XML APOSTROPHE ({@code '}). */
CONVERT_APOS_ENTITY( "'", "entity" );
Thoughts?
[1]: https://whitemagicsoftware.com/keenquotes/ https://whitemagicsoftware.com/keenquotes/
[2]: https://tedclancy.wordpress.com/2015/06/03/which-unicode-character-should-represent-the-english-apostrophe-and-why-the-unicode-committee-is-very-wrong/ https://tedclancy.wordpress.com/2015/06/03/which-unicode-cha...
[3]: http://www.unicode.org/L2/L1999/n2043.pdf http://www.unicode.org/L2/L1999/n2043.pdf
- crazygringo 2y agoBut Unicode code points aren't supposed to represent semantics, they're supposed to represent... well, "abstract characters" that are either glyphs or get combined into glyphs, which can have multiple (ambiguous) semantic meanings. That's why there aren't two different period characters to represent the end of a sentence vs. a decimal point, or two different em dashes where one represents a pause while the other comes at the end of dialog to indicate the sentence was interrupted (literally the opposite of a pause, it's being cut off). So since the apostrophe and the right single quote are visually identical, Unicode stays consistent in recommending that they be the same character. The name Unicode gives to a character is intended to represent one of its semantic meanings, not all of them. (Unicode does have plenty of visually identical characters, but they generally belong to totally different languages, like the English "o" and the Greek omicron "ο".)
- thangalin 2y agoThanks for shedding a little more light. Ignoring the semantics, in this case, between an apostrophe and a right single quote has resulted in many documents containing information that can not be parsed unambiguously because we have to pick a glyph and doing so with an ambiguously defined glyph loses contextual information. As a side-effect, since GPTs are based on the examples we give, they can't encode the proper punctuation for many phrases that use British English quotation mark styles, making them unable to "curl" the quotation mark properly. For example, none can curl this paragraph correctly: ''E's got a 'ittle box 'n a big 'un,' she said, 'wit' th' 'ittle 'un 'bout 2'×6". An' no, y'ain't cryin' on th' "soap box" to me no mo, y'hear. 'Cause it 'tweren't ever a spec o' fun!' I says to my frien'. The other downside to using ' or ' is that most fonts treat them as straight quotes, making for "improper" English typography when typeset into a book.
- crazygringo 2y ago> many documents containing information that can not be parsed unambiguously Well, and I'd suggest the unambiguous information was usually never there in the first place. It's less of an encoding problem, and more of an input "problem". People type either a single quote/apostrophe, or a double quote, and let smart quotes sort it out. And sure, smart quotes will fail spectacularly with your spectacularly pathological example! Heck, it took me a few seconds to figure out what on earth was going on with the first 5 characters. :) Your example would usually be typeset properly in a physical published book because it's done with professionals manually reviewing the typography. Just throw it in the bucket of hyphens vs. minuses vs. dashes em and en, x's versus multiplication signs... our symbols are full of ambiguities, it's not just apostrophes.
- thangalin 2y ago> let smart quotes sort it out. Smart quotes fail in simple cases, too. https://gitlab.com/DaveJarvis/KeenQuotes/-/tree/main/src/test/resources/com/whitemagicsoftware/keenquotes/texts https://gitlab.com/DaveJarvis/KeenQuotes/-/tree/main/src/tes... I've developed a lexer/parser that can disambiguate most cases, but wow was it a chore to write. > Well, and I'd suggest the unambiguous information was usually never there in the first place. Interesting. Isn't the text ambiguous because the glyphs lack the semantics to capture the usage of apostrophes versus closing single quotes? It's a Catch-22, isn't it? If UNICODE had semantics for apostrophes versus right single quotes, then our documents would be unambiguous. But we can't make them unambiguous because UNICODE doesn't capture these semantics.