10 ms·
Unicode 15.0 Slide Show
- deleted 4y ago[deleted]
- TheRealPomax 4y agoAndrew's Babelmap [1] is one of those applications that, if you do anything text or typography related, is basically required owning. With a donation, of course. [1] https://www.babelstone.co.uk/Software/BabelMap.html https://www.babelstone.co.uk/Software/BabelMap.html
- virtualritz 4y agoI'm usually ok with what macOS Character Viewer offers. I am rarely on Windows and didn't know about BabelMap. It looks like it fills the gap there. I work mostly on Linux so I hacked a Character Viewer clone in Rust over a weekend recently[1]. It just does what I need but I'm planning to add features to it if I find them useful. So I am curious: what functions does BabelMap offer that you can't live without, especially as a typographer? [1] https://github.com/virtualritz/glyphana https://github.com/virtualritz/glyphana
- arm 4y agoSince you mentioned macOS, it would be remiss of me to not mention UnicodeChecker: https://earthlingsoft.net/UnicodeChecker/index.html https://earthlingsoft.net/UnicodeChecker/index.html
- lpghatguy 4y agoI am a big fan of unicode.link[1] for inspecting existing strings, searching for codepoints, and comparing how strings get encoded. [1] https://unicode.link/ https://unicode.link/
- phkahler 4y agoGNU unifont has the entire MBP hut is a bitmap font. Is there an equivalent monospaced scalable font we can use in GPL software?
- politelemon 4y agoNoto Sans? https://fonts.google.com/noto/specimen/Noto+Sans+Mono https://fonts.google.com/noto/specimen/Noto+Sans+Mono
- phkahler 4y agoUnfortunately I also want technical symbols Depth, Counterbore, and Countersink and perhaps more.
- troymc 4y agoSome random characters didn't render in my browser. Upon inspection: font-family: Georgia, Serif; I don't think those fonts support all of Unicode. Google created their Noto fonts [1] for this purpose; I wonder why those aren't being used. [1] https://en.wikipedia.org/wiki/Noto_fonts https://en.wikipedia.org/wiki/Noto_fonts
- mistrial9 4y agogentium is a font with a very large number of glyphs also https://software.sil.org/gentium/ https://software.sil.org/gentium/
- jfk13 4y agoBut only for Latin/Greek/Cyrillic scripts; it makes no claim to be a pan-Unicode font (family).
- mistrial9 4y agook I did review the Gentium characters after this.. Gentium is an odd set of every dot, egrave, umlaut and circumflex, in many, many variations. You are right that Gentium does not make an effort to represent unicode code pages. Instead, it seems that the designers know about very old and odd language systems of Eastern and Central Europe, and have reproduced those.
- jfk13 4y agoBrowsers will generally do "fallback" to some other font, if the font(s) named in the CSS don't support the characters present in the text. But for some of the rarer characters, you may not have any available font that supports them.
- Someone 4y agoSee https://en.wikipedia.org/wiki/Fallback_font https://en.wikipedia.org/wiki/Fallback_font. It typically isn’t a browser feature, but an OS one. Since 1998 MacOS has a “last resort” font that has glyphs (not necessarily unique) for every Unicode code point. They donated it to Unicode (https://en.wikipedia.org/wiki/Fallback_font#Unicode_Last_Resort_font https://en.wikipedia.org/wiki/Fallback_font#Unicode_Last_Res...), so I expect most OSes running full-blown modern browsers to have it or something similar (those running smaller browser engines may be too space constrained to have room for it)
- nanis 4y agoAnd yet there is still no unambiguous lower case "I" or upper case "i".
- ClumsyPilot 4y agothats the job of a font, not encoding
- nanis 4y agoNo, the fact that there is no codepoint that makes those mappings ambiguous is due to the way Unicode decided to save to codepoints for seemingly no good reason. What should be _the_ value of `"I".lower()`? Or, "i".upper()? And please don't bring up locales. The whole point of accepting the complexity of Unicode is to be able to take a document which stands on its own without external references. > Early character encodings also conflicted with one another. That is, two encodings could use the same number for two different characters, or use different numbers for the same character. > The Unicode Standard provides a unique number for every character, no matter what platform, device, application or language.[1] Those statements are outright lies: Unicode does not provide a unique number fpr "upper case Turkish dotless i". Nor does it provide one for "lower case Turkish dotted i". If it did, it would be possible to correctly map "i" to "I" or "İ" and "I" to "i" or "ı" without having to know anything other than the source codepoint. The font does not even come into play here. [1]: https://unicode.org/standard/WhatIsUnicode.html https://unicode.org/standard/WhatIsUnicode.html
- deleted 4y ago[deleted]
- Kwpolska 4y agoUnicode is for representing text, not allowing arbitrary manipulation of it. It isn’t the job of Unicode to encode those relationships. Also, the Turkish `i` stuff is just the tip of the iceberg. Should Unicode be able to round-trip `'ß'.upper().lower()`? Keeping the existing capitalization of ß → SS, you need to define a "uppercase S that used to be ß" character. Then there’s the Dutch `ij`, in which both characters are either uppercase or lowercase (`Ij` at the start of a word is incorrect). There’s a ligature in Unicode, but it’s only for compatibility with some legacy keymaps. But is there a point in adding a new version of "S" that a lot of software would not recognize as equivalent to the plain old ASCII "S" (and one might end up far away from a ß due to copy-pasting or stuff), bringing weird bugs and security issues? Should the Dutch throw out all their keyboards just so they get a new key for the special IJ ligature?
- paulirish 4y agoPerhaps a better match for some folks' expectations, the Unicode Consortium's YouTube has plenty of talks on https://youtube.com/@unicode https://youtube.com/@unicode. Low view counts but often quite fascinating
- pavlov 4y agoWikipedia has the “no original research” rule. Unicode really should have had a similar “no original designs” rule. Too late now — it’s become an annually refreshed collection of fun fashionable clip art instead of an impartial repository of humankind’s symbols.
- chungy 4y agoUnicode's stuck to that principle better than Wikipedia has stuck to its principle.
- tokinonagare 4y ago[flagged]
- Name_Chawps 4y agoThis split already exists. Emoji are kept in a separate block from other characters.
- mcswell 4y agoNot exactly. It's true that they are (mostly) not found in the same block as other (alphabetic, numeric, punctuation etc.) characters. But there is no single emoji block, which makes it hard to write regular expressions to search for emoji, unless you are using a regex language that allows reference to Unicode properties (like Perl, or the Python3 regex library, not to be confused with the more commonly used re library).
- jurimasa 4y agoWhat makes a writing system such, and why emoji are not a writing system?
- tokinonagare 4y agoLet me introduce 2 websites that can answer your questions : Google and Wikipedia. The former allow you to search website on the called web platform, and the latter is an online collaborative encyclopedia. By using those two websites, you can easily find the definition of a writing system, which is "[...] a method of visually representing verbal communication" and reflect on the fact that since emoji aren't used to encode an oral message, they don't form a writing system.
- willm 4y agoI found this more entertaining than the new Avatar movie.
- dhosek 4y agoI kind of expected this to be more an overview of the new stuff in Unicode 15.0. As the author of a Rust Unicode crate (finl_unicode), I always like to dig through the release notes to see what sort of strange new stuff is on offer.
- mycall 4y agoI would love to see someone make an image to unicode "curve fitting" algorithm or converter, similar to ANSIDRAW.
- hollasch 4y agoSee https://shapecatcher.com/ https://shapecatcher.com/.
- PostOnce 4y agoTangent: I recognized the domain and tried to remember why, and now I remember. I'm working on a game, and babelstone.co.uk has probably the world's most comprehensive (and high quality) set of runic fonts: https://www.babelstone.co.uk/Fonts/ https://www.babelstone.co.uk/Fonts/ https://www.babelstone.co.uk/Fonts/Runic.html https://www.babelstone.co.uk/Fonts/Runic.html https://www.babelstone.co.uk/Fonts/AngloSaxon.html https://www.babelstone.co.uk/Fonts/AngloSaxon.html
- einpoklum 4y agoThis is the most important part of Unicode for me: https://www.unicode.org/reports/tr9/tr9-46.html https://www.unicode.org/reports/tr9/tr9-46.html because I speak a right-to-left language. Whoever wants to write an application involving text entry, and truly support localization or internationalization, should take the time to read at least section 3: https://www.unicode.org/reports/tr9/tr9-46.html#Basic_Display_Algorithm https://www.unicode.org/reports/tr9/tr9-46.html#Basic_Displa...
- raffy 4y agoI have tool which lets you scroll through all characters in a large document. It also shows names, scripts, and IDNA-like status. https://adraffy.github.io/ens-normalize.js/test/chars.html https://adraffy.github.io/ens-normalize.js/test/chars.html
- another2another 4y agoBlock 309 : Chess Symbols Wow, Unicode seems to be expanding way out of the original goal of representing written characters - how often would these symbols be even used? And how many font sets would bother including them?
- an_ko 4y agoWikipedia has some rationale: https://en.m.wikipedia.org/wiki/Chess_symbols_in_Unicode https://en.m.wikipedia.org/wiki/Chess_symbols_in_Unicode Probably the most notable: > Use figurine algebraic notation, which replaces the letter that stands for a piece by its symbol, e.g. ♘c6 instead of Nc6. This enables the moves to be read independent of language (the letter abbreviations of pieces in algebraic notation vary from language to language).