12 ms·
Text rendering hates you (2019)
- chrismorgan 3y ago(2019) Previous discussions: • https://news.ycombinator.com/item?id=21105625 https://news.ycombinator.com/item?id=21105625 (29 September 2019, 542 points, 169 comments) • https://news.ycombinator.com/item?id=30330144 https://news.ycombinator.com/item?id=30330144 (14 February 2022, 399 points, 153 comments) —⁂— If you find this article interesting, you’ll probably also like the article someone else published a month later: “Text Editing Hates You Too” <https://lord.io/text-editing-hates-you-too/ https://lord.io/text-editing-hates-you-too/>. It has also been discussed here a couple of times: • https://news.ycombinator.com/item?id=21384158 https://news.ycombinator.com/item?id=21384158 (29 October 2019, 875 points, 282 comments) • https://news.ycombinator.com/item?id=27236874 https://news.ycombinator.com/item?id=27236874 (21 May 2021, 384 points, 182 comments)
- putlake 3y agoAnother problem with text rendering that's not mentioned in TFA is hyphenation. I have a table on a webpage that can have some long text in some cells. I have 2 simple requirements: # If a word will fit by moving it to the next line, then move it to the next line and do not hyphenate i.e., do not break the word. # If the word is too long to fit in a single line, then break it up. Hyphenate at will. There is no incantation combination of CSS properties word-break, word-wrap, overflow-wrap, hyphens and white-space that will do this. In 2023. I believe word-break: break-word does #1 but it's not hyphenating for me. And MDN says word-break: break-word is deprecated.
- sings 3y agoSounds like a use case for soft hyphens. If you don’t mind where the words break, you could sprinkle soft hyphens through the text to get it to break more often than the default renderer would otherwise.
- chrismorgan 3y agoSoft hyphens are no use in this case, because browsers take them as a break opportunity of equal standing with a space. In my experience, liberal sprinkling on soft hyphens makes things worse, not better, because you end up with loads of gratuitous hyphens. What’s needed is a different algorithm for breaking paragraphs into lines, something better than the greedy algorithm that all browsers use. Something like Knuth-Plass, which applies a penalty to hyphens so that it’ll use them if the alternatives are bad enough, but won’t be too eager about using them.
- DonHopkins 3y agoSupple hyphens, then?
- sings 3y agoI definitely agree that soft hyphens are a simplistic workaround if used in the way I suggested, and are indeed inferior to a complex and well-thought out hyphenation system. Still, if someone is considering `word-wrap: break-word;` for narrow table columns, soft hyphens are worth knowing about.
- hinkley 3y agoI have learned this, forgotten it, and then stepped on this garden rake again a handful of times in my career. HTML should treat hyphens as hints and nobody does.
- tronical 3y agoThe closest I can think of is for you to insert one or multiple soft-hyphens.
- chrismorgan 3y agoWhat should be done with hyphenation and indeed breaking paragraphs into lines in general is largely just undefined. There are mild movements from time to time, but overall no one’s sufficiently interested in implementing the really good stuff, so we’re left with the simple, easy and bad that everyone has grown used to. I’m glad to say that Chromium has just shipped `text-wrap: balance`, which is at least one step in directions of goodness. I hold out hope that some day some browser will implement a `text-wrap: pretty` backed by something like Knuth-Plass. https://bugzilla.mozilla.org/show_bug.cgi?id=630181 https://bugzilla.mozilla.org/show_bug.cgi?id=630181 is relevant, shows that some thought has gone into how it could be achieved in Firefox. And while talking of hyphenation, what happens if you try to hyphenate in the middle of what would otherwise be a ligature? e.g. at “af-fable”. Alas, in this instance no one has got really enthusiastic about fixing it in Firefox: https://bugzilla.mozilla.org/show_bug.cgi?id=479829 https://bugzilla.mozilla.org/show_bug.cgi?id=479829, you get the ff glyph split in half, much like the mixed-colour handling shown in this article.
- taeric 3y agoAmusingly, folks usually seem hesitant to go with justified due to fears of "rivers" in the text. I can't claim that won't happen, but it seems largely overblown in concerns. Picking "affable" was an incredible nerd snipe! I would have split the syllables wrong on that, as I have yet to convince myself that I pronounce "fa" in the middle there. Similar problems come in when you have words that hyphenate differently depending on their use.
- chrismorgan 3y agoHyphenation points are a funny thing. People would commonly go for “aff-able” (aff-a-ble), but in such cases, you tend to get better results for reading by splitting in the middle of the repeated letter, or more generally the consonant sequence (af-fa-ble). I’m not certain if this is to do with phonemes (that is, the hyphenation matching how you speak) or to do with aiding with continuing (simply making it easier for you to pick up on the next line); I’ve thought about it a little, but not that much. There’s a similar but more aggressive form of the problem in lyrics on sheet music, where you’re declaring a hyphenation point between all syllables (though good engraving will avoid placing unnecessary hyphens), and people often break syllables incorrectly or suboptimally. Pulling up one score I noticed this with when I received it last week or so, some examples: strang-ers, runn-ing, en-em-y, anch-or, surr-end-er. I’d split (or hyphenate) them as stran-gers, run-ning, en-e-my, an-chor, sur-ren-der. I’m not certain if I’d always hyphenate and lyrics-syllable-break in the same places, but in the cases I’ve contemplated I would, though I had to think about sur-ren-der versus sur-rend-er for a moment. Also I can’t say exactly why my mind lighted on the word “affable” (it was the second word my mind came to, after I discarded “affiance”), but you made me think about it more deliberately and then I was curious to see what the first word in a dictionary would be. In my /usr/share/dict/words, the first (excluding proper nouns) is “affability”. (The last is “whiffs”, with “whiffletrees” as the last word I’d maybe consider hyphenating between the fs.)
- aidos 3y agoHave just skimmed it and this is a great intro to the weird and wonderful world of rendering text. It gets even more wild as you descend down into the myriad ways this information is specified within fonts. For bonus points dig into how fonts are embedded within PDFs. There’s definitely something wrong with me because I find the whole thing fascinating. Picking through the specs and existing open-source code feels more like an archaeological crusade than anything else.
- rizky05 3y ago[dead]
- rob74 3y agoI think this is true for most specs that are sufficiently old and sufficiently complex: there will be some parts that were specified, but never widely used, some parts that were once used, but are not relevant anymore, etc. etc. For an onlooker, these curiosities might seem fascinating. If you are setting out to implement the spec and don't have an infinite amount of time however, the multitude of corner cases, some important, some not (sometimes without an obvious way of distinguishing between the two) will get on your nerves pretty quickly...
- uncletaco 3y agoOne spec I've been diving into recently is emacs' skeleton auto typing engine. It has been in the codebase and documented for 30 years at this point but never saw widespread use. It was apparently was conceived as away to write not just snippets but also autoclose parenthesis and interactively fill in function or form data. I've been trying to figure out how I would go about using it as an alternative to tempel or yasnippet.
- plg 3y agoThis is one reason why I respect LaTeX so much, and why I can often spot documents written using MS Word a mile away. The idea that it doesn’t matter, that humans are not sensitive to or affected by these things, is plainly false. If I’m reading a long document, it matters. It may not matter ‘objectively’ but … I am human and it does affect me. The corollary is, if I’m writing a document that I know others will be judging (e.g. a research grant application or a scientific publication) I will absolutely 100% do what I can to make it a more pleasant experience for the reader, including using LaTeX for text rendering. I may change (or not) the default font, but for text rendering and all the tiny decisions about spacing, kerning, etc, I will trust LaTeX over MS Turd any day. Sure, “it shouldn’t matter”, a reader ought to judge my ideas and not be affected by this other stuff —- but, alas, the reader (for now ;) ) is human, and we are affected by this stuff.
- chrismorgan 3y ago> I may change (or not) the default font Please do. Computer Modern is awful; its optical sizing is all wrong, far too thin and with too much stroke thickness variation. On some common system and rendering configurations, it’s genuinely not far off illegible at common sizes. Its shapes were designed for printers that (a) had much higher resolution than most screens do, and (b) exhibited at least a moderate degree of glyph dilation (e.g. ink bleed). These days, it’s OK in print (still too thin, since printers are more precise in this way than they used to, which essentially makes everything a bit thinner than it used to be; but not too troublesomely), but really bad on screen for many renderers. Vaguely similar story on Courier New, which was badly digitised in a way that made it far thinner than it should be (due to not taking into account extra weight added by the typewriter’s ink ribbon), and so Windows includes all kinds of hacks to make it somewhat less intolerable (radical changes in hinting, and special-casing in ClearType), but half of those hacks don’t work any more (or work inconsistently) due to how things are done these days. Before anyone jumps on me saying they love Computer Modern and it looks fine to them: I’m focusing on the objective parts. On some screens with some software, it looks tolerable, despite clearly not matching most fonts’ ideas of what a regular weight should be. But on more configurations, it’s dreadful. I’m not talking about the rigid shapes and weird curly bits and such, which I also incidentally dislike, because that’s more subjective. But the weight stuff and the extreme stroke thickness variation stuff, that’s objective.
- PaulHoule 3y agoI’ve got this suspicion that sans-serif fonts are fashionable because something happening to of kerning, maybe somebody got a patent that prevents common software from implementing good automatic kerning. Today I can’t set a serif font for display in print and stand to look at the results if I don’t kern manually, this is true if I use Powerpoint, it is true if I use Adobe Illustrator and also on PC or Mac. I think most people just give up and use a sans-serif font. The thing is I don’t remember having the problem with desktop publishing in the 1990s, maybe I was less picky then but I’d say the results I get setting serif headlines today are so bad it is beyond being picky, also when I see posters that other people make people seem to just not use serif fonts anymore so I think they feel the same way.
- dredmorbius 3y agoSans-serif fonts are clearer on low-resolution displays. This includes both conventional desktop displays, with a low pixel density (96 dpi was the standard for CRTs and LCDs until Apple disrupted the field with "Retina" at ~200 dpi+, though in colour, which reduces effective density by 1/3. For mobile devices, DPI is typically high, but the overall display is small, and hence, individual glyphs tend to be small. On a large e-ink display, with monochrome density of 200--300 DPI, serif fonts are far easier and more comfortable to read for me, as the serifs provide additional cues as to the glyph shape and form. In particular, difficult-to-distinguish ASCII glyphs, such as O0, lI1, 5S, gq t+, ji, 4A, are much easier to distinguish. That describes my principle mobile device of the past 2+ years. I literally cannot see the individual pixels without a high-power magnifying glass. Though at lower-quality settings, font rendering remains somewhat glitchy. Those settings offer higher refresh rates and less flicker, so it's a bit of a trade-off. For DTP, the end result was almost always paper, and so long as you were using a laserprinter at 300+ dpi, serif would have been a preferred font choice. The turn to the Web saw a huge increase in use of sans fonts. Including, inexplicably and to my great annoyance, in published books, I suspect because the font carried connotations of "online" or "modern".
- PaulHoule 3y agoMyself I set a lot of posters, but when I was thinking about this I found a book printed in the 1990s that claimed to be set in “Palatino” and the kerning was so good I could swear the serif on the lowercase “r” was making love to the curve of an “s” next to it. That inspired me to take a look at the “Palatino” font on my PC and, sure enough, the “Palatino” on my PC looks nothing like that “Palatino” on that book.
- RcouF1uZ4gsC 3y agoI am convinced that this was one of the secrets to the USA’s rise to computing dominance: It had a language and alphabet that was amenable to a relatively simple encoding combined with a massive market of people that didn’t care if anything else worked. Thus even the very slow and limited memory computers of that time could actually do useful text manipulation (like sorting phone books) in a reasonable amount of time.
- dmytrish 3y agoMy perspective, from a non-native speaker: Latin alphabet has happened to overlap with majority of world capital and economy. That is all to it. It's a never-ending baroque timesink to learn English spelling properly, but it did not matter. Computing based on Chinese or Arabic would be a hindrance. Computing based on Hebrew, Cyrillic, Greek or Hindi would not be in any way (although for Slavic languages I do find flexibility of word order and plethora of word forms a hindrance, but it's a linguistic one).
- xeonmc 3y agoand Korean?
- chris_wot 3y agoActually, it took a long time to get to Unicode. Computers started with 7-but character encodings (and many rivals, including EBCDIC). Thankfully, text processing became pretty critical so they formed ASCII, when they went from 7-bit encodings to 8-bit encodings things improved, but then committees realised Latin dialects together had more than 255 characters… hence for some time we were all in code page hell. That’s why you got ISO-8599 - the most famous of which is ISO-8599-1. The genius of Unicode was they took all this junk and mapper into character planes, with ISO-8599-1 plastered into the BMP. I went into it in detail years ago: http://randomtechnicalstuff.blogspot.com/2009/05/unicode-and-oracle.html?m=1 http://randomtechnicalstuff.blogspot.com/2009/05/unicode-and...
- foobiekr 3y agoI think that’s right. In practice it just took time for memory to increase a bit and for statistical input methods to show up. Certainly actively bad writing systems delayed IT adoption. In fact, famously, the underlying theory of the Fifth Generation Conputing Project was the belief that advanced intelligent machines were needed to provide a working word processor for Japanese. No small amount of cultural supremacy was involved in that era, as well, now mostly forgotten.
- deleted 3y ago[deleted]
- Izkata 3y ago> If you’re in Safari or Edge, this might still look ok! If you’re in Firefox or Chrome, it looks awful, like this: I'm guessing this was written around when Edge still used its own rendering engine, and nowadays it looks like Firefox and Chrome.
- phkahler 3y ago>> The shape of a character depends on its neighbours: you cannot correctly draw text character-by-character. Well, for the languages used in newspapers for hundreds of years I'd say "sure you can". Just because we can do better than that with computers doesn't mean it's wrong. Sure text rendering is very complex, but this is also a first world problem.
- yakubin 3y agoLigatures were used centuries before computers were a thing. How it was done on printing presses: <https://www.dreamstime.com/royalty-free-stock-photo-ligature-letterpress-printing-blocks-image24494305 https://www.dreamstime.com/royalty-free-stock-photo-ligature...>. Also, this problem applies more to countries outside of the first world.
- chris_wot 3y agoThat works great with Latin based scripts. Now try it with Indic scripts: https://r12a.github.io/scripts/indic-overview/ https://r12a.github.io/scripts/indic-overview/
- minionnn 3y ago[dead]
- z3t4 3y agoI often see people complaining about "it's just text, why is it so slow to render", you can render a high res 3d environment faster then you can render the same screen full of text. That said, the text rendering engines of today are very optimized. So it's relative fast considering the amount of work. And humans are very sensible when it comes to text, a human can notice if one pixel has the wrong color in text, but wont notice it in a 3d scene.
- WhereIsTheTruth 3y agospecially google chrome (chromium and all the electron mess) on linux https://github.com/ryuukk/linux-improvements/blob/main/chromium.md https://github.com/ryuukk/linux-improvements/blob/main/chrom...
- cyclotron3k 3y agoOff-topic: Did any major OS successfully implement sub-pixel anti-aliasing on monitors that had been rotated by 90°?
- rpigab 3y agoNice, hadn't thought of that issue!
- int_19h 3y agoNot automatically, but in Linux at least you could set it FreeType for vertical RGB or BGR arrangement.
- Sunspark 3y agoWhich is useful, because the Steam Deck's panel is actually a portrait tablet RGB display rotated to landscape, which means when you are using the desktop mode of the Deck, the subpixel layout is actually V-BGR! Of course, if you connect it to an external monitor, you need to change that back to RGB since there is no "if monitor then pixel layout" function.
- ziml77 3y agoI wish they all supported arbitrary pixel layouts. My QD-OLED has a triangular layout that leads poor looking text no matter what ClearType settings are selected. (Between GDI and DirectWrite, the font rendering system in Windows is a mess). WOLED has similar issues due an RWBG layout.
- lostmsu 3y agoIn Windows you tweak sub-pixel anti-aliasing per monitor to whatever you like.
- ziml77 3y agoIf only. There's a ticket open with Microsoft's PowerToys to improve the anti-aliasing situation: https://github.com/microsoft/PowerToys/issues/25595 https://github.com/microsoft/PowerToys/issues/25595 And heres an explanation from the dev of MacType about how DirectWrite can cause different applications to perform text rendering differently from each other: https://github.com/snowie2000/mactype/wiki/DirectWrite-vs-GDI https://github.com/snowie2000/mactype/wiki/DirectWrite-vs-GD...
- chris_wot 3y agoWhen text rendering it goes through multiple stages. I remember going through the LibreOffice text rendering code. Text rendering basically goes through the following stages: * segment up text into paragraphs. You’d think this would be easy, but Unicode has a lot of seperators. Heck in html you have break and paragraph tags, but Unicode has about half a dozen things that can count as paragraph seperators. * parse text into style runs - each time text font, color slant, weight, or anything like this changes you add it to a seperate run * parsing the text into bidirectional runs - the text must work out the points at which it shifts text direction and place them into a new run at each shift of direction * you need to figure out how to reconcile the two types of runs into a single bidi and style run list. Do t forget that you might need to handle vertical text! And Japanese writing has ruby characters that are characters between columns. * fun bit of code - working out kashida length in Arabic. Took one of the real pros of the LO dev team to work out how to do this. Literally took them years! * you then most work out what font you are actually going to use - commonly known as the itemisation stage. This is a problem with Office suites when you don’t have the font installed. There is a complex font substitution and matching algorithm. Normally you get a stack of fonts to chose from and fallback to - everybody has their own font fallback algorithm. The PANOSE system is one such system when they literally take a bunch of text metrics and use distance algorithms to work out the best font to select. This is not universally adopted and most people have bolted on their own font selection stack, in general it’s some form of this. LibreOffice has a buggy matching algorithm that frankly doesn’t actually work due to some problems with logical operators and a running font match calculation metric they have baked in. At one point I did extensive unit testing around this in an attempt to document and show existing behaviour, I submitted a bunch of patches and tests piecemeal but they only decided to accept half of them because they kept changing how they wanted the patches submitted and then eventually someone who didn’t understand the logic point blank refused to accept the code - I just gave up at this point and the code remains bug riddled and unclear. On top of this, you need to take into account Unicode normalisation rules. Tricky. * now you deal with script mapping. Here you take each character (usually a code point) and work out the glyph to use and where to place it. You get script runs - some languages have straight forward codepoint to glyph rules, others less so. By breaking it into script runs it makes it far easier to work out this conversion. * now you get the shaping of the text into clusters. You’ll get situations where a glyph can be positioned and rendered in different ways. - in Latin-based languages and example is the “ff” - this can be two “f” characters but often it’s just a single character. It gets weirder with Indic characters which change based on the characters before and after… my mind was blown when I got to this point. This gets complex, and fast - luckily there are plenty of great quality shaping engines that handle this for you. Most open source apps use HalfBuzz, which gets better and better with each iteration. * now you take these text runs, and a lot of the job is done. However, paragraph separation is not line separation. You have a long enough paragraph of text and you must determine where to add breaks i lines of the text - basically word wrapping. Whilst this seems very simple, it’s not because then you get text hyphenation. This can vary based on language and script. A lot of this I worked out from reading LO code, but I did stumble onto an amazing primer here: https://raphlinus.github.io/text/2020/10/26/text-layout.html https://raphlinus.github.io/text/2020/10/26/text-layout.html The LO guys are pretty amazing, for the record. They are dealing with multiple platforms and using each of the platforms text rendering subsystems where possible. Often they start to standardise - but certainly it’s not an easy feat. Hats off to them!
- thworp 3y agoIf this article made you think about how to get better text rendering here are a couple of suggestions for devs: Many terminal emulators have greyscale anti-aliasing as an option or default, I know and have tested: - Windows Terminal (Windows) where you can set it in settings.json - kitty (linux) where it is default - You can probably get it globally in Linux through fontconfig settings, but I've never loked into that. If you're on a display at or above 200 ppi (4k at 27") you can also disable font anti-aliasing completely. The effect that has on letter shapes (esp. in Windows) is pretty striking. In chromium-based browsers and firefox there is a changing collection of about:config setting that control their Direct (edited to fix line breaks)
- lostmsu 3y ago> Windows Terminal (Windows) where you can set it in settings.json Wow thank you. I just switched from GrayScale antialiasing to ClearType, and the text quality seem to have improved. For the people who are interested, in UI it is in Settings -> Defaults -> Advanced.
- tayistay 3y agoText rendering doesn't hate you as much as article implies, AFAICT. 1. "Retina displays really don’t need [subpixel AA]" So eventually that hack will be needed less often. Apple has already disabled subpixel AA on macOS AFAICT. 2. Style changes mid-ligature: "I’m not aware of any super-reasonable cases where this happens." So that doesn't matter. 3. Shaping is farmed out to another library (Harfbuzz) 4. "Oh also, what does it mean to italicize or bold an emoji? Should you ignore those styles? Should you synthesize them? Who knows." ... yes, you should ignore those styles. 5. "some fonts don’t provide [bold, italic] stylings, and so you need a simple algorithmic way to do those effects." Perhaps this should be called "Text Rendering for Web Browsers Hates You." TextEdit on macOS simply doesn't allow you to italicize a font that doesn't have that style. Pages does nothing. Affinity won't allow it. 6. Mixing RTL and LTR in the same selection also seems like a browser problem, but I guess it could maybe happen elsewhere. 7. "Firefox and Chrome don’t do [perfect rendering with transparency] because it’s expensive and usually unnecessary for the major western languages." Reasonable choice on their part. Translucent text is just a bad idea. Probably could go on. It's a good discussion of the edge cases if you really need to support everything, I suppose.
- wffurr 3y ago>> because often you have control over the fonts you use Good luck with that if you want to support non-Latin scripts. Now you're shipping 100s of MBs of fonts with your app.
- tayistay 3y agoNo. You use the system fonts, which tend to be good
- birdsarentreal 3y ago9 out of 10 designers disagree
- wffurr 3y agoAnd then you no longer have control over the fonts you are using, and it’s worse if you want to write cross platform code. Web and iOS (as far as I can tell), don’t have any way of querying which fonts actually end up getting used to lay out text either, which makes rendering some entered text very difficult. I will probably end up having to ship entire text layout, font handling, and itemization stack in my wok just to do this correctly. It’s pretty frustrating down in the weeds of it.
- jokoon 3y agoI made several bitmap non-monospace font with an online tool, I'm quite happy with them. I use them in a few places, nothing more, and it's a bit difficult to make a good looking bitmap font, but it's an important part of having software that use simple things.
- barbariangrunge 3y agoI wrote an openGL font renderer once. It was a lot of fun. Bezier curves are such an elegant technique. The difference between what I wrote and what you'd use in a proper environment is pretty big, but I recommend it sometime. Fonts are pretty much just third or fourth degree beziers, plus a way to shade between the lines iirc (i may have my terminology wrong). Try it out sometime, I did my curves using tessellation shaders. Btw, you'll never find a better guide on beziers than here: https://pomax.github.io/bezierinfo/ https://pomax.github.io/bezierinfo/
- viggity 3y agoI work for a major foundry and fonts are ludicrously complex for a ton of reasons, but for the most part TTFs use quadratics, but our designers exclusively use cubics (postscript outlines) at design time and there is a conversion process to fit them as quadratics. On ultra lowend devices, having quadratics renders faster and the hinting works better, but if you're a designer consuming fonts, you should stick with OTFs as it is what the designer intended.
- Sunspark 3y agoI am an end-user, but I pay a lot of attention to rendering methods and hinting so I consider myself educated on the subject. Saying OTF is tricky.. because OTF file format can be postscript outlines OR truetype outlines and you don't know what is inside until you check. Sticking to what the designer intended is not always the right approach because in the end, it's about the end-user experience. If your output medium is a printing press, absolutely, every time what the designer intended. If your medium is a screen then the resolution and how the OS works matters. Using IBM Plex Sans on a standard definition screen as an example, the unhinted OTF will end up at some sizes squishing the letters into the wrong shapes and render lighter while the TTF will render the shapes better due to better grid fitting and darker. Basically and unfortunately, you have to test it for every environment and not assume that a given file is the best viewing experience. I dislike AA in general as a blurry mess (and disable it on my desktop, which yes, does mean I have to hand-pick fonts that were hinted for no-AA), and fail to understand why as more and more high-DPI screens are used that AA is still being used seems redundant and counter-intuitive.
- n6h6 3y agoWow, this answered some questions I've wondered about for a while yet never thought of looking up. Especially the section on subpixel antialiasing; very neat. I didn't know that was an intentional thing, I figured it was just something that happened on its own.
- knuckleheadsmif 3y agoThis brings back memories for me. I’m retired now but I helped develop and implemented many of the algorithms for multilingual text and Unicode back in the 80s first at Xerox and later at Apple. What this article misses is a lot of the complex formatting issues also with text rendering. For example even implementing tabs can be complex (remember that you have different tab types and conflicting constraints like keeping something decimal or left aligned but now having enough space to do that with out overlapping characters.) In languages like German if you have a computed or soft, hyphens the spelling of a word can change. Good paragraph breaking, widow orphan control, page breaking, computed header/footers that can change heights… are also complex issues that have to be dealt with. Back when I worked on this stuff we also had much slower computers that made things even more difficult so that you could type anywhere in the paragraph and still have responsive and correct output (you can’t delay formatting if the chargers change on context although some formatting can be delayed.)
- apples_oranges 3y agoNice, must have been really painful at times.. btw. I just noticed that when I zoom in Safari I get a pixel wise zoom and then, just a brief moment later, the fonts are re-rendered sharply. And this is I think the fastest laptop they sell maxed out in specs. So still there are optimisations in place to make text appear smoothly.
- thechao 3y agoThere's a group who're trying to get Mayan orthography into Unicode. Mayan has some challenges; the two I like best are: 1. Rendering is in pairs-of-columns, with a "Z" layout; and, 2. Glyphs have a concept of depth, and you must render subglyphs with the correct depth and occlusion to maintain meaning. Mayan also has the same feature as Egyptian Hieroglpyhs (but not the same extent) of multiple glyphs per word/logograph/letter. Here's a 2018 proposal, that covers all of this in more detail: https://www.unicode.org/L2/L2018/18038-mayan.pdf https://www.unicode.org/L2/L2018/18038-mayan.pdf
- qingcharles 3y agoI see this as the canary in the coal mine in case our Alien Overlords™ turn up. At least we'll be able to glyph what they're saying.
- chungy 3y agoThere's an irony in here, he's trying to demonstrate Firefox draws text wrong, but his example of "wrong" rendering doesn't match how Firefox actually renders it for me.
- AtNightWeCode 3y agoNot text rendering. Web. CSS still does not have proper support for kerning. Amazing how some things can be so bad for decades.
- einpoklum 3y agoIf you sympathize with the travails of people working on text rendering in applications, please consider supporting (among other projects): 1. The LibreOffice project (libreoffice.org), the free office application suite. This is where the rubber hits the road and developers deal with the extreme complexities of everything regarding text - shaping, styling, multi-object interaction, multi-language, you name it. And - they/we absolutely need donations to manage a project with > 200 million users: https://www.libreoffice.org/donate https://www.libreoffice.org/donate 2. harfbuzz (https://harfbuzz.github.io https://harfbuzz.github.io), and specifically Behdad Esfahood the main contributor. Although, TBH, I've not quite figured out whether you can donate to that or to him. At least star the project on GitHub I guess.
- heleninboodler 3y agoAdobe takes text rendering very seriously, because all their products were at one time primarily about targeting high-end printing applications. Back when I was there, there was one dude who was "the text renderer guy" for all the applications products (e.g. not printers and systems stuff) and when things got gnarly and an application had a text-related ship-blocker, they'd call him in and he'd patiently fix the ways in which some noobs had fucked up text in their particular application by not knowing a lot of the stuff that's in this blog post.
- OnlyMortal 3y agoThank you guys for your excellent RIP and Display Postscript on NeXT.
- heleninboodler 3y agoI accept these thanks on behalf of those who actually deserve it. :)
- hsn915 3y agoI'm not sure about noe but for a very long time Adobe Photoshop did not support complex bidirectional text.
- lannisterstark 3y ago"Hey so what do you do?" "I unfk other people's fonts" that's funny af.
- aatharuv 3y agoOne of the examples for Devanagari that is claimed to be broken is a nonsensical combination that really isn't supposed to work either as per any of the Devanagari script using languages or the Unicode standard. Devanagari (roughly speaking) has Consonants (with an inherent vowel A), dependent vowels signs added to a consonant, and independent vowel letters. And a few other signs for aspiration, nasalization, and cancelling the inherent vowel to combine consonants. न्हृे Starts with N (the consonant NA with the inherent vowel A cancelled and ends with the Devanagari Consonant HA with _two_ vowel signs added to it - DEVANAGARI VOWEL SIGN VOCALIC R and DEVANAGARI VOWEL SIGN E CV is the standard for Devanagari "syllables". When you do want to write two vowels after each other, you would write CV, and then another independent Vowel letter, so it would look like न्हृए (which would end with DEVANAGARI LETTER E instead of DEVANAGARI VOWEL SIGN E) *These are syllables as per the script definition rather than linguistic syllables. https://www.unicode.org/versions/Unicode15.0.0/UnicodeStandard-15.0.pdf https://www.unicode.org/versions/Unicode15.0.0/UnicodeStanda... section 12.1 has more details on the specifics of implementing Devanagari script, but not necessarily all of the conjunct forms between consonants, which are used especially when rendering Sanskrit.
- mnutt 3y agoI spent a good portion of this weekend tracking down a minor text rendering issue I was having in WebKit. It is mind-boggling just the mountains of code that go into text rendering, and pretty amazing that they have managed to keep such a complex system so performant. (the issue ended up involving harfbuzz, but it wasn't clear initially since it dealt with HTML whitespace collapse)
- jml7c5 3y agoOn a related note, there's a proposal for a Unicode project to improve complex script support in terminals: https://www.unicode.org/L2/L2023/23107-terminal-suppt.pdf https://www.unicode.org/L2/L2023/23107-terminal-suppt.pdf Many layout and shaping issues are present even in monospace text, at least if you want to do anything beyond ASCII.
- dahwolf 3y agoRecently at work we were testing a new font for the brand, won't mention the name, it's unimportant. We couldn't get it to look right on Windows. Only at a few select font sizes and weights did it look OK, at every other setup it looked like somebody took random bites out of the glyphs, meaning the individual characters had inconsistent weights. On Apple devices, we had no such issues. It seems their anti-aliasing algorithm intervenes more deeply, prioritizing a good and consistent result over theoretical integrity. You might have this same issue with thin fonts (this wasn't a thin issue) looking good on Mac and unreadable on Windows. A simple way to put it is that Mac renders font thicker (more black). I've faced the issues above for multiple custom fonts. It's why I no longer believe in the supposed best practice of fluid font sizes. I hardcode a limited set of font sizes and only those proven to look good across systems. Oh, another fun one is the glyphs of fonts having internal padding or asymmetric padding. Which messes with line heights and vertical centering. Or how about needing to tweak letter-spacing as characters overlap. I've come across all of this and more for some widely used fonts. Most are anything but ready-to-go.
- mjevans 3y agoOffhand, this might be HiDPI stuff. I haven't interacted with it myself, but upscaling for unaware applications is the first thing I would think to look at.
- mananaysiempre 3y agoThe Raster Tragedy[1] creates the impression that mangling font outlines to the extent that they’re utterly unrecognizable is the fundamental mode of operation of Windows font hinting. Granted, it’s a miracle you can meaningfully render vector fonts at 96 dpi at all. I just wish a closer look at that miracle did not reveal countless eyeballs and writhing tentacles. [1] http://rastertragedy.com/ http://rastertragedy.com/
- kalleboo 3y agoI think the typical summary of the difference in Mac and Windows font rendering is that the Mac tries to get the fonts to look closer to the printed page and Windows tries to fit the lines to the pixel grid more exactly. The Mac does it's job with no regard to the sharpness or fuzziness of the rendering - e.g. it will happily render you a "one-pixel line" as two gray lines just so that it ends up aligned beautifully with the rest of the letterform, but people without HiDPI displays will complain that this is "blurry". Windows tries to fit the pixels more exactly into the pixel grid of the display to ensure sharpness and a 90's sensibility of readability. The problem here is that for any font that is not hand-hinted well at the required size, this often falls flat on its face and looks terrible. And the art of font hinting seems to have gotten lost in the 90's.
- Asooka 3y agoAt this point I almost want to say that asking computers to render text in any way other than blitting pixmaps to a regular grid is a cruel treatment of the machine spirit. Joking aside, I am one of those people that completely disagree with subpixel anti-aliasing. I wish I could turn it off completely everywhere on Windows. I can do it on GNU/Linux and macOS never had that. It always looks wrong to me, regardless of monitor used or settings I pick, like the letters have colour fringing. I hated it when Microsoft introduced Clear Type in Windows XP and I still can't stand it.
- ggm 3y agoThis fragment: <<So [emoji of a dark woman] may literally appear as [emoji of a person] [emoji of dark brown] [emoji of the female symbol]>> the "may literally appear as" is actually EASIER TO UNDERSTAND its a person, of colour, who is a female. The single emoji is almost unviewable on my monitor at any precision and I could not ascribe feminine qualities to it, I almost was unable to ascribe personhood qualities. This to me is the central problem of emoji: They actually suck at conveying meaning, where alphabets and words do really really well.
- Calzifer 3y agoText rendering and transparency (not always together) torture me to often :( Quite recently I noticed that Java2D text rendering cheats the same way as Firefox and Chrome described in the overlapping example. The character Ø can be expressed in Unicode in two ways. Either as single code point U+00D8 or as an O + combining slash (U+0338). Since the first one is one code point Java2D renders it correctly with transparency but the second variant as two characters with notable overlap since Java is lazy on the 'combining' part.
- Rapzid 3y agoThinking of writing a text editor from scratch scratch? Oh, sweet summer child. Like seriously, font rednering is bonkers complicated.