4 ms·
One interesting/fun problem with text layout is that: width(A) + width(B) != width(A+B) ...which some basic text layout engines assume. The post touches on
by bfgeek 4y ago
One interesting/fun problem with text layout is that:
width(A) + width(B) != width(A+B)
...which some basic text layout engines assume. The post touches on this, and is one reason why line-breaking is so difficult.
If you add some text to a line, the width of the line may be longer or shorter(!) than if you measured (shaped) the parts separately.
This occurs for many reasons, the post mentions splitting a ligature, can also happen with kerning (spaces can have kerning applied for example).
E.g. many text layout engines will incrementally "add" to a line. However to do this correctly for all cases you need to re-shape (measure) the entire line (or from a known "safe to break" point in the shape result see: https://github.com/harfbuzz/harfbuzz/issues/224 https://github.com/harfbuzz/harfbuzz/issues/224 ). This is typically fine for latin cases, but begins to become slow for more complex scripts (Thai).
Chromium kinda works backwards. It tries to fit everything in the paragraph on a single line, shapes (measures), then finds a line-break opportunity within the (potentially large) shape result which would "fit" on the line without re-shaping.
Then given that line-break will reshape the entire line (using the "safe to break" API), (and remaining content for the next line), see if it will fit and repeats the process[1].
By always "removing" content from the line you end up with a correct implementation which works for all the crazy things which can happen with text rendering.
[1] The content in the line may become bigger after taking a line-break, hence this needs to happen in a loop until there is one "unbreakable" piece of content in the line. This loop typically only goes through one iteration.
- amelius 4y agoDidn't Knuth solve this problem in his Computers & Typesetting books?
- raphlinus 4y agoNo. Most line breaks in English are at spaces, and the Computer Modern fonts don't have any complex shaping behavior across spaces (it's mostly complex scripts such as Nastaliq that have this behavior). I haven't checked to see whether it takes into account the kerning value with the hyphen added. The first system I know of that does solve this very well is DirectWrite. Sergey Malkin describes the solution here: https://github.com/harfbuzz/harfbuzz/issues/1463#issuecomment-505592189 https://github.com/harfbuzz/harfbuzz/issues/1463#issuecommen... It sounds like Chromium has since also adapted the "safe to break" API. Great to see people engaging these hard problems!
- taeric 4y agoI doubt Knuth would claim for "solving" the problem, but I do get the impression that it is a solution. In particular, it does tackle hyphenation of words. Is where I learned of the re-cord versus rec-ord problem. I'm curious what you mean regarding complex shapes across spaces. If it is related to the displayed bug regarding how words like "office" would be split, I don't think TeX ever had the bug as described. Though, I could be wrong, easily. (Incidentally, what a fun bug. Kudos to who found that one!) Edit: I also had the impression that TeX would adjust intercharacter spacing as a whole to keep things looking a bit more uniform. Essentially, the goal was to act a lot like a human would when setting out the characters. If needed, you could adjust spacing on all characters within a margin of "not going to be noticed" so that you could expand a line of text to take up the full width. Without just having more space between a few words.
- bfgeek 4y agoA lot of folks think that Knuth-Plass is the optimal solution for good looking text, for all text. It really only considers Latin, and even then has restrictions. Some fonts have the space character in the kerning table, and possibly (although super rare) in the "ligature" table (gsub). (https://www.sansbullshitsans.com/ https://www.sansbullshitsans.com/ is an example of a font with a space in gsub, "paradigm shift" maps to a single glyph). If you want to go really deep on how complex some scripts are take a look at: https://r12a.github.io/scripts/arabic/arb.html#shaping https://r12a.github.io/scripts/arabic/arb.html#shaping
- taeric 4y agoFair, a lot of folks do treat it with probably more praise than makes sense. I'm sure I'm one of those people. :D That said, my understanding is that there really was no "optimal" for the best way to arrange text. You can, effectively, make an objective function that you can optimize on; but there is no global "this will make text look good" algorithm. Indeed, even ligatures are... debatable in utility. I personally like them; but I would scoff at any claims for any objective superiority of them. They are fun and a bit of a "flex" for laying out text on a computer. Any other claim is going to be tough to hold up.