23 ms·
Text Editing Hates You Too
- jen_h 7y agoReally appreciate this post. But I’ve gotta admit that I CTRL-Fed it for ligatures in bidirectional text and Czech accents and found no result...and couldn’t figure out whether I was happy or sad about that.
- dvt 7y agoThis post reminds me of one of Jon Skeet's seminal SO answers[1] on subtracting dates. Things we take for granted are hard. Dates are hard, text is hard, graphics are even harder. We stand on the shoulders of giants. It's important to stop and smell the roses every now and then. [1] https://stackoverflow.com/questions/6841333/why-is-subtracting-these-two-times-in-1927-giving-a-strange-result https://stackoverflow.com/questions/6841333/why-is-subtracti...
- AmericanChopper 7y agoI remember trying to debug this exact “issue” in a java project years ago. I poked around the library to figure out what was going on, which only raised more questions, before I eventually stumbled upon that exact SO thread. Turns out dates have an enormous amount of complexity.
- Joeri 7y agoDate and time are complex because they are social and historical constructs. Everything about them is arbitrary and exists because of arcane historical reasons. No regular system of organizing time maps cleanly onto the movement of heavenly bodies because those movements are not regular (hence: leap seconds). You can describe time in a way that is precise, or in a way that is intuitive, but not both, and woe to anyone trying to convert between the representations. Time zones are the worst because they are political tools, changed for political reasons, often with very little lead time. I think of them as a one way function mapping universal time onto the local political reality.
- ggm 7y agoI know its not the answer, but I always convert time into time in seconds since the epoch and perform the arithmetic. epoch time understands negative numbers: % date -ur -28501440617 Mon Oct 29 05:09:43 UTC 1066 %
- eesmith 7y agoIt's easy for me to write something in Russian, I just hire a Russian translator for me. That is, "convert time into time in seconds since the epoch" is the hard part. All you've done is make it someone else's problem, while the question Skeet answered asked specifically about the time difference between 1927-12-31 23:54:08 and 1927-12-31 23:54:09 in Singapore.
- ggm 7y agoYes. This is beautifully expressed. And 'hiring a russian interpreter' is almost always the right way to fix this kind of problem. POSIX would seem to me to be the russian interpreter of choice here. Is java a language which has no word for the colour green? in the sense where the eponymous russian interpreter has to say: "concept which cannot adequately be expressed in russian" or something The POSIX spec exists for this reason. This feels like a variation of "omg utf-8 is hard" which is why people embed utf handling in the language python2 -> python3
- eesmith 7y agoIt seems less than graceful to respond to articles which point out the existence and difficulty of problems by saying it's fixed. The entire point is to show how it's more difficult than many think, not elide the issues. What does POSIX have to do with anything? Eg, the tz database isn't POSIX.
- ggm 7y agoThe POSIX spec of date time specified conversions into a neutral form. Maybe it's C centric thinking but it felt like a problem which has existed and been solved by writing a 'russian interpreter' solution: convert the values on the edge to a canonical form and use functions which manipulate the canonical form. The TZ files i only mention because within the limit of modern epoch Times and dates they do a pretty good job of handling the minutia of what a local time means in any given period. I'm sorry if I came across as churlish. Do you feel this is a problem of concept in abstract, or in java, or in java implementation or what? I feel it's a problem which I solve, by standing on the shoulders of smarter people and using the code they wrote. I don't reimplement the wheel because I distrust my own wheelmaking skills. I'm sure there are heaps of non-trivial corner cases I don't have to deal with and I do not mean to minimise work done to explain a problem.
- gambler 7y ago> Things we take for granted are hard. Dates are hard, text is hard, graphics are even harder. We stand on the shoulders of giants. More like we stand on piles of broken shit. A lot of "hard" things in programming have very little fundamental complexity. They are hard because at some point in time someone made a compromise that no longer is relevant or makes any sense. It is possible to redesign those things to be simple, but modern developers are conditioned to fight this at every step. The idea that the first thing you should try to do when faced with extraordinary complexity is to bypass it is completely alien to people whose daily jobs consist of continuously banging their heads against convoluted frameworks and deficient APIs. Time handling is a great example. Until you get into general relativity, time itself is extremely simple and uniform. But human representations of time are extremely convoluted, because historically they are based on astronomy. Here is the kicker: 99.99% of all software has nothing to do with astronomy, and yet people chose to inject those complexities in core structures that track time. (For example, Unix timestamps have leap seconds, which is a source of innumerable bugs and ridiculous convolutions in code to prevent those bugs.)
- stnmtn 7y agoAre you arguing for trying to create a new method of tracking time instead of just including moment.js?
- hrktb 7y agoI read parent’s comment as a plea to be simple where it’s allowed. For instance use straight timestamps where there is no need for calendar manipulations. I am sympathic to this argument and I’d also totally join a movement to ignore DST or uncouple years from earth rotation
- m45t3r 7y agoThis post sounds kinda like a satire: "we have problem X, did you try something.js to fix it?"
- rovolo 7y agoIf by 'based on astronomy' you mean 'based on the sun' because the sun is a star, then I agree. But, plenty of software has lots to do with the sun because they have to interact with people, and people like to be synchronized with the sun. If you don't want software to interact with humans, or you only care about duration rather than human datetime, then go ahead and use UT1. However, as long as your interacting with people, you're going to need to fix people first. You should start with these primers if you're going to fix how humans interact with time: - https://qntm.org/abolish https://qntm.org/abolish - https://qntm.org/continuous https://qntm.org/continuous
- Thorentis 7y agoInteresting! I hadn't thought this deeply about text editing before. I disagree that the "Bad #3" example in the "Emoji Modifiers" section is actually bad though - it's the outcome I would expect of an editor.
- mjevans 7y agoIt's probably bad because there isn't an additional kludge added: decomposition of combined character entities for editing. This would involve a concept of sub-character (code-point) 'tabs' (which ideally would be distinct UTF-8 entities).
- jakear 7y agoVSCode does the "correct" behavior of Bad #3, but doesn't even need to do the "bad" part about pushing the bytewise carat position around, as it logically maintains two characters, but visually coalesces both the middle position and the front together. Wonder why it wasn't mentioned.
- Mathnerd314 7y agoAnother option would be that deleting the 'a' doesn't completely delete it but instead replaces it with a zero-width space or zero-width non-joiner, so it looks like Bad #1 (but is Unicode-compliant) and hitting delete again gives Bad #3. The whole example is a bit contrived though, nobody is going to enter in a skin tone modifier character by hand in daily use. They'll select an appropriately-colored emoji.
- jakear 7y agoIMO, having delete trigger a zero-width space insertion would be the worst option. I mention in a sibling comment that VSCode gets around this by having two separate logical carat positions combined into a single visual position. So the byte offset of the cursor changes as expected, while still maintaining "Unicode correctness", for whatever thats worth.
- pcr910303 7y agoFor people who are interested in the HN post mentioned in the article: here it is. Text Rendering Hates You[0] [0]: https://news.ycombinator.com/item?id=21105625 https://news.ycombinator.com/item?id=21105625
- throwaway_n 7y agoI once tried to customize a rich text editor to support SVG text effects (like https://www.smashingmagazine.com/2015/05/why-the-svg-filter-is-awesome/ https://www.smashingmagazine.com/2015/05/why-the-svg-filter-...) and damn near lost my mind. If I had known that you can write an entire book on just text layout I wouldn't have bothered: https://www.oreilly.com/library/view/svg-text-layout/9781491933817/ https://www.oreilly.com/library/view/svg-text-layout/9781491...
- arpa 7y agoTo be honest, I can't help but think that we're overcomplicating things for ourselves. I mean: think about a typewriter, hell, even typesetting. And now let's consider modern text editing. It's crazy! And yet, the only thing I need is something even simpler than vi: ability to move between the lines when doing cat - > file. That being said, i do appreciate the effort the good people of text editing make.
- saagarjha 7y agoRight, but you probably mostly deal with a documents written in a language with a Latin script. Those who don't or can't do this need better.
- arpa 7y agoWhat fits all, fits none. Do one thing and do it well. Meaning: those who need it, need something else. Also, original point about typesetting stands, even in Thai case. Also i find it ironic that we care a lot about let's say caret movement, when whole classes of devices are unavailable for the blind.
- mntmoss 7y agoI agree with the ironic sentiment. We've pushed up the complexity of a lot of features because we can, not because the value is there. Unicode itself is a great example of such: it's a dumping ground for "everything that resembles text" - and while it's easy to consume and there is some benefit to be had, it's also hard to comprehend on any level beyond the very basics. It's quite a ways away from the early telegraphy encoding systems. With respect to selection, selection is key to all forms of interactive editing because without it you don't have interactivity, you have a linear workflow. As such it does deserve some respect as a fundamental feature, even if it's not critical to fix every last bug. Even if you are working via screen reading and voice to text, you can still benefit from having selection tools available to you.
- tsimionescu 7y agoIt depends a lot on your goal. If your goal is, for example, to edit a book, you often need much more than a typewriter can give you. You need figures, equations, centralized text, footnotes, rich text styles, tables. If your goal is even more ambitious, say, publishing the same book online, such that others can use the contents of the book and select it, look it up, share it, edit it etc., you need many of the above to be understood by the text engine, so you end up with shaping, support for all human scripts with all their accumulated idiosyncrasies etc. Without any of this, you can barely have an academic life for example on the internet,or engineering disciplines and so on.
- Aperocky 7y agoAs almost an exclusive vim user (large scale java doesn't work very well on vim), I find it funny when text editing is a huge pain - if there was pain for me that's about a month of overall time sunk into getting really good at vim. After the transition text editing has mostly been a joy.
- wonnage 7y agoThe article isn't about how to use a text editor
- defanor 7y agoWhile spending a lot of time with monospace fonts and mostly ASCII characters (programming and writing, terminal emulators, IRC/mail/MUDs/feeds in Emacs) and working on a hobby project involving text rendering and selection (with potentially proportional fonts and Unicode), I keep wondering whether it's even worth all the trouble. In addition to what's mentioned in those articles, on larger documents rendering speed also matters; word cache (in addition to mentioned glyph cache) helps somewhat, but complicates other things even further, and calculating total document height (for scrolling) still requires to calculate all the positions, wrap the lines, etc, which in turn requires to render all the text first. Perhaps one can introduce yet another mechanism to only estimate that at first, and then elaborate the estimation in background, but that's yet another opportunity for bugs to creep in (and additional complication, of course). On the contrast, the programs that just use character grids (and perhaps don't handle bidirectional texts well) can be both faster and simpler, while still are quite capable of rendering texts, even with various styles/decorations and occasional images, in many languages. There seems to be plenty of accidental complexity even in handwriting, but once it is used for computing and text manipulations are involved, the task appears to be much harder than it has to be.
- kccqzy 7y agoHave fun in your anglocentric world then. I don't think you have sympathy towards users for whom monospaced fonts do not work for their language, ASCII can't encode all the code points in their language. Rendering speed hasn't mattered since years ago. Browsers can render megabytes of pure text (hours of reading material) in seconds. LaTeX can typeset hundreds of pages in seconds.
- defanor 7y ago> I don't think you have sympathy towards users for whom monospaced fonts do not work for their language, I do, that's one of the reasons I'm trying to handle Unicode and proportional fonts in the first place. Given the circumstances, it is indeed preferable to do so, but given that it's basically encoding of approximated sounds (well, potentially information in general) on a 2D plane, it is a rather complicated (both code-wise and computations-wise) way to achieve such a task. > Rendering speed hasn't mattered since years ago. Browsers can render megabytes of pure text (hours of reading material) in seconds. Chrome only started using "complex path" for English texts a few years ago, employing mentioned word cache. Rendering megabytes of text in seconds is indeed achievable (especially if the words are cached, and it's not megabytes of integer sequences), but still complicates the rendering, and still slower than to open and navigate through the same megabytes of text in, say, less or vim, or even links -g. It's not an issue if you don't have to do it (or to use a program where it wasn't optimized well), but if you do, it is. Edit: added a reply to the second part.
- jakear 7y agoNot mentioned in the article, but an interesting point of difficulty in editors is the character transposition command (Ctrl-T on most MacOS applications, stemming from emacs I imagine): TextEdit and VSCode are the only editors I know of to handle transposing emoji in a reasonable way. VSCode will separate the emoji from the color modifier. TextEdit keeps them together, other editors I know of corrupt the data by slicing surrogate pairs or just no-op the command.
- Pxtl 7y agoWait, there are people who use the ctrl-T command? That thing is the bane of my existence because I have muscle-memory for "new tab" in firefox but when I do it in a text editor for "new file" out of habit it just jumbles text.
- jakear 7y agoI’m a pretty bad typist, so it’s helpful for when I need to fix a word and I can do it in one keystroke. Also my new tab binding is command-t.
- pcr910303 7y agoAs a person living with the CJK languages (I’m specifically a Korean), I find that some of these problems are prominent, even in the 21th century. There are an excessive amount of programs that conflate key presses & text input, and ones that don’t consider input methods. I use macOS & Linux, and while the default text handling system called Cocoa Text System in macOS handles input methods well, almost all applications that implement it’s own, like big apps like Eclipse and Firefox, don’t get this right. On Linux, it’s terrifying; I’ve never seen any app that allows input systems to work naturally, and after a week of use you get used to pressing space & backspace after finishing every Hangul word. The Unix-style composability they want (apps should work whether or not input methods are used - and looks like Linux users that use Latin characters don’t use any input methods (opposed to macOS where Latin characters are input by a Latin input system), so looks like this state will persist. About the emoticons, I’m not that concerned with that since most (if not all) users won’t really input the color modifier separately (or even encounter files that have a separate one), so you can just select a sensible behavior like the #2 or 3 or 4. Users who understand the color modifiers, and other Unicode fiasco will understand what is happening under the hood, and ones that don’t will just think the file is broken and none of the behaviors will make sense, whatever you do.
- hsivonen 7y ago> I use macOS & Linux, and while the default text handling system called Cocoa Text System in macOS handles input methods well, almost all applications that implement it’s own, like big apps like Eclipse and Firefox, don’t get this right. What specific IME problems do you have with Firefox on Mac? > On Linux, it’s terrifying; I’ve never seen any app that allows input systems to work naturally, and after a week of use you get used to pressing space & backspace after finishing every Hangul word. Do you mean you have to press space twice and erase the second space? With IBus?
- dfcowell 7y agoThere are no spaces between words in Chinese or Japanese. Pressing space confirms the current selection in the Japanese IME, which is expected behavior. Where some Linux implementations get it wrong is they also insert a space after the word, meaning the user has to select the desired word in the IME with the space bar and then remove the erroneous inserted space. Edit: Correction based on feedback below. Previously stated that Hangul does not have spaces.
- saagarjha 7y agoI once wrote a simple LaTeX renderer for a project, and boy was it hard. And all I had to do was support a subset of the whole thing! Even with clear biases towards left-to-right languages and just a subset of the LaTex, it was a nightmare: you'd do something, but realize that it wouldn't render correctly in a certain situation, rewrite the code to include more context or do another layout pass, and hit another issue somewhere else. It was slow work, complicated to debug, and utterly stymied many refactoring efforts. I can't even imagine how much work must go into doing this quickly and correctly for a language that's much more complex…
- indentit 7y agoThis is the first time I've ever seen it explained how input methods allow users to write in languages like Chinese using the Latin letters A-Z: phonetics. I had always wondered about this, it seemed like some arcane magic to me lol, how people knew what to type. It's now clear that for someone like me - whom only understands a few Latin based languages - a super comprehensive tutorial would be required if one wanted to understand enough about other languages to be able to work on a text editor etc. I guess ideally all teams working on such projects would have an experienced team member whose native language isn't Latin based.
- weeb_throwaway 7y agoThis barely scratches the complexity of IMEs. For example Hong Kong mainly speak a different dialect of chinese, Cantonese. So pinyin (mandarin phonetics) is useless. Most people type with Cangjie instead: https://en.wikipedia.org/wiki/Simplified_Cangjie https://en.wikipedia.org/wiki/Simplified_Cangjie
- alpaca128 7y agoYou can try out how it works with Google Translate, for languages like Chinese, Japanese etc. they have a button in the bottom right corner of the text field where you can toggle to an IME input system. It's pretty intuitive, basically you just write the word phonetically and then press space, and it'll automatically replace the characters and let you choose alternatives. Smartphones have similar systems but with different layouts.
- hyh1048576 7y agoChinese here. To be a little picky, phonetics is just one of the several methods to input Chinese. (Pinyin is the dominant one, and yes it is phonetics, other ways uses shapes of the characters and maybe a lot faster, see https://en.wikipedia.org/wiki/Wubi_method https://en.wikipedia.org/wiki/Wubi_method for example.) Speaking of input methods I always with there are good English input methods, it will be useful too! For example if user enters "compre" the input methods goes 1. comprehension 2. comprehensive (with an order depending on the conditional probability -- there is some interesting math behind input prediction, and with their help input can be a lot smooth, however I only see it used on phones not computers.)
- be5invis 7y agoLet me count all the features that a proper rich text box should support: · Input methods, including... · Composing IME · Speech · Handwriting · Emoji picker · Complex text rendering, including... · Bidi · Shaping · Complex-script-aware editing · Complex-script-aware find, replace, etc. · Accessibility And what they can choose to support: · Advanced typography, including... · About line breaking... · Hyphenation · Optimized line break (Knuth-plass, etc.) · About microtypography · OpenType features (ligatures, etc.) · OpenType variations · Multiple master fonts · Proper font fallback ← It is not a simple lookup-at-each-character process · Advanced Middle East features, including... · Kashida · Advanced Far East features, including... · Kinsoku Shori · Auto space insertion · Kumimoji · Warichu · Ruby · Proper vertical layout, including... · Yoko-in-Tate · Inline objects and paragraph-like objects, including... · Images · Hyperlinks · Math equations ← This is really hard.
- poooood 7y agoThis guy went to the Harvard of the Midwest. I feel it in my soul.
- poooood 7y agoThis guy seems p bright. Prolly went to Harvard (of the Midwest).
- Mikhail_Edoshin 7y ago"Hate" is a strange choice of word to describe what is merely a complex subject. Lots of different models (emoji + skin tone modifiers is not a bad model in isolation, very much like a letter + combining accent) are developed independently and then when we test them with the rest of the players on the field, they don't integrate too well. Yes, they don't, but what would you expect? :) And of course decisions made by different developers will vary in logic, consistency, and completeness. This is normal state for the integration part.
- TeMPOraL 7y agoIt's a phrase. You approach a seemingly simple problem that reveals itself to be a fractal of complexity - at some point you start to feel as if the problem itself had agency and it wanted to make you miserable.
- dandare 7y agoTangential: Microsoft's automatic selection of whole words and sentences is the single biggest productivity drain in my professional career (all my attempts to disable it failed so far).
- spongeb00b 7y agoFor anyone having to deal with this on Microsoft Word on the Mac it can be disabled through Word’s preferences.
- dandare 7y agoThe setting does not apply to Outlook when writing/editing emails, only when reading emails. Then, from my experience, the setting will be reverted with the next Word update :(.
- pbhjpbhj 7y agoGah, I've just started with MS Word (after a long hiatus) and was wondering why it always grabs surrounding punctuation into the selection (even when the punctuation 'points' the wrong way, it seems).
- jraph 7y agoI remember being annoyed by that years ago the few times I used Windows. I am used to triple-click+drag to achieve this when I need it. When I actually don't want to select words, it gets in the way. One can probably use the keyboard for precise things like that though. I can see triple-clicking+drag being tricky and that most people probably want to select entire words most of the time anyway, so I don't know what the best solution is. I'd argue that selections need handles like on mobile operating systems so it can be edited afterwards.
- z3t4 7y agoI've naively implemented text rendering, selection, etc from scratch in a text editor, and if you only support monospace, it's pretty simple. The hardest part in software is to not implement features. A trick is to only implement the complex features for the users that need it - in another VCS branch. Or you make an abstraction layer. But abstraction layers are hard, and if they are not air tight (non leaky) you are worse off, because then the developer need to know both layers. Then there is the third strategy, which I would call an anti-pattern, that when enough features have built up, in order to get rid of the relics you rewrite it from scratch, preferably in a new framework. I also believe most advancements in science are accidental, like penicillin, and lexically scoped modules - so if we do not do things, like new frameworks, and re-invent the wheel, the CS field wouldn't advance.
- NickGerleman 7y agoI until very recently worked on Microsoft Word. The whole problem gets even more complicated when you add support for richer content like formatting, images, comments, etcetera. Sprinkle in requirements for things like three-way merge, simultaneous editing from multiple authors, undo behavior on top of that, and the amount of cross-cutting complexity for something seemingly simple can be absolutely astonishing.
- edwinyzh 7y agoHey Nick, I'm glad to hear about that you are in the Word team:) I'm working on a document manager and static site builder based on Mirosoft Word (https://docxmanager.com/ https://docxmanager.com/) (The upcoming version will have a more user-friendly tabbed UI)
- faeyanpiraat 7y agoInteresting concept, I've started working on something similar, but I like working in plaintext, soo: Started creating a custom text editor, then I ran into all kinds of weird difficulties like the original article mentions. Ended up integrating Notepad++, like you integrate Word.
- pitaj 7y agoHe specifically said that he is not longer on the Word team.
- alkonaut 7y agoHow much is poor text input eroding the use of non-latin script worldwide? I.e. how many just give up instead of using poor input? As an example: I use a Swedish keyboard. If I'm in a situation where ö is missing from the input (because the app is stuck in an en-US keyboard layout say, so I have to hit a modifier ¨+o to type an ö) then there is a very good chance my communication with my colleagues would just naturally be in english instead. I'd sigh and just give up using it. Rather than using an o for my ö I'd just type it in english instead. Is this a thing in e.g. asia, israel, or the arab world? Do kids that speak english communicate more in english in cases where the input doesn't let them communicate easily using their preferred script? Are there new "hybrid" languages popping up in electronic communication where languages that use non-latin scrpit are written in latin in e.g. text messages? (You could argue that emoji is just that but the other way around I suppose)
- bonoboTP 7y agoFor Hungarian, people stuck with an English layout usually just leave off the diacritical marks (áéíóöőúüű -> aeiooouuu). While this theoretically leaves some of the meaning ambiguous, and pedants can craft examples that may be ambiguous even with context, it works well enough in practice. Switching to English is way overblown a reaction. Two Hungarians chatting in English (unless there are non-Hungarian speakers involved) seems extremely weird to me. It may be partially that English is really foreign for us, while it's pretty close linguistically to Swedish, both being Germanic.
- Vinnl 7y ago> It may be partially that English is really foreign for us, while it's pretty close linguistically to Swedish, both being Germanic. That wouldn't explain it, because we do the same in Dutch: just leave off the diacritics. Diacritics are pretty rare in Dutch, though.
- alkonaut 7y agoYeah in swedish leaving them off isn't working. The diacritics aren't for accentuation, they are distinct letters. An ö is as different from o as u and e are.
- lykahb 7y agoI did quite a lot of hacking around the text selection when extending CKEditor for XML editing. In my opinion there are several reasons why text input and selection are so difficult: 1. The behavior is complex. 2. The behavior resists being formally defined. It differs across the user interfaces, such as URL bar, rich text editor, terminal, etc. 3. The requirements are evolving. The recent big change - the emojis - is still not universally supported. 4. The reasons above tend to add accidental complexity to the API that will only grow over time. I feel like the best approach would be to create a mathematical theory describing the text input semantics. This approach worked quite well for other complicated areas in CS, such as concurrency or memory management.
- greggman2 7y agoa common issue in browsers are pages looking for the ESC key. I'm in some modal dialog on a page. I'm typing CJK in the IME. I press ESC to exit the IME or to cancel a conversion, the IME exits but the page gets the ESC key and closes the model dialog. Any text I had entered to that point is lost.
- aasasd 7y agoEvery time I hear someone is implementing even a part of text editing, I begin to snigger. (Ahem Workflowy cough.) I'll gladly put these two articles in my reading list, to go over some evening while reclined comfortably and sipping wine—so I then can laugh and slap my knees even harder when a new text editor is mentioned.
- danShumway 7y agoI've lightly advocated for a while that emoji shouldn't be part of the Unicode standard at all. I'm sure there are things some advantages, I'm sure there are other considerations I'm not thinking of, but it just seems like a really bad idea to stuff the Unicode standard. I don't know the official name or who came up with it, but I use Slack's entry format exclusively in every application. :thumbsup: :pelipper_blushing: :angry_cat: :apple: If the application can detect that as an emoji and swap it out, fine. If it can't, I don't change my format. My preference would be if applications left emoji in that format, and just rendered them differently at display time. The advantage of having emoji just be a purely clientside rendering feature, and behind the scenes all fall back to normal text is: a) they can be easily aliased across multiple languages (:cat: :gato:) b) if you paste an emoji into an application that doesn't support them, you don't get an unrecognizable character. Progressive enhancement! c) it's accessible when copied and pasted into a pure UTF-8 text format. It's just better blind-accessible in general. d) it's more forward compatible. I can use :cthulhu: right now without waiting for it to get added to the standard. e) get rid of modifiers. Like, seriously, just get rid of them. Emoji aren't programmatically generated, you still need to draw one image for each modifier combination, and you still need to program support for each one. So, what's the advantage of using modifiers over just adding multiple glyphs? They're just there to save space in the character list, which is only a problem because emoji are in Unicode. :smile: :fake_smile: f) better support for custom emoji in general. Basically every platform has custom emoji, and it's weird because half of your emoji are standardized and half aren't. And then whenever new emoji get added to the standard, if they conflict with your custom emoji your app breaks. I would hesitate to standardize emoji at all, beyond having a consortium that says, "this is what :thumbsup: means, you can extend on top of this as you see fit." It feels like extra complexity for no benefit other than, "we need a standard".
- philplckthun 7y agoI think this works great for apps like Slack, in user-land so to speak, but isn't realistic for the Unicode standard, not only because these entry formats are in English. Modifiers and combinators aren't exclusive to emojis, but apply to all kinds of glyphs in other languages and writing systems as well. Arabic script even has some common ligatures for common expressions. A lot of complexity simply doesn't stem from emoji in Unicode, a lot of the complexity comes from all the writing systems that Unicode supports. Admittedly, emoji are kind of an oddball addition to Unicode, but they're by far not the most complex part of it.
- derefr 7y ago> Our carets will need an extra bit that tells them which line to tend towards. Most systems call this bit “affinity”. Is it just me, or does the author disprove themselves with the figure in the same section? I feel like the best possible solution is the one depicted: that you just see the cursor split between both lines.
- greenshackle2 7y agoIt's not a split cursor, it's two cursors, one on each line.
- derefr 7y agoI know, but two cursors sort of looks like a split cursor, and you could visually tweak cursor rendering to make “one cursor split between two lines” a visually-distinct case from having actual multiple cursors (because some text editors do indeed support multiple cursors.) What I’m saying is that I’d prefer a text-edit control that gives you a visual indicator for “one cursor split between two lines”, to one that pretends the cursor is on one line or the other, when it really will act with the navigation semantics of being split between two lines.
- greenshackle2 7y agoAh right, I misunderstood you. Yeah that sounds like a sensible solution.
- bandrami 7y agoPersonally I've always thought RTL switches in an editor were a mistake. An editor displays and lets you edit a character stream, and those characters do not need to be in their final spatially rendered position for editing.
- billconan 7y agothese type of things seem not to be documented anywhere. If I want to start from scratch, what can I use as a reference?
- phil9987 7y agoThis is incredible, thanks for sharing. I had the exact same thoughts the author is describing in the beginning - how hard can it be? It turns out super hard. And one shouldn't forget that it is one of the few input methods to our computers. With voice not working properly and drawing on the mouse pad / touchscreen being too slow I would argue it is still the number one input method. So the expectations towards UX are extremely high and there should be no faults. I admire anybody who works on some kind of text input mechanisms from this day on.
- 6510 7y agoOn c64 the text is mono-spaced and empty areas are filled with spaces. I didn't appreciate the genius of that until now.
- tomaszs 7y agoSome years ago i was working on a project to autocorrect text with high level written text recognition. I have used standard text editor. Man, i thought 1 day about rewriting text field. But angels stopped me from it. People dont know how hard it is!
- azhenley 7y agoThis is one of the best things I have seen on HN in a long time! It really makes me want to go work for a company with a text editor product or to teach a course on implementing simple text editors.
- knolax 7y agoThe problem the author states in Vim is overstated. All you need is to write two functions that switch the input method to English when entering normal mode and back to your original input method when entering insert/replace mode. It's only 24 lines of vimscript including whitespaces and comments. Out of all the TUI programs out there, Vim with it's modal design is probably the one program least affected by a conflation between keypresses and text input.
- big_chungus 7y agoThis is not really cross-platform. This is also not cross-application, so it's often necessary to do os-level, hence prior problems. At least for Spanish there are a limited number of non-ascii chars, so I jus bind directly to a modifier.