19 ms·
CRLF is obsolete and should be abolished
- steeeeeve 2y agoOnce in a long while, people start looking into things and wanting them to make sense. Like hey - why don't we start using the field separator and record separator characters when exporting/importing data. But then you end up realizing that even when you are right, the energy it would take to push a change like that is astounding. Those who successfully create an RFC and find a way push it through all the way to it becoming a standard are admirable people.
- nuancebydefault 2y agoIt is a bit like saying, let's forget about AM radio since FM radio is much better. Oh wait forget FM since DAB exists. Oh wait forget about broadcast protocols since everything is point to point or multicast now. The reality is that existing protocols CANNOT be changed. Only new versions are released and the old ones (which might rely on CRLF) will never die.
- eska 2y agoI’ll join your revolution under the condition that I’m allowed to ignore all lines of text that don’t end in a newline :) (POSIX)
- justin66 2y agoNice. I think that's the most energized I've seen Richard Hipp on a topic.
- michaelmior 2y ago> various protocols (HTTP, SMTP, CSV) still "require" CRLF at the end of each line What would be the benefit to updating legacy protocols to just use NL? You save a handful of bits at the expense of a lot of potential bugs. HTTP/1(.1) is mostly replaced by HTTP/2 and later by now anyway. Sure, it makes sense not to require CRLF with any new protocols, but it doesn't seem worth updating legacy things. > Even if an established protocol (HTTP, SMTP, CSV, FTP) technically requires CRLF as a line ending, do not comply. I'm hoping this is satire. Why intentionally introduce potential bugs for the sake of making a point?
- javajosh 2y ago>What would be the benefit... It is interesting that you ignore the benefits the OP describes and instead present a vague and fearful characterization of the costs. Your reaction lies at the heart of cargo-culting, the maintenance of previous decisions out of sheer dread. One can do a cost-benefit analysis and decide what to do, or you can let your emotions decide. I suggest that the world is better off with the former approach. To wit, the OP notes for benefits " The extra CR serves no useful purpose. It is just a needless complication, a vexation to programmers, and a waste of bandwidth." and a mitigation of the costs "You need to search really, really hard to find a device or application that actually interprets U+000a as a true linefeed." You ignore both the benefits assertion and cost mitigating assertion entirely, which is strong evidence for your emotionality.
- perching_aix 2y ago> you ignore the benefits the OP describes Funnily enough, the author doesn't actually describe any tangible benefits. It's all just (in my reading, semi-sarcastic) platonics: - peace - simplicity - the flourishing of humanity ... so instead of "vague and fearful", the author comes on with a "vague and cheerful". Yay? The whole shtick about saving bandwidth, lessening complications, and reducing programmer vexations are only ever implied by the author, and were explicitly considered by the person you were replying to: > You save a handful of bits at the expense of a lot of potential bugs. ... they just happened to be not super convinced. Is this the kind of HackerNews comment I'm supposed to feel impressed by? That demonstrates this forum being so much better than others?
- YZF 2y agoWhat's your estimate for the cost of changing legacy protocols that use CRLF vs. the work that will be done to support those? My intuition (not emotion) agrees with the parent that investing in changing legacy code that works, and doesn't see a lot of churn, is likely a lot more expensive than leaving it be and focusing on new protocols that over time end up replacing the old protocols anyways. OP does not really talk about the benefit, he just opines. How many programmers are vexed when implementing "HTTP, SMTP, CSV, FTP"? I'd argue not many programmers work on implementations of these protocols today. How much traffic is wasted by a few extra characters in these protocols? I'd argue almost nothing. Most of the bits are (binary, compressed) payload anyways. There is no analysis by OP of the cost of not complying with the standard which potentially results in breakage and the difficulty of being able to accurately estimate the breakage/blast radius of that lack of compliance. That just makes software less reliable and less predictable.
- deltaknight 2y agoAs an implementation detail, I assume many programs simply ignore the CR character already? Whilst of course many windows programs (and protocols as mentioned) still require CRLF, surely the most efficient way to make something cross-platform if to simply act on the LF part of CRLF, that way it works for both CRLF and LF line ends. The fact that both CRLF and LF used the same control character in my eyes in a huge bonus for this type of action to actually work. Simply make everything cross platform and start ignoring CR completely. I’m surprised this isn’t mentioned explicitly as a course of action in the article, instead it focuses on making people change their understanding of LF in to NL which is as unnecessary complication that will cause inevitable bikeshedding around this idea.
- phkahler 2y ago>> instead it focuses on making people change their understanding of LF in to NL which is as unnecessary complication that will cause inevitable bikeshedding around this idea. Not really. In order to ignore CR you need to treat LF as NL.
- deltaknight 2y agoFair point, although I’d suggest that many programs already treat LF as NL (e.g. unix text files), so this understanding of the meaning of LF already exists in the world. If you’re writing anything generic/cross-platform, you have to be able to treat LF as NL. So there isn’t really a change to be made here.
- djha-skin 2y ago> Let's make CRLF one less thing that your grandchildren need to know about or worry about. The struggle is real, the problem is real. Parents, teach your kids to use .gitattribute files[1]. While you're at it, teach them to hate byte order marks[2]. 1: https://stackoverflow.com/questions/73086622/is-a-gitattributes-file-really-necessary-for-git https://stackoverflow.com/questions/73086622/is-a-gitattribu... 2: https://blog.djhaskin.com/blog/byte-order-marks-must-diemd/ https://blog.djhaskin.com/blog/byte-order-marks-must-diemd/
- nsnshsuejeb 2y agoThe letters after the dot in my filename don't map 1 to 1 with the file format.
- Kwpolska 2y agoNope. Git should not mess with line endings, the remote repository not matching the code in your local clone can bite you when you least expect it. On Windows, one should disable the autocrlf misfeature (git config --global core.autocrlf false) and configure their text editor to default to LF.
- zulu-inuoe 2y ago3000% agree. I have been bitten endlessly by autocrlf. It is absolutely insane to me that anyone ever considered your having your SOURCE CONTROL tool get/set different content than what's in the repo
- layer8 2y agoThis is impractical in many situations, because tools that process build-source files (for example XML files that control the build, or generated source files) inherently generate CRLF on Windows. These are many, many, many tools, not just one’s text editor. The correct solution is to use .gitattributes.
- Kwpolska 2y agoIf everyone’s on Windows, or if the tool always generates/requires CRLF, then you should store the files with CRLF line endings in the repository. In a mixed Windows/Linux environment, I would still prefer to handle this myself rather than expecting Git to mangle line endings.
- fortran77 2y agoThe article had some major gaffes. Teletypes never had a ball. The stationary platen models had type boxes and cylinders, but never balls.
- refset 2y agoNot sure whether this changes anything about your critique, but note that the IBM 2741 terminal embedded a Selectric typewriter: > Selectric-based mechanisms were also widely used as terminals for computers, replacing both Teletypes and older typebar-based output devices. One popular example was the IBM 2741 terminal https://en.wikipedia.org/wiki/IBM_Selectric https://en.wikipedia.org/wiki/IBM_Selectric
- wrs 2y agoWell, it says right there, the 2741 replaced Teletypes. It wasn't a Teletype. (Not sure I'd call this a "major gaffe", though!)
- refset 2y agoNot a capital-T Teletype but it seems like it was widely used as a teleprinter and had similar mechanical constraints/requirements. The post does touch on this language ambiguity: > Teletypes (technically "teleprinters" - "teletype" was just the most popular brand name)
- perching_aix 2y agoWell, at least the title is honest. Straight up asking people to break standards out of sheer conviction is a new one for me personally, but it's definitely one of the attitudes of all time, so maybe it's just me being green. Can we ask for the typical *nix text editors to disobey the POSIX standard of a text file next, so that I don't need to use hex editing to get trailing newlines off the end of files?
- 201984 2y agoWhat's wrong with trailing newlines?
- perching_aix 2y agoOther than select software being pissy about it, not much. Just like how there's nothing wrong with CRLF, except for select software being pissy about that too.
- bmitc 2y agoYep. Select software being Unix command line tools.
- tryauuum 2y agoI do like concatenating files with cat, and if a file has its final line not ending in newline symbol the result is ugly. I know it's just me but my worldview is that the world would be better if all editors had "insert final newline" behavior
- perching_aix 2y agoMy problem is that what I input (and observe!) doesn't match what's persisted. Worse still, editors lie about it to me until I close the file and reopen it. And just to really turn the knife, various programs will then throw a fit that a character that I did not input and my editor lies about not being present, is present. I hope it's appreciable why I find this frustrating. I expect my editor to do what I say, not secretly(!) guess what I might have wanted, or will potentially want sometime in the future. Having to insert a newline while concatenating files is a chore, but a predictable annoyance. Having to hunt for mystery bytes, maybe less so.
- fweimer 2y agoSMTP <https://datatracker.ietf.org/doc/html/rfc2821#section-4.1.1.4 https://datatracker.ietf.org/doc/html/rfc2821#section-4.1.1....> is pretty clear that the message termination sequence is CR LF . CR LF, not LF . LF, and disagreements in this spot are known to cause problems (include undesirable message injection). But then enough alternative implementations that recognize LF . LF as well are out there, so maybe the original SMTP rules do not matter anymore.
- deleted 2y ago[deleted]
- deleted 2y ago[deleted]
- theginger 2y agoRidiculous! We need to develop 1 universal standard that covers everyone's use cases. Yeah!
- Ekaros 2y agoI think I can offer most reasonable compromise here. Decide upon on new UTF-8 code point. Have the use mandated and ignore and ban all end-points that do not use this code-point instead of CRLF or just LF alone.
- phkahler 2y agoSo break everything.
- whizzter 2y agohttps://xkcd.com/927/ https://xkcd.com/927/
- kps 2y agoYou mean U+2028 LINE SEPARATOR?
- bear8642 2y ago> Decide upon on new UTF-8 code point. Unicode have already done so - (NEL) https://www.compart.com/en/unicode/U+0085 https://www.compart.com/en/unicode/U+0085
- deleted 2y ago[deleted]
- 2y ago
- shadowgovt 2y agoDefine "abolish." We could certainly try to write no new software that uses them. But last I checked, there are terabytes and terabytes of stored data in various formats (to say nothing of living protocols already deployed) and they aren't gonna stop using CRLF any time soon.
- eviks 2y agoIs defined in 4 points at the end
- tedunangst 2y agoNo mention of what happened the last time we mixed and matched line endings? https://smtpsmuggling.com/ https://smtpsmuggling.com/
- deltaknight 2y agoDoesn’t this show that ignoring CR and only processing LFs is a good idea? If I’m understanding right (probably wrong), this vuln relied on some servers using CRLF only as endings, and others supporting both CRLF and LF. If every server updated to line-end of LF, thereby supporting both types, this vuln wouldn’t happen? Of course if there’s is a mixed bag then I guess this is still possible, if your server only supports CRLF. At least in that scenario you have some control over the issue though.
- hifromwork 2y agoYes, if every server/middleware implemented parsing in the same way this kind of vulnerability wouldn't happen. Same goes for HTTP smuggling and other smuggling attacks. Unfortunately, asking more people to ignore the currently estabilished standards makes the problem worse, not better.
- dwattttt 2y agoAs I mentioned else-thread: it doesn't matter as much which option is chosen, so long as everyone agrees. If everyone agrees that LF on its own is enough (and we stop sending CR's to make sure it's not part of whatever comes before LF), that's fine. But it's just as fine for everyone to agree that CRLF is right, and reject plain LF.
- WesolyKubeczek 2y ago> Even if an established protocol (HTTP, SMTP, CSV, FTP) technically requires CRLF as a line ending, do not comply. Send only NL. Now just go pound sand. Seriously. And you owe me 5 minutes of my life wasted on reading the whole thing. My god, I would have thought all those “simplification” ideas die off once you have 3 years of experience or more. Some people won’t learn. P. S. Guess even the most brilliant people tend to have dumb ideas sometimes.
- elcritch 2y agoConversely, I'd argue most brilliant people tend to have more dumb ideas than others, usually on oddly specific topics which most people would find inconsequential.
- the_gorilla 2y agoIt's true. Smart people tend to have a lot of novel ideas, most of which are going to be retarded. Most people just have no ideas.
- rgmerk 2y agoOf all the stupid and obsolete things in standards we use to interoperate, CRLF is one of the least consequential.
- moomin 2y agoCounterpoint: Unix deciding on a non-standard line ending was always a mistake. It has produced decades of random incompatibility for no particular benefit. CRLF isn’t a convention: it’s two different pieces of the base terminal API. You have no idea how many programs rely on CR and LF working correctly.
- matheusmoreira 2y agoYeah. It's weird how Unix picked LF given its love of terminals. CRLF is the semantically correct line ending considering terminal semantics. It's present in the terminal subsystem to this day, people just don't notice because they have OPOST output post processing enabled which automatically converts LF into CRLF.
- eqvinox 2y agoI'd argue (but have no historical context) that it's a distinction between storage format and presentation interface, and IMHO that makes a lot of sense. A terminal has other operations too, backspaces and deletes being the most basic. Which coincidentally are one hell of a mess across different terminal types between ^H / 0x08 and DEL / 0x7f as well… (And these distinctions predate UNIX — if I were confronted with an inconsistent mess I'd go for simplicity too, and a 2-byte newline is definitely not simple just by merit of being 2 bytes. I personally wouldn't have cared whether it was CR or LF, but would have cared to make it a single byte.)
- fanf2 2y agoIt is a standard line ending. ANSI X3.4-1968 says: 10 LF (Line Feed). A format effector that advances the active position to the same character position on the next line. (Also applicable to display devices.) Where appropriate, this character may have the meaning “New Line” (NL), a format effector that advances the active position to the first character position on the next line. Use of the NL convention requires agreement between sender and recipient of data. ASCII 1968 - https://www.rfc-editor.org/info/rfc20 https://www.rfc-editor.org/info/rfc20 ASCII 1977 - https://nvlpubs.nist.gov/nistpubs/Legacy/FIPS/fipspub1-2-1977.pdf https://nvlpubs.nist.gov/nistpubs/Legacy/FIPS/fipspub1-2-197...
- lynx23 2y agoCan OP please tell me how to abolsih CR while in Raw Mode? Did he forget about it, or am I just unimaginative?
- samatman 2y agoRight, you don't need to search that hard for a device which interprets 0xA as a line feed, just set your terminal to raw mode, done. But given the very first sentence: > CR and NL are both useful control characters. I'm willing to conclude that he doesn't intend A Blaste Against The Useless Appendage of Carriage Return Upon a New Line, or Line Feed As Some Style It, to apply to emulators of the old devices which make actual use of the distinction.
- lynx23 2y agoI know that we're technically emulating old devices... But that time is so long gone. I actually never worked on a hardware terminal in my entire career, which is already almost 30 years. I think it is about time to stop calling it emulation, because thats no longer what it is. Its simply the way how text mode applications do I/O. It has become so ubiquitous that ncurses is slowly going out of fashion, because you might as well just use the common ANSI escape sequences, because they're supported everywhere anyways. IOW, raw mode isn't just an emulation required to get a 50 year old peripheral device to work, its necessary for almost everything that sits between an CLI and a GUI.
- samatman 2y agoI agree with this, and generally take pains to refer to the programs as terminals, not terminal emulators. But at the same time, when a user presses the enter key and stdin provides CR, if you're in raw mode, you can get NL semantics by emitting CRLF to stdout, and that does in fact emulate the behavior of DEC terminals, which do that because teletypes did. > I actually never worked on a hardware terminal in my entire career I used to look books up at the library using a VT220. In the late 1990s they replaced that with an ASPX web browser endpoint running on PC hardware, and it was terrible. But I'm also not quite old enough to have used them for programming. You're completely correct that it's no longer emulation of hardware terminals, there are dozens of input and output sequences which no hardware terminal ever used or understood. In many ways it's now emulation of XTerm, but even that era is slowly being left behind.
- forrestthewoods 2y agoI could not possibly disagree with this more strongly or violently. In short - shutup and deal with it. Is it an extremely mild and barely inconvenient nuisance to deal with different or mixed line endings? Yes. Is this actually a hard or difficult problem? No. Stop trying to force everyone to break their backs so your life is inconsequentially easier. Deal with it and move on.
- Avamander 2y agoWhy do we _have to_ keep bringing this legacy baggage with us for the next decades though? Allowing CRFL-less operation intentionally, especially in new implementations. Abusing protocol tolerance is (just a bit) to switch current ones. Should allow relatively gradual progress towards Less Legacy:tm: with basically no cost. Not every change is "breaking your back" especially if you should be updating your systems anyways to implement other, larger and more important changes.
- forrestthewoods 2y agoBecause it’s literally fine and a non-issue. Only whiny Linux babies cry about it. It’s trivial for tools to support. Trivial. Like this is easiest, least harmful baggage in the history tech debt baggage. There will always be tech debt. Always and forever. Burn cycles on one that matters.
- Avamander 2y agoSo what's the issue with getting rid of this debt slowly? It costs basically nothing, yet makes it cleaner for those in the future. Debts matter at a larger scale and the long run.
- wongogue 2y agoThey carried the debt. Why shouldn’t everyone else? Regarding this issue…I don’t think the author is advocating for patching standards. Just consider CR as deprecated and use it only for backward compatibility. I do it similarly. I don’t convert line endings but any new project uses LF irrespective of the OS and configured as such in the editor.
- Animats 2y agoNow convince Microsoft. It's really the legacy of DOS that keeps this alive.
- nycdotnet 2y agoEven Notepad.exe supports LF only text files now.
- A4ET8a8uTh0 2y agoI feel it necessary to have an obligatory 'Would someone think of banking?' before we 'abolish'(however we eventually arrive at defining it )anything. I mean it is all cool to have this idea, but real world implications, where half the stuff dangles on a text file, appear to be not considered here. For clarity's sake, I am not saying don't do it. I am saying: how will that work? edit: spaces, tabs and one crlf
- anonymousiam 2y agoThis article seems like it was written to troll people into a flame war. There is no such character as NL, and the article does not at all address that fact that the "ENTER" key on every keyboard sends a CR and not a LF. Things work fine the way they are.
- o11c 2y agoU+0085 is sometimes called NL (it is the standard in EBCDIC), but more often NEL in the ASCII world.
- TacticalCoder 2y ago> There is no such character as NL ... More specifically the Unicode control character U+000a is, in the Unicode standard, named both LF and NL (and that comes from ASCII but in ASCII I think 0x0a was only called LF). It literally has both names in Unicode: but LINEFEED is written in uppercase while newline is written in lowercase (not kidding you). You can all see for yourself that U+000a has both names (and eol too): https://www.unicode.org/charts/PDF/U0000.pdf https://www.unicode.org/charts/PDF/U0000.pdf > and the article does not at all address that fact that the "ENTER" key on every keyboard sends a CR and not a LF. what a key on a keyboard sends doesn't matter though. What matters is what gets written to files / what is sent over the wire. ... $ cat > /tmp/anonymousiam<ENTER> <ENTER> <CTRL-C> ... $ hexdump /tmp/anonymousiam 00000000 000a When I hit ENTER at my Linux terminal above, it's LINEFEED that gets written to the file. Under Windows I take it the same still gets CRLF written to the file as in the Microsoft OSes of yore (?). > Things work fine the way they are. I agree
- anonymousiam 2y agoTry the cat example again with your tty in raw mode instead of cooked mode. (stty raw) Note that your job control characters will no longer function, so you will need to kill the cat command from a different terminal, then type: stty sane (or stty cooked) to restore your terminal to "normal" operation. You will then see the 0d hex carriage return characters in the /tmp/anonymousiam file, and no 0a hex linefeed characters present.
- sunk1st 2y ago> Nobody ever wants to be in the middle of a line, then move down to the next line and continue writing in the next column from where you left off. No real-world program ever wants to do that. Is this true?
- anamax 2y agoNo, it's not true. It was used for "graphics" on character-only terminals.
- numpad0 2y agoisn't CR without LF how CLI progress bars work?
- cowsandmilk 2y agoHe says there are good usages of CR, he only argues for getting rid of LF.
- samatman 2y agoReplacing LF behavior with NL behavior in terminal raw mode is a non-starter, decades worth of software will break. The enter/return key only sends CR anyway, software has to decide what to do with that: sometimes the answer is emit CRLF, often there are program-defined margins so the CR gets translated to direct cursor movements. I'm pretty sure drh is making a case only against the use of CRLF in protocols, not trying to redesign terminal in the process. If you're emulating a machine which understands LF then you're kinda stuck with line feed semantics, for better and for worse.
- Ekaros 2y agoNope.
- gfody 2y agowe should leave it for backwards compatibility and adopt U+0085 as the standard next line codepoint. and utf8 libraries could unofficially support every combination of 0A 0D as escape sequences.
- NelsonMinar 2y agosqlite is a work of absolute genius. But every once in awhile something comes along to remind us how weird its software background is. Fossil. The build system. The TCL test harness. And now this, a quixotic attempt to break 50+ years of text formatting and network protocols. Yes CRLF is dumb. No, replacing it is not realistic.
- bmitc 2y agoDoes anyone besides poorly designed Unix tools and Git actually get confused by any of this? I configure my editor to just use LF on whatever OS to appease Linux and configure Git to never mess with them. And in dealing with serial protocols, it's never an issue.
- midnitewarrior 2y agoIf you'd like to break every system, and nearly every protocol, start abolishing arbitrary line endings that have been used for decades. That will make things better.
- jftuga 2y agoI wrote a command line program to determine/detect the end-of-line format, tabs, bom, and nul characters https://github.com/jftuga/chars https://github.com/jftuga/chars Stand-alone binaries are provided for all major platforms.
- nunobrito 2y agoXKCD has graphically replied to this topic: https://xkcd.com/927/ https://xkcd.com/927/
- eviks 2y agoIt hasn't like it never does. In this case your mistake is that number of standards doesn't change
- Eduard 2y ago> Stop using "linefeed" as the name for the U+000a code point. stop reinventing terms. it's literally standardized with the name "LF" / "line feed" in Unicode.
- wongogue 2y agoHe is not reinventing anything. Unicode also defines 0a as LF, NL and EOL. In modern software, 0a is used as NL anyway. https://www.unicode.org/charts/PDF/U0000.pdf https://www.unicode.org/charts/PDF/U0000.pdf
- lifthrasiir 2y agoJust in case... Unicode doesn't define anything about C0 control characters. Everything you see from the code chart is from ISO/IEC 6429 and only shown there for information. Some parts of Unicode and related standards do assign a special meaning to U+000A, but often also to U+000D for the obvious reason.
- ericyd 2y agoI'm not trying to be obtuse but I am actually confused how a modern machine correctly interprets CRLF based on the description in this post. If a modern machine interprets LF as a newline, and the cursor is moved to the left of the current row before the newline is issued, wouldn't that add a newline _before_ the current line, i.e. a newline before the left most character of the current line? Obviously this isn't how it works but I don't understand why not.
- chowells 2y agoLine feed is "move the cursor down one line". It's irrelevant what is currently on the line. These are printer/terminal control instructions, not text editing instructions.
- ericyd 2y agoOk, I conflated terminal instruction with text editing instruction. I thought the post made them sound like they behave the same but it sounds like I misunderstood, thank you.
- BlueTemplar 2y agoIf you are thinking of it being more like pressing "Home" then "Enter", it would seem that "Enter" actually works more like LFCR ?
- WillAdams 2y agoFWIW, I actually find CRLF handy in a database export I work in --- it exports cells with multiple lines by using LF for the linebreaks --- I open it in a text editor, replace all LFs w/ \\ (so as to get a single line for each data record and to cause the linebreaks to happen in LaTeX), and it's ready for further processing.
- zac23or 2y ago> Even if an established protocol (HTTP, SMTP, CSV, FTP) technically requires CRLF as a line ending, do not comply. Send only NL. Insane. First i think it was a April 1st joke, but is not. Let's break everything because YES.
- pathartl 2y agoLet's all move to little-endian while we're at it. Don't accept anything else!
- Quekid5 2y agoIndeed. Very strange to hear a break-the-world suggestion from a person leading a company famous for never breaking the world. I'm kind of confused by this whole post. I do understand the desire for simplification (let's ignore the argument of whether this is one), but...
- MatthiasPortzel 2y agoThey acted on these words, updating their HTTP server to serve just \n. => https://sqlite.org/althttpd/info/8d917cb10df3ad28 https://sqlite.org/althttpd/info/8d917cb10df3ad28 Send bare \n instead of \r\n for all HTTP reply headers. While browser aren't effected, this broke compatibility with at least Zig's HTTP client. => https://github.com/ziglang/zig/issues/21674 https://github.com/ziglang/zig/issues/21674 zig fetch does not work with sqlite.org
- abhinavk 2y agoIt has been reverted.
- dankwizard 2y ago"Call to action" my god guy get a grip youre upset about some unicode
- zulu-inuoe 2y agoOf all the hills to die on. What an unbelivably silly one. CRLF sucks, suck it up. As many others have noted, there are millions of devices this idea puts in jeopardy for absolutely no reason. We should be reducing the exceptions, not creating them
- pdonis 2y agoFor extra fun, the original Mac OS used CR by itself to mean newline.
- ripe 2y agoHa, ha, ha! I love it. I believe the author is serious, and I think he's on to something. OP clearly says that most things in fact don't break if you just don't comply with the CRLF requirement in the standard and send only LF. (He calls LF "newline". OK, fine, his reasoning seems legit.) He is not advocating changing the language of the standard. To all those people complaining that this is a minor matter and the wrong hill to die on, I say this: most programmers today are blindly depending on third-party libraries that are full of these kinds of workarounds for ancient, weird vestigial crud, so they might think this is an inconsequential thing. But if you're from the school of pure, simple code like the SQLite/Fossil/TCL developers, then you're writing the whole stack from scratch, and these things become very, very important. Let me ask you instead: why do you care if somebody doesn't comply with the standard? The author's suggestion doesn't affect you in any way, since you'll just be using some third-party library and won't even know that anything is different. Oh bUT thE sTandArDs.
- abhinavk 2y ago> (He calls LF "newline". OK, fine, his reasoning seems legit.) He is not advocating changing the language of the standard. The Unicode standard does call it NL along with LF. 000A <control> = LINE FEED (LF) = new line (NL) = end of line (EOL) Source: https://www.unicode.org/charts/PDF/U0000.pdf https://www.unicode.org/charts/PDF/U0000.pdf
- truetraveller 2y agoI agree 100%. This is the cause of endless confusion, especially in crossplatform text files. Not to mention parsing programmatically.
- SQLite 2y agoAuthor here: My title was imprecise and unclear. I didn't mean that you should raise errors if CRLF is used as a line terminator in (for example) HTTP, only that a bare NL should be allowed as an acceptable line terminator. RFC2616 recommends as much (section 19.3 paragraph 3) but doesn't require it. The text of my proposal does say that CRLF should continue to be accepted, for backwards compatibility, just not required and not generated by default. I failed to make that point clear. My initial experiments suggested that this idea would work fine and that few people would even notice. Initially, it appeared that when systems only generate NL instead of CRLF, everything would just keep working seamlessly and without problems. But, alas, there are more systems in circulation that are unable to deal with bare NLs than I knew. And I didn't sell my idea very well. So there was breakage and push-back. I have revised the document accordingly and reverted the various systems that I control to generate CRLFs again. The revolution is over. Our grandchildren will have to continue dealing with CRLFs, it seems. Bummer. Thanks to everyone who participated in my experiment. I'm sorry it didn't work out.
- AndyKelley 2y agoCopying my comment from lobste.rs in case you didn't see it: I really appreciate this attitude. As programmers, we love to complain and grumble to each other about how the state of things suck, or that things are over complicated, but then too often the response is the software engineering equivalent of “I paid my student loans, so you should have to, too”. A new person joins the project, and WTFs at something, and the traumatized veterans say, “haha oh boy welcome, yeah everything sucks! You’ll get used to it soon.” I hate that attitude. We are at the very, very beginning of software protocols that could potentially last for millennia. From that perspective, you would look back at this situation and think of Richard’s blog post as super obvious, the clear voice of reason, and the reaction of everyone here as myopic. Even if our software protocols for whatever reason don’t last that long, we need to be working on reducing global system complexity. Beauty and elegance aside, there is such a thing as complexity budget which is limited by the laws of information theory, the computer science equivalent of the laws of physics. People like Richard understand this intuitively, and actively work towards reconstructing our world to regain complexity currency so that it can be spent on more productive things. I would have backed you 100%.
- fijiaarone 2y agoLine feed is exactly what you do when you are editing text. But nobody uses it. CR + LF was meant as an instruction for teletype printers, so it is outdated, and looks like he withdrew the proposal (which couldn’t have ever been serious) after some feedback. Fossil SCM, btw, was written by the creator of SQLite, so his opinion shouldn’t be discounted as some random nobody.
- fracus 2y agoI read your article and am now fully indoctrinated to your noble cause. I propose an official chant. "Death to LF!"
- webprofusion 2y agoNext you'll be telling us to use spaces instead of tab.
- M95D 2y agoBut adopting this new standard means we'll have to re-tool entire industries! /s
- srg0 2y agoI would also like to point out that English spelling is obsolete and should be abolished (/s). The text of the CRLF abolition proposal itself contains more digraphs, trigraphs, diphthongs, and silent letters than line-ending sequences. The last letter of the word "obsolete" is not necessary. "Should" can be written as only three letters in Shavian "𐑖𐑫𐑛". According to ChatGPT, the original proposal had: Number of sentences: 60 Number of diphthongs: 128 (pairs of vowels in the same syllable like "ai", "ea", etc.) Number of digraphs: 225 (pairs of letters representing a single sound, like "th", "ch", etc.) Number of trigraphs: 1 (three-letter combinations representing a single sound, like "sch") Number of silent letters: 15 (common silent letter patterns like "kn", "mb", etc.) For all intents and purposes, CRLF is just another digraph.
- ksp-atlas 2y agoI'm a big fan of English spelling reform and know Shavian and sometimes write in it, but I feel shavian is limited due to how heavily it uses letter rotation. Dyslexics already have trouble with b, d, p and q, having most letters have a rotated form would be challenging