3 ms·
This comes up every so often. I think there are good reasons it hasn't caught on. Here's a relevant comment from ten years ago: > whole point is to be roughly
by bariumbitmap 1mo ago
This comes up every so often. I think there are good reasons it hasn't caught on. Here's a relevant comment from ten years ago:
> whole point is to be roughly human-readable and using non-printing characters defeats that. You can't even easily enter these things via the command line.
> If we're abandoning human-readability, why even bother with ASCII? Just use a binary format. Has anyone actually used ASCII unit and record separator delimiters successfully? I'd be curious about what advantages they had over a binary format, even just a protobuf or Thrift serialized form. If we want to preserve schemalessness, there's stuff like Sereal.
--- arjie, June 8, 2016
https://news.ycombinator.com/item?id=11862769 https://news.ycombinator.com/item?id=11862769
Here's another one from more than twelve years ago:
> I've done this.
> Everybody hated it. Most text editors don't display anything useful with these characters (either hiding them altogether or showing a useless "uknown" placeholder), and spreadhseet tools don't support the record separator (although they all let you provide a custom entry separator so the "unit" separator can work). Besides the obvious problem that there's no easy way to type the darned things when somebody hand-edits the file.
--- Pxtl, March 26, 2014
https://news.ycombinator.com/item?id=7474600 https://news.ycombinator.com/item?id=7474600
- toast0 1mo agoYahoo access logs used control characters to separate fields beyond a fixed width field, but it wasn't the record separators. It used ctrl-E to separate fields, which had a one character identifier. Some of the fields had sub-fields which were separated ctrl-F. Slide 22 https://www.radwin.org/michael/talks/yapache-oscon2006.pdf https://www.radwin.org/michael/talks/yapache-oscon2006.pdf
- kirb 1mo agoOn displaying control characters, as of now, at least VSCode and Sublime Text do show them clearly. VSCode uses Unicode control pictures “␄” with a red background, while Sublime shows plain-text “<EOT>” in grey. You can also copy-paste - the real control character lands on your clipboard. Otherwise the rest of it stands, especially for a non-developer using Notepad or TextEdit. This is trading off ongoing usability for one-time developer convenience. The example given in the post would also struggle with a large file as it loads the entire contents into memory, while having Python feed you lines allows it to read in chunks. Plenty of accurate, unit-tested CSV parsers exist in every language, it’s fine to use one and be done with it. Tabular data formats are a solved problem.
- xg15 1mo agoI think there is a deeper question: "record separator" and "unit separator" chars are a good idea, but why are they invisible? I agree that in the current form, they are mostly unusable. What was the idea of the designers how they should be used?
- RiverCrochet 1mo agoASCII was created in the 1960s. That's around the time when the industry wasn't really sure if all files were just going to be a stream of bytes that were completely up to programs to interpret, or if it was the OS's responsibility to enforce a database-like structure on all disk-like I/O.