5 ms·
Any time you have a character with a special meaning you have to handle that character turning up in the data you're encoding. It's inevitable. No matter what o
by niccl 2y ago
Any time you have a character with a special meaning you have to handle that character turning up in the data you're encoding. It's inevitable. No matter what obscure character you choose, you'll have to deal with it
- noosphr 2y agoThe difference is that the coma and newline characters are much more common in text than 0x1F and 0x1E, which if you restrict your data to alphanumeric characters (which you really should) will never appear anywhere else.
- dietr1ch 2y agoThe characters would likely be unique, maybe even by the spec. Even if you wanted them, we use backslashes to escape strings in most common programming languages just fine, the problem CSV is that commas aren't easy to recognize because they might be within a single or double quote string, or might just be a separator. Can strings in CSV have newlines? I bet parsers disagree since there's no spec really.
- wvenable 2y agoExcept we have all these low ASCII characters specifically for this purpose that don't turn up in the data at all. But there is, of course, also an escape character specifically for escaping them if necessary.
- Brian_K_White 2y agoYou can't type any of those on a typewriter, or see them in old or simple simple editors, or no editor like just catting to a tty. If you say those are contrived examples that don't matter any more then you have missed the point and will probably never acknowledge the point and there is no purpose in continuing to try to communicate. One can only ever remember and type out just so many examples, and one can always contrive some response to any single or finite number of examples, but they are actually infinite, open-ended. Having a least common denominator that is extremely low that works in all the infinite situations you never even thought of, vs just pretty low and pretty easy to meet in most common situations, is all the difference in the world.
- kapep 2y agoEven if you find a character that really is never in the data - your encoded data will contain it. And it's inevitable that someone encodes the encoded data again. Like putting CSV in a CSV value.
- oever 2y agoIt's evitable by stating the number of bytes in a field and then the field. No escaping needed and faster parsing.
- pasc1878 2y agoBut not human editable/readable
- kevincox 2y agoI understand this argument in general. But basically everyone has some sort of spreatsheet application that can read CSV installed. In some alternate worked where this "binary" format caught on it would be a very minor issue that it isn't human readable because everyone has a tool that is better at reading it than humans are. (See the above mentioned non-local property of quotes where you may think you are reading rows but are actually inside a single cell.) Makes me also wonder if something like CBOR caught on early enough we would just be used to using something like `jq` to read it.
- jonathanberi 2y agohttps://github.com/wader/fq https://github.com/wader/fq is "jq for binary formats."
- mjw_byrne 2y agoExactly. "Use a delimiter that's not in the data" is not real serialisation, it's fingers-crossed-hope-for-the-best stuff. I have in the past does data extractions from systems which really can't serialise properly, where the only option is to concat all the fields with some "unlikely" string like @#~!$ as a separator, then pick it apart later. Ugh.
- dietr1ch 2y ago> Exactly. "Use a delimiter that's not in the data" is not real serialisation, it's fingers-crossed-hope-for-the-best stuff. It's not doing just this, you pick something that's likely not in the data, and then escape things properly. When writing strings you can write a double quote within double quotes with \", and if you mean to type the designated escape character you just write it twice, \\. The only reason you go for something likely not in the data is to keep things short and readable, but it's not impossible to deal with.
- mjw_byrne 2y agoI agree that it's best to pick "unlikely" delimiters so that you don't have to pepper your data with escape chars. But some people (plenty in this thread) really do think "pick a delimiter that won't be in the data" - and then forget quoting and/or escaping - is a viable solution.
- solidsnack9000 2y agoIn TSV as commonly implemented (for example, the default output format of Postgres and MySQL), tab and newline are escaped, not quoted. This makes processing the data much easier. For example, you can skip to a certain record or field just by skipping literal newlines or tabs.