4 ms·
I think it's too bad that we didn't choose delimters that had no chance of appearing in actual text. Commas are a big one, when parsing csv there's always the p
by version_five 3y ago
I think it's too bad that we didn't choose delimters that had no chance of appearing in actual text. Commas are a big one, when parsing csv there's always the problem of having commas in the text of a field. One hack I have used if I don't want to delete them is to swap them for a character I imagine will never be in the text, such as | (pipe). It all could have been avoided if we had some standard delimiters that were not part of common text.
- meep0l 3y agoThe tab character (U+0009) is a good candidate for this. Many CSV parsers already support it.
- mtizim 3y agohttps://en.m.wikipedia.org/wiki/C0_and_C1_control_codes#Field_separators https://en.m.wikipedia.org/wiki/C0_and_C1_control_codes#Fiel... The problem is that any delimiter that has no chance of appearing in actual text will be hard to discover, and cannot appear on a standard keyboard (so it is not easily human-writable). So we are kind of stuck with the comma for human readable formats.
- Symbiote 3y agoThe standard characters are ASCII/Unicode field separator, group separator, record separator and unit separator. https://en.wikipedia.org/wiki/C0_and_C1_control_codes#Field_separators https://en.wikipedia.org/wiki/C0_and_C1_control_codes#Field_...
- jeroenhd 3y agoI have actually used two of these in some kind of hell SQL query that needed to concatenate strings for the fastest path between database and frontend as possible. I was surprised to see how well it actually worked, I assumed surely some kind of step in the middle of the chain would break non-printable characters. Not using these is wasting so many bits of one-byte character encodings, I don't understand why we even need to bother escaping CSV files if we could just use the appropriate control characters instead.
- tannhaeuser 3y agoThe more immediate problem with Tim Berners-Lee's choice of delimiters in URL syntax is that ampersand starts an entity reference in SGML default concrete syntax and thus <a href="bla&x=y"> will be rejected as a reference to an undeclared entity "x" in SGML (whereas in XML it will be rejected as incomplete entity reference missing a terminating ";" character).
- im3w1l 3y agoThis doesn't work. There will always be that one guy thinking "what if we put a csv in another csv?!" So escaping it is.
- version_five 3y agoI've done a lot of csv parsing with shell scripts (usually using awk or cut). My rule of thumb is that if i know the csv may have some commas in quoted text, I can just remove them and work with the shell script. If it's going to have newlines or anything more complicated, I need a real csv parser.