6 ms·
Tab always made more sense as a column delimiter, but CSV is much better standard. However through simplicity, TSV outshines when used for high volume data pro
by guidedlight 2y ago
Tab always made more sense as a column delimiter, but CSV is much better standard.
However through simplicity, TSV outshines when used for high volume data processing.
- SahAssar 2y agoHow is TSV more simple when it just uses another delimiter? Seems like the exact same processing is needed.
- willf 2y agoBecause commas are much more common than tabs and so there is less escaping needed
- SahAssar 2y agoBut you always need to have the code for escaping and you always need to check for it. I don't see how an implementation can be simpler without dropping support for tabs within columns, which would make it non-conformant to the spec.
- RadiozRadioz 2y agoBecause, while you always _should_ implement the proper escaping, that takes work. Not a large amount of work, but more than zero. In many cases your data doesn't contain commas or tabs, so you can do it the super simple way and get back that time. There are more cases where data is tabless than commaless, so using TSV affords you more opportunities to get this quick and dirty timesave when you need a fast solution.
- SahAssar 2y agoAh, so what you actually mean is more performant (for a subset of uses), not simpler? So if I have a TSV and a CSV containing either pure numbers or complex data (say the contents of each file in a codebase where each row likely contains both commas and tabs), they would be equivalent in both performance, right? If I have a TSV and a CSV containing natural written language TSV might be more performant since there are likely much more commas than tabs (I'm guessing this is your point?). Regardless of the input data the encoding/decoding code would be equally simple (since they need to account for the same edge cases), right?
- galleywest200 2y agoIt is simply easier to take basic written human text and put it into a TSV than a CSV as humans use commas far more than tabs. You can replace a TAB with spaces and keep things legible but replacing a comma in a sentence can literally change the meaning. Maybe you could wrap everything in quotation marks but that is ugly.
- SahAssar 2y agoIt's not easier or simpler since you need the exact same checks and steps, just with a different delimiter. It doesn't matter if you need to do them less times for certain inputs since you need the same checks, the same encoding/decoding steps and so on. Do you think you would have an easier time writing a TSV parser than a CSV one? If so, why? And wrapping in quotes does not solve anything since now you need to both check for escaped quotes and tabs/commas. It's the same but one level deeper.
- dafelst 2y agoI think what they're saying is that with some minor control over the data in your dataset, you don't need to care about escaping _in your parser_ at all. The same might be said of CSV but I would argue that in the majority of situations tabs are less semantically meaningful than commas and newlines, so it is generally fine just to strip them out. Obviously this is not robust solution, but in cases I've seen, it works adequately. If one were to be doing it "the right way" then I agree with you wholeheartedly.
- SahAssar 2y agoI get what you are saying, but my point is that is not CSV or TSV. It's a homemade format with its own rules that just happens to be inspired by TSV or CSV.
- RadiozRadioz 2y agoSorry, I probably chose my words incorrectly. I meant "work" and "time" in relation to human work to produce the parser. With CSV, it's more likely you'll encounter data where you need to implement the escaping. With TSV, you can get away with the simple parser for much longer, as it's comparatively rare to find data that contains tabs.
- kjkjadksj 2y agoOnly if you anticipate your data having commas in fields, which is somewhat rare in my experience (as in I’ve never seen it at all).
- setopt 2y agoComma is the decimal separator in many European languages. It’s also not uncommon in strings.
- kjkjadksj 2y agoPandas can handle that with a flag
- setopt 2y agoSure. Another option is semicolon-delimited files which are also in use in Europe, and Pandas handles that fine too. I was responding to your comment that you had never seen commas in data fields in CSV files, and wanted to point out that this is a quite common issue in Europe. (It also often wreaks havoc with Excel files btw, as Excel will then only casts strings to decimal numbers when a file is opened in some locales...)
- mike256 2y agoBecause its strictly using tabs and only tabs. Not comma or semicolon or whatever else...
- constantcrying 2y agoTo be honest, both are used for exactly one reason, because a parser which vaguely works is extremely simple to write. In basically every case a superior data format exists.
- setopt 2y agoIt is also the original purpose of TAB: the tabulation key, used to delimit tabular data. I also like that TSV files are both readable in plain text and easy to manipulate directly in editors like Vim and Emacs via rectangle operations or macros. That’s often a quick way to e.g. delete or transpose columns, or even do simple math operations on columns, without needing to go via Pandas, Excel, etc.
- mike256 2y agoAs long as none of your csv producers are using microsoft products or live outside the US. In europe csv is terrible with different delimiter depending on system language and such nasty things just because of excel.