3 ms·
> Notably if I see anyone parsing CSVs with cut again I'm going to die inside. Try unpicking a problem where someone put in the name field "Smith, Bob"... How
by Bluecobra 2y ago
> Notably if I see anyone parsing CSVs with cut again I'm going to die inside. Try unpicking a problem where someone put in the name field "Smith, Bob"...
How do you tackle this? Would you count the numbers of commas in each line then manually fix the lines that contain more fields?
- hggigg 2y agoYou parse them and reject CSVs that do not conform. There is absolutely no way to reason about a malformed CSV.
- harry8 2y agohttp://www.catb.org/~esr/writings/taoup/html/ch05s02.html http://www.catb.org/~esr/writings/taoup/html/ch05s02.html paragraph titled DSV style. (Yeah esr, not a fan, whatever...) Csv sucks no matter what, there is no one csv spec. Then even if you assume the file is "MS Excel style csv" you can't validate it conforms. There's a bunch of things the libraries do that cope with at least some of it that you will not replicate with cut or an awk one liner.
- ReleaseCandidat 2y agoYes, as in "check the number of parsed fields for each line" and don't forget about empty fields. Throw an error and stop the program if the number of columns isn't consistent. Which doesn't mean that you can't parse the whole file and output all errors at once (which is the preferred way, we don't live in the 90s any more ;), just don't process the wrong result. And with usable error messages, not just "invalid line N".