4 ms·
Thanks, really interesting to see some numbers. When I parse this on the server I use a composable parser that doesn't need to do the string manipulation you h
by KayEss 10y ago
Thanks, really interesting to see some numbers.
When I parse this on the server I use a composable parser that doesn't need to do the string manipulation you have to in Python. I have no reason to believe that the parsing speed would be different at all with my code. I'd like to know though, so I'll probably do some tests -- that is if I ever get time :)
There is one other benefit (which is actually quite important for what we're using it for, but YMMV) and that is that it's very nearly CSV and for many simple files it is the same as CSV. This makes it very familiar looking to users of the system and means that for the most part these files can be easily previewed using a spreadsheet.
- dalke 10y agoYour "composable parser" is conceptually similar to my "Either I have to write my own JSON tokenizer" alternative, though you already have such an alternative. You are right that a well-optimized implementation will not be slower than a JSON Lines implementation on top of a stock JSON parser. The difficulty is in convincing people to build and deploy a high-performance CSJ parser. Why not convince them to use a high-performance CSV reader, conforming to whichever variant of CSV you wish to prefer, which presumably could include UTF-8? Your point about similarity with CSV is a good one, but it fails in exactly the same ways as existing CSV portability fails. That is, consider the following CSJ: "question","answer" "Was it Nicholson, Duvall, or Lloyd who said \"Here's Johnny!\"?","Nicholson" A trivial CSV parser which splits on "," will fail because it doesn't understand the quoting rules. A CSV parser that does understand quoting rules, like the one in Python, will almost certainly default to Excel's quoting rules, which is different than what JSON does: >>> import csv >>> r = csv.reader(open("tmp.csv")) question|answer >>> print("|".join(next(r))) Was it Nicholson, Duvall, or Lloyd who said \Here's Johnny!\"?"|Nicholson It can be fixed for Python by changing the convention: >>> r=csv.reader(open("tmp.csv"), escapechar='\\') >>> print("|".join(next(r))) question|answer >>> print("|".join(next(r))) Was it Nicholson, Duvall, or Lloyd who said "Here's Johnny!"?|Nicholson but someone previewing it in a spreadsheet may see what appears to be garbage characters. It's even worse for CSJ files which embed JSON objects, because the commas in those objects are not subject to quoting rules. And there's the question of how the spreadsheet handles non-ASCII characters. Here's an example using the tiny CSJ file from the spec page: "name", "age", "job" "Kirit S\u00e6lensminde", 45, "Minister Without Portfolio" "Freyja S\u00e6lensminde", 5, null Using Python's default CSV reader I get: >>> print("|".join(next(r))) name| "age"| "job" >>> print("|".join(next(r))) Kirit S\u00e6lensminde| 45| "Minister Without Portfolio" >>> print("|".join(next(r))) Freyja S\u00e6lensminde| 5| null You can immediately see that the space after the comma caused a portability problem! I'll regenerate the CSJ file so it doesn't have a space: "name","age","job" "Kirit S\u00e6lensminde",45,"Minister Without Portfolio" "Freyja S\u00e6lensminde",5,null Even then, the Unicode escapes cause a problem: >>> print("|".join(next(r))) name|age|job >>> print("|".join(next(r))) Kirit S\u00e6lensminde|45|Minister Without Portfolio >>> print("|".join(next(r))) Freyja S\u00e6lensminde|5|null Thus, almost the only way you get perfect fidelity is with the CSV subset that is already portable. That is, "for the most part", CSV is portable enough. (BTW, if imperfect fidelity is okay, you can open the JSON-Lines-for-table as a CSV file. The only problem will be the '[' as the first character in the first column, and the ']' as the last character in the last column.)