4 ms·
> rather than make a format which is harder to read Do you really believe that this is harder to read than JSON lines? We use both, and I've never heard a user
by KayEss 10y ago
> rather than make a format which is harder to read
Do you really believe that this is harder to read than JSON lines? We use both, and I've never heard a user complain that CSJ is hard to read, but the JSON lines has a lot of clutter.
{"age": 45, "name": "Kirit Sælensminde", "job": "Minister Without Portfolio"}
{"job": null, "age": 5, "name": "Freyja Sælensminde"}
That seems fine for a low number of columns, but gets much harder as the column count goes up. Especially so as you lose control of the key order in the line.
- dalke 10y agoThe difference between this format and JSON lines is 1) the leading and trailing '['/']' for each line, 2) the convention that each line encodes an array and all arrays have the same length, and 3) the convention that the first line contains string names for the keys for the values on the remaining lines. To clarify, JSON Lines says "Each Line is a Valid JSON Value", "The most common values will be objects or arrays, but any JSON value is permitted." It does not need to encode an object, so there needn't be "a lot of clutter". Here's a sketch of a table parser in Python built on top of JSON Lines using conventions 2) and 3): import json def read_table(infile): line = infile.readline() keys = json.loads(line) for line in infile: yield dict(zip(keys, json.loads(line))) The equivalent for a CSJ parser replaces the two "(line)" occurrences with "('[' + line + ']')". As such, no, it's not hard to write a CSJ parser, but the performance will be slower. I tested it with 3 columns (string, int, int) and 1456021 lines (including the header). Best of 4 for the JSON Lines + convention = 11.3 seconds. Best of 4 for CSJ = 11.8 seconds, which was slower than any of the JSON Lines parser times. Perhaps the 8% slowdown isn't important After all, the uncompressed file size is smaller. #bytes gzip size JSON Lines 34316877 6561939 CSJ 31404835 6576623 savings: ~10% ~0.2% So if the uncompressed savings is important, then use CSJ. But that's a small niche. Anyone concerned about space will use a custom format, while those who need to embed JSON objects or lists aren't going to be worried about 2 extra characters per line.
- KayEss 10y agoThanks, really interesting to see some numbers. When I parse this on the server I use a composable parser that doesn't need to do the string manipulation you have to in Python. I have no reason to believe that the parsing speed would be different at all with my code. I'd like to know though, so I'll probably do some tests -- that is if I ever get time :) There is one other benefit (which is actually quite important for what we're using it for, but YMMV) and that is that it's very nearly CSV and for many simple files it is the same as CSV. This makes it very familiar looking to users of the system and means that for the most part these files can be easily previewed using a spreadsheet.
- dalke 10y agoYour "composable parser" is conceptually similar to my "Either I have to write my own JSON tokenizer" alternative, though you already have such an alternative. You are right that a well-optimized implementation will not be slower than a JSON Lines implementation on top of a stock JSON parser. The difficulty is in convincing people to build and deploy a high-performance CSJ parser. Why not convince them to use a high-performance CSV reader, conforming to whichever variant of CSV you wish to prefer, which presumably could include UTF-8? Your point about similarity with CSV is a good one, but it fails in exactly the same ways as existing CSV portability fails. That is, consider the following CSJ: "question","answer" "Was it Nicholson, Duvall, or Lloyd who said \"Here's Johnny!\"?","Nicholson" A trivial CSV parser which splits on "," will fail because it doesn't understand the quoting rules. A CSV parser that does understand quoting rules, like the one in Python, will almost certainly default to Excel's quoting rules, which is different than what JSON does: >>> import csv >>> r = csv.reader(open("tmp.csv")) question|answer >>> print("|".join(next(r))) Was it Nicholson, Duvall, or Lloyd who said \Here's Johnny!\"?"|Nicholson It can be fixed for Python by changing the convention: >>> r=csv.reader(open("tmp.csv"), escapechar='\\') >>> print("|".join(next(r))) question|answer >>> print("|".join(next(r))) Was it Nicholson, Duvall, or Lloyd who said "Here's Johnny!"?|Nicholson but someone previewing it in a spreadsheet may see what appears to be garbage characters. It's even worse for CSJ files which embed JSON objects, because the commas in those objects are not subject to quoting rules. And there's the question of how the spreadsheet handles non-ASCII characters. Here's an example using the tiny CSJ file from the spec page: "name", "age", "job" "Kirit S\u00e6lensminde", 45, "Minister Without Portfolio" "Freyja S\u00e6lensminde", 5, null Using Python's default CSV reader I get: >>> print("|".join(next(r))) name| "age"| "job" >>> print("|".join(next(r))) Kirit S\u00e6lensminde| 45| "Minister Without Portfolio" >>> print("|".join(next(r))) Freyja S\u00e6lensminde| 5| null You can immediately see that the space after the comma caused a portability problem! I'll regenerate the CSJ file so it doesn't have a space: "name","age","job" "Kirit S\u00e6lensminde",45,"Minister Without Portfolio" "Freyja S\u00e6lensminde",5,null Even then, the Unicode escapes cause a problem: >>> print("|".join(next(r))) name|age|job >>> print("|".join(next(r))) Kirit S\u00e6lensminde|45|Minister Without Portfolio >>> print("|".join(next(r))) Freyja S\u00e6lensminde|5|null Thus, almost the only way you get perfect fidelity is with the CSV subset that is already portable. That is, "for the most part", CSV is portable enough. (BTW, if imperfect fidelity is okay, you can open the JSON-Lines-for-table as a CSV file. The only problem will be the '[' as the first character in the first column, and the ']' as the last character in the last column.)