3 ms·
Line-delimited JSON works, but imposes a high size overhead (field names are repeated for every record) and is not well supported in non-programming environment
by rav 6y ago
Line-delimited JSON works, but imposes a high size overhead (field names are repeated for every record) and is not well supported in non-programming environments.
The article only discussed line-delimited JSON objects. How about line-delimited JSON arrays? You put the field names in an array on the first line, and the values in arrays on the subsequent lines.
Although in my day-to-day, the overhead of field names is rarely problematic.
- dragonwriter 6y agoYeah, line delimited JSON where each entity is a JSON array seems to hit exactly the point being aimed at. Also, like XML (definitely not space efficient) with SAX parsers, plain JSON has streaming processors available. They are naturally more complex than streaming processors for line-oriented formats, but they exist, so plain JSON, either array of arrays objects with the main data in array of array is an alternative, e.g.: { "headers": [...], "body": [ [...], [...], ... ] } the last form is also trivially extensible to support other keys for additional metadata, without harming the ability of applications without knowledge of the other keys to process the main data
- CraftThatBlock 6y agoI think the OP was talking more about: ["name","address"] ["Bob","1 Apple Street"] ["John","2 Apple Street"] It's compatible with JSON Lines, and mostly equivalent to CSVs, at least how they are most commonly used.
- jrimbault 6y agoThat was also my thought getting to this line. First line is a list of the records fields names, then every line is a tuple of a record, convert tuple to struct and done.
- mcswell 6y agoI've seen very large JSON files passed around in zipped form, and there's a Python library that reads them in this form (afaict without first creating an uncompressed version of the entire zipfile). I would think if you zipped a JSON file where the fields are repeated, the compression algorithm would pick up on this and compress those field names (including their quotes and trailing colon etc.). Or do I misunderstand how compression works?