3 ms·
Just to be clear, are you talking about a file where each line is its own json object rather than the entire file being one large array? If so, it’s easy in Py
by aobdev 4y ago
Just to be clear, are you talking about a file where each line is its own json object rather than the entire file being one large array?
If so, it’s easy in Python to do the following:
with open(“file.json”) as src:
for line in src:
json.loads(line)
- eatonphil 4y agoI'm talking about with simdjson. Lemire suggested reading line-by-line is not a good idea [0]. So I'm asking about the ideal approach using simdjson, not JSON parsers in general. [0] https://github.com/simdjson/simdjson/issues/188#issuecomment-503292225 https://github.com/simdjson/simdjson/issues/188#issuecomment...
- MikeDelta 4y agoWe had to parse thousands of multi GB zipped jsons for financial data. I don't have any code but it involved (in C++) boost gzip to unpack chunks into 100MB blocks and the simdjson iterate_many function to parse the block into a stream. The json files were not line-delimited but every row had a newline (json: [\n{...},\n{...}\n....\n] ), so we had to clean the commas in order for simdjson to process it as if it were newline-delimited. Whatever remained at the end of the buffer (few hundred bytes of incomplete json) we would copy just before the boost-unzip block so that there was a continuation of json data. We also reused the parser and string buffer for performance. Also try to parse the json fields in the right order for fastest performance, if possible.
- capableweb 4y ago> Also try to parse the json fields in the right order for fastest performance, if possible. What do you mean with this precisely, not sure I understand? What is the right order? AFAIK, JSON attributes/keys are not ordered, so there is no "right" order, or order at all.
- MikeDelta 4y agoAgreed on jsons not being ordered. The docs on simdjson mention that if you parse the fields in the order as they appear on file, then simdjson can do it in one iteration. If you get fields in random order, simdjson will have to loop the data a few times. If your json on disk is { "field1": "val1", "field2": "val2" } then it is faster to say (pseudo C++) simdjson::document_reference elem; std::string_view tmp; elem["field1"].get_string().get(tmp); elem["field2"]... than the other way around. This works only if you know what you are looking for and it is consistent on disk. Note that you are reading from an iterator that consumes the value: getting a field twice gives an error; totally different than reading from a dict. See also [1] under "Extracting Values: You can cast a JSON element to a native type..." [1] https://github.com/simdjson/simdjson/blob/master/doc/basics.md#using-the-parsed-json https://github.com/simdjson/simdjson/blob/master/doc/basics....
- capableweb 4y agoAh yes, that makes sense. Thanks a lot for explaining!