3 ms·
I had a very similar problem with Ruby and some of our workloads. I ended up utilizing this evented parser [1] to process smaller subsets of the JSON document a
by aabaker99 11y ago
I had a very similar problem with Ruby and some of our workloads. I ended up utilizing this evented parser [1] to process smaller subsets of the JSON document and then clear them from memory. The result is a memory usage that is at most the size of one of these subsets (plus some overhead).
I have been wanting to generalize this approach with some open source software for some time, but I haven't gotten around to it. I think an attractive solution is for a user to specify a JSONPath [2] upfront about the types of subsets they are interested in. In the case of Ruby these chunks could then be yielded by an iterator and cleared, but I wonder if a more general-purpose CLI program which prints the subsets to stdout would also work if the subsets are small enough. But this adds another serialization/deserialization step.
[1] https://github.com/dgraham/json-stream https://github.com/dgraham/json-stream
[2] http://goessner.net/articles/JsonPath/ http://goessner.net/articles/JsonPath/
- lobster_johnson 11y agoI did something very similar with XML in Ruby. All the libraries are either event-based (SAX, which is awful for actual parsing code) or snarf the whole document into RAM. I wrote a simple SAX wrapper [1] that applies a (so far, extremely simplified) subset of XPath to the stream, then parses each match into a DOM using Nokogiri. Works exceedingly well, though it's not faster than reading everything into RAM. [1] https://github.com/t11e/nokogiri-simple-streaming https://github.com/t11e/nokogiri-simple-streaming