17 ms·
Analyzing multi-gigabyte JSON files locally
- isoprophlex 4y agoNice writeup, but is jq & GNU parallel or a notebook full of python spaghetti the best (least complex) tool for the job? DuckDB might be nice here, too. See https://duckdb.org/2023/03/03/json.html https://duckdb.org/2023/03/03/json.html
- nojito 4y agoCalling Dask python spaghetti is quite hilarious. That spaghetti can auto scale to hundreds of machines without skipping a beat. Which is far more useful than the other tools you mentioned which are only useful for one off tasks.
- isoprophlex 4y agoParallelizing a turd across hundreds of machines doesn't mean you're doing something genius, it just means you now have a hundred machines that have to deal with your shit.
- samwillis 4y agoDuckDB is awesome. As a comparison, I have a dataset that starts life as a 35gb set of json files. Imported into Postgres it's ~6gb, and a key query I run takes 3 min 33 seconds. Imported into DuckDB (still about ~6gb for all columns), the same SQL query takes 1.1 second! The key thing is that the columns (for all rows) the query scans total only about 100mb, so DuckDB has a lot less to scan. But on top of that it's vectorised query execution is incredibly quick. https://mobile.twitter.com/samwillis/status/1633213350002798593 https://mobile.twitter.com/samwillis/status/1633213350002798...
- pletnes 4y agoI found that exporting big tables as a bunch of parquet files is faster and uses less memory than duckdb’s internal format.
- samus 4y agoPostgreSQL would probably be way faster if you add proper indexes.
- mattpallissard 4y agoAnd not store the data as Json/jsonb to begin with.
- pletnes 4y agoDuckdb is fantastic. Doesn’t need a schema, either.
- e12e 4y agoThere's also clickhouse/clickhouse local - eg: https://clickhouse.com/blog/worlds-fastest-json-querying-tool-clickhouse-local https://clickhouse.com/blog/worlds-fastest-json-querying-too... https://clickhouse.com/docs/en/operations/utilities/clickhouse-local https://clickhouse.com/docs/en/operations/utilities/clickhou... https://clickhouse.com/blog/getting-data-into-clickhouse-part-2-json https://clickhouse.com/blog/getting-data-into-clickhouse-par...
- cpuguy83 4y agoJq does support slurp mode so you should be able to do this using that... granted I've never attempted this and the syntax is very different. --- edit --- I used the wrong term, the correct term is streaming mode.
- Groxx 4y agoIt does work, but it is a huge headache to use, in part because the documentation around it is nowhere near enough to understand how to use it. If I used it regularly I'd probably develop a feel for it and be much faster - it is reasonable, just abnormal and extremely low level, and much harder to use with other jq stuff. But I almost always start looking for alternatives well before I reach that point.
- hprotagonist 4y agoi would seriously consider sqlite-utils here. https://sqlite-utils.datasette.io/en/stable/cli.html https://sqlite-utils.datasette.io/en/stable/cli.html
- qbasic_forever 4y agoWas going to post the same thing, I suspect converting the dataset to a SQLite db would be infinitely more fast and productive than pecking away at it with pandas and such.
- philwelch 4y agoSQLite is great for datasets that fit comfortably into memory, but otherwise it starts to struggle.
- hprotagonist 4y agohappily, i have multiple gigabytes of memory …
- philwelch 4y agoSure, but a 40 GB SQLite database on a machine with 16 GB of RAM is not gonna be happy
- qbasic_forever 4y agoYou're not going to do better with pandas or similar tools. If it can't fit in memory, it's going to be painful. SQLite is the least painful in my experience, and it sets you up for working with the data in a proper DB like postgres or similar for when you get fed up with the memory constraints.
- philwelch 4y agoI wouldn’t use pandas in that situation either.
- thakoppno 4y agoWould sampling the JSON down to 20MB and running jq experimentally until one has found an adequate solution be a decent alternative approach? It depends on the dataset one supposes.
- epalm 4y agoYeah, I do this when querying sql databases. I limit the data to some small/local range, iteratively work on the query, and when I'm happy with the local results, I remove the filter and get the big results.
- version_five 4y agoFor a hacky solution, I've often just used grep, tr, awk, etc. If it's a well structured file and all the records are the same or similar enough, it's often possible to grep your way into getting the thing you want on each line, and then use awk or sed to parse out the data. Obviously lots of ways this can break down, buy 9GB is nothing if you can make it work with these tools. I have found jq much slower.
- philwelch 4y agoYeah, if the JSON is relatively flat, converting to TSV makes the data fairly trivial to consume using awk and other classic command line tools. I did a lot of this when a past employer decided they couldn’t afford Splunk.
- hamilyon2 4y agoClickhouse is the best way to analyze 10GB sized json by far. Latest bunch of features add near-native json support. Coupled with ability to add extracted columns make the whole process easy. It is fast, you can use familiar SQL syntax, not constrainted to RAM limits. It is a bit hard if you want to iteratively process file line-by line or use advanced SQL. And you have one-time cost of writing schema. Apart from that, I can't think of any downsides. Edit: clarify a bit
- fdajojiocsjo 4y ago[dead]
- mastax 4y agoDask looks really cool, I hope I remember it exists next time I need it. I've been pretty baffled, and disappointed, by how bad Python is at parallel processing. Yeah, yeah, I know: The GIL. But so much time and effort has been spent engineering around every other flaw in Python and yet this part is still so bad. I've tried every "easy to use" parallelism library that gets recommended and none of them has satisfied. Always: "couldn't pickle this function" or spawning loads of processes that use up all my RAM for no visible reason but don't use any CPU or make any indication of progress. I'm sure I'm missing something, I'm not a Python guy. But every other language I've used has an easy to use stateless parallel map that hasn't given me any trouble.
- isoprophlex 4y agoI've been seeing python at least once every week for a looooong time. Years. A decade maybe. You are not missing anything. It's a big steamy pile of horse manure.
- xk3 4y agoThreadPoolExecutor if IO-bound ProcessPoolExecutor if CPU-bound for example with ThreadPoolExecutor(max_workers=4) as e: e.submit(shutil.copy, 'src1.txt', 'dest1.txt') e.submit(shutil.copy, 'src2.txt', 'dest2.txt') but yeah if you're truly CPU bound then move to something lower level like C or Rust
- dermesser 4y agoI can recommend Julia for easier parallelization while being reasonably Python-like. It's compiled, too, which helps even with single-threaded throughput.
- zeitlupe 4y agoSpark is my favorite tool to deal with jsons. It can read as many jsons – in any format located in any even nested folder structure – as you want, offers parallelization, and is great to flatten structs. I've never run into memory issues (or never ran out of workarounds) so far.
- pidge 4y agoYeah, given that everything is now multi-core, it makes sense to use a natively parallel tool for anything compute-bound. And Spark will happily run locally and (unlike previous big data paradigms) doesn’t require excessive mental contortions. Of course while you’re at it, you should probably just convert all your JSON into Parquet to speed up successive queries…
- iknownothow 4y agoHow much memory would a spark worker need to process a single JSON file that is 25GB? To clarify, this is not JSONL or NDJSON file. Just a single JSON object.
- jmmv 4y agoSome random comments: * A few GBs of data isn't really that much. Even /considering/ the use of cloud services just for this sounds crazy to me... but I'm sure there are people out there that believe it's the only way to do this (not the author, fortunately). * "You might find out that the data doesn’t fit into RAM (which it well might, JSON is a human-readable format after all)" -- if I'm reading this right, the author is saying that the parsed data takes _more_ space than the JSON version? JSON is a text format and interning it into proper data structures is likely going to take _less_ space, not more. * "When you’re ~trial-and-error~iteratively building jq commands as I do, you’ll quickly grow tired of having to wait about a minute for your command to succeed" -- well, change your workflow then. When tackling new queries, it's usually a good idea to reduce the data set. Operate on a few records until you have the right query so that you can iterate as fast as possible. Only once you are confident with the query, run it on the full data. * Importing the data into a SQLite database may be better overall for exploration. Again, JSON is slow to operate on because it's text. Pay the cost of parsing only once. * Or write a custom little program that streams data from the JSON file without buffering it all in memory. JSON parsing libraries are plentiful so this should not take a lot of code in your favorite language.
- bastawhiz 4y ago> JSON is a text format and interning it into proper data structures is likely going to take _less_ space, not more. If you're parsing to structs, yes. Otherwise, no. Each object key is going to be a short string, which is going to have some amount of overhead. You're probably storing the objects as hash tables, which will necessarily be larger than the two bytes needed to represent them as text (and probably far more than you expect, so they have enough free space for there to be sufficiently few hash collisions). JSON numbers are also 64-bit floats, which will almost universally take up more bytes per number than their serialized format for most JSON data.
- vlovich123 4y agoI think even structs have this problem because typically you heap allocate all the structs/arrays. You could try to arena allocate contiguous objects in place, but that sounds hard enough that I doubt that anyone bothers. Using a SAX parser is almost certainly the tool you want to use.
- cube2222 4y agoOctoSQL[0] or DuckDB[1] will most likely be much simpler, while going through 10 GB of JSON in a couple seconds at most. Disclaimer: author of OctoSQL [0]: https://github.com/cube2222/octosql https://github.com/cube2222/octosql [1]: https://duckdb.org/ https://duckdb.org/
- Nihilartikel 4y agoIf you're doing interactive analysis, converting the json to parquet is a great first step.. After that duckdb or spark are a good way to go. I only fall back to spark if some aggregations are too big to fit in RAM. Spark spills to disk and subdivides the physical plans better in my experience..
- pradeepchhetri 4y agoWell if you need to convert json to parquet to do anything fast, then what is the meaning ? You will end up wasting way more resource in that conversion itself that your benefit is all equalized in the cost of extra storage utilization (since now you have json and parquet files both). The whole point is to do fast operations in json itself. Try out clickhouse/clickhouse-local.
- closeparen 4y agoIf you're doing interactive analysis, generally you're going to have multiple queries, so it can be worthwhile to pay the conversion cost once upfront. You don't necessarily retain the JSON form, or at least not for as long.
- deleted 4y ago[deleted]
- lmeyerov 4y agoYep! We do the switch to parquet, and then as they say, use dask so we can stick with python for interesting bits as SQL is relatively anti-productive there Interestingly, most of the dask can actually be dask_cudf and cudf nowadays: dask/pandas on a GPU, so can stay in the same computer, no need for distributed, even if TBs etc of json
- berkle4455 4y agoJust use clickhouse-local or duckdb. Handling data measured in terabytes is easy.
- tylerhannan 4y agoThere was an interesting article on this recently... https://news.ycombinator.com/item?id=31004563 https://news.ycombinator.com/item?id=31004563 It prompted quite some conversation and discussion and, in the end, an updated benchmark across a variety of tools https://colab.research.google.com/github/dcmoura/spyql/blob/master/notebooks/json_benchmark.ipynb https://colab.research.google.com/github/dcmoura/spyql/blob/... conveniently right in the 10GB dataset size.
- Groxx 4y agotbh my usual strategy is to drop into a real programming language and use whatever JSON stream parsing exists there, and dump the contents into a half-parsed file that can be split with `split`. Then you can use "normal" tools on one of those pieces for fast iteration, and simply `cat * | ...` for the final slow run on all the data. Go is quite good for this, as it's extremely permissive about errors and structure, has very good performance, and comes with a streaming parser in the standard library. It's pretty easy to be finished after only a couple minutes, and you'll be bottlenecked on I/O unless you did something truly horrific. And when jq isn't enough because you need to do joins or something, shove it into SQLite. Add an index or three. It'll massively outperform almost anything else unless you need rich text content searches (and even then, a fulltext index might be just as good), and it's plenty happy with a terabyte of data.
- kosherhurricane 4y agoWhat I would have done is first create a map of the file, just the keys and shapes, without the data. That way I can traverse the file. And then mmap the file to traverse and read the data. A couple of dozen lines of code would do it.
- jeffbee 4y agoOne thing that will greatly help with `jq` is rebuilding it so it suits your machine. The package of jq that comes with Debian or Ubuntu Linux is garbage that targets k8-generic (on the x86_64 variant), is built with debug assertions, and uses the GNU system allocator which is the worst allocator on the market. Rebuilding it targeting your platform, without assertions, and with tcmalloc makes it twice as fast in many cases. On this 988MB dataset I happen to have at hand, compare Ubuntu jq with my local build, with hot caches on an Intel Core i5-1240P. time parallel -n 100 /usr/bin/jq -rf ../program.jq ::: * -> 1.843s time parallel -n 100 ~/bin/jq -rf ../program.jq ::: * -> 1.121s I know it stinks of Gentoo, but if you have any performance requirements at all, you can help yourself by rebuilding the relevant packages. Never use the upstream mysql, postgres, redis, jq, ripgrep, etc etc.
- jeffbee 4y agoI guess another interesting fact worth mentioning here is the "efficiency" cores on a modern Intel CPU are every bit as good as the performance cores for this purpose. The 8C/8T Atom side of the i5-1240P has the same throughput as the 4C/8T Core side for this workload. I get 1.79s using CPUs 0-7 and 1.82s on CPUs 8-15.
- ginko 4y agoThis is something I did recently. We have this binary format we use for content traces. You can dump it to JSON, but that turns a ~10GB into a ~100GB file. I needed to check some aspects of this with Python, so I used ijson[1] to parse the JSON without having to keep it in memory. The nice thing is that our dumping tool can also output JSON to STDOUT so you don't even need to dump the JSON representation to the hard disk. Just open the tool in a subprocess and pipe the output to the ijson parser. Pretty handy. [1] https://pypi.org/project/ijson/ https://pypi.org/project/ijson/
- maCDzP 4y agoI like SQLite and JSON columns. I wonder how fast it would be if you save the whole JSON file in one record and then query SQLite. I bet it’s fast. You could probably use that one record to then build tables in SQLite that you can query.
- funstuff007 4y agoAnyone who's generating multi-GB JSON files on purpose has some explaining to do.
- ghshephard 4y agoLogs. jsonl is a popular streaming format.
- funstuff007 4y agoI guess, but you can grep JSONL just like you can a regular log file. As such, you don't need any sophisticated tools as discussed in this article. > 2. Each Line is a Valid JSON Value > 3. Line Separator is '\n' https://jsonlines.org/ https://jsonlines.org/
- ghshephard 4y agoYes - 100% I spend hours a day blasting through line json and I always pre-filter with egrep, and only move to things like jq with the hopefully (dramatically) reduced log size. Also - with linejson - you can just grab the first 10,000 or so lines and tweak your query with that before throwing it against the full log structure as well. With that said - this entire thread has been gold - lots of useful strategies for working with large json files.
- ddulaney 4y agoI really like using line-delimited JSON [0] for stuff like this. If you're looking at a multi-GB JSON file, it's often made of a large number of individual objects (e.g. semi-structured JSON log data or transaction records). If you can get to a point where each line is a reasonably-sized JSON file, a lot of things gets way easier. jq will be streaming by default. You can use traditional Unixy tools (grep, sed, etc.) in the normal way because it's just lines of text. And you can jump to any point in the file, skip forward to the next line boundary, and know that you're not in the middle of a record. The company I work for added line-delimited JSON output to lots of our internal tools, and working with anything else feels painful now. It scales up really well -- I've been able to do things like process full days of OPRA reporting data in a bash script. [0]: https://jsonlines.org/ https://jsonlines.org/
- stonecolddevin 4y agoIsn't this pretty much what JSON streaming does?
- ddulaney 4y agoYep, it’s a subset of JSON streaming (using Wikipedia’s definition [0], it’s the second major heading on that page). I like it because it preserves existing Unix tools like grep, but the other methods of streaming JSON have their own advantages. [0]: https://en.m.wikipedia.org/wiki/JSON_streaming https://en.m.wikipedia.org/wiki/JSON_streaming
- klabb3 4y ago+1. While yes, you can have a giant json object, and you can hack your way around the obvious memory issues, it’s still a bad idea, imo. Even if you solve it for one use case in one language, you’ll have a bad time as soon as you use different tooling. JSON really is a universal message format, which is useful precisely because it’s so interoperable. And it’s only interoperable as long as messages are reasonably sized. The only thing I miss from json lines is allowing a type specifier, so you can mix different types of messages. It’s not at all impossible to work around with wrapping or just roll a custom format, but still, it would be great to have a little bit of metadata for those use cases.
- zerop 4y agoOther day i discovered duckdb on HN which allows firing SQL on JSON. But i am not sure if that can take this much volume of data.
- jahewson 4y agoI had to parse a database backup from Firebase, which was, remarkably, a 300GB JSON file. The database is a tree rooted at a single object, which means that any tool that attempts to stream individual objects always wanted to buffer this single 300GB root object. It wasn’t enough to strip off the root either, as the really big records were arrays a couple of levels down, with a few different formats depending on the schema. For added fun our data included some JSON serialised inside strings too. This was a few years ago and I threw every tool and language I could at it, but they were either far too slow or buffered records larger than memory, even the fancy C++ SIMD parsers did this. I eventually got something working in Go and it was impressively fast and ran on my MacBook, but we never ended up using it as another engineer just wrote a script that read the entire database from the Firebase API record-by-record throttled over several days, lol.
- simonw 4y agoI've used ijson in Python for this kind of thing in the past, it's pretty effective: https://pypi.org/project/ijson/ https://pypi.org/project/ijson/
- mlhpdx 4y agoBack in the bad old days when XML consumers hit similar problems we’d use and event based parser like SAX. I’m a little shocked there isn’t a mainstream equivalent for JSON — is there something I’ve missed?
- jahewson 4y agoOh yes, some time ago I wrote a nodejs module to handle large xml files like that https://www.npmjs.com/package/big-xml https://www.npmjs.com/package/big-xml For JSON, given that large files are generally record-based ndjson is the solution I’ve encountered http://ndjson.org/ http://ndjson.org/ and it works nicely with various tools out there using the .ndjson file extension
- taspeotis 4y ago.NET has this built in with Utf8JsonReader [1]. > Utf8JsonReader is a high-performance, low allocation, forward-only reader for UTF-8 encoded JSON text, read from a ReadOnlySpan<byte> or ReadOnlySequence<byte> Although it's a bit cumbersome to use with a stream [2]. [1] https://learn.microsoft.com/en-us/dotnet/standard/serialization/system-text-json/use-dom-utf8jsonreader-utf8jsonwriter?pivots=dotnet-7-0#use-utf8jsonreader https://learn.microsoft.com/en-us/dotnet/standard/serializat... [2] https://learn.microsoft.com/en-us/dotnet/standard/serialization/system-text-json/use-dom-utf8jsonreader-utf8jsonwriter?pivots=dotnet-7-0#read-from-a-stream-using-utf8jsonreader https://learn.microsoft.com/en-us/dotnet/standard/serializat...
- 19h 4y agoTo analyze and process the pushshift Reddit comment & submission archives we used Rust with simd-json and currently get to around 1 - 2GB/s (that’s including the decompression of the zstd stream). Still takes a load of time when the decompressed files are 300GB+. Weirdly enough we ended up networking a bunch of Apple silicon MacBooks together as the Ryzen 32C servers didn’t even closely match its performance :/
- xk3 4y agozstd decompression should almost always be very fast. It's faster to decompress than DEFLATE or LZ4 in all the benchmarks that I've seen. you might be interested in converting the pushshift data to parquet. Using octosql I'm able to query the submissions data (from the begining of reddit to Sept 2022) in about 10 min https://github.com/chapmanjacobd/reddit_mining#how-was-this-made https://github.com/chapmanjacobd/reddit_mining#how-was-this-... Although if you're sending the data to postgres or BigQuery you can probably get better query performance via indexes or parallelism.
- zX41ZdbW 4y ago[flagged]
- 19h 4y agoUnfortunately we're not just searching for things but extracting word frequencies of every user for stylometric analysis, so we need to do custom crunching. Spreading this task into many sub-slices of the files is annoying because the frequencies per user add up quite a lot, which results in quite a massive amount of data.
- Animats 4y agoRust's serde-json will iterate over a file of JSON without difficulty, and will write one from an iterative process without building it all in memory. I routinely create and read multi-gigabyte JSON files. They're debug dumps of the the scene my metaverse viewer is looking at. Streaming from large files was routine for XML, but for some reason, JSON users don't seem to work with streams much.
- rvanlaar 4y agoRecently had 28GB json of IOT data with no guarantees on the data structure inside. Used simdjson [1] together with python bindings [2]. Achieved massive speedups for analyzing the data. Before it was in the order of minutes, then it became fast enough to not leave my desk. Reading from disk became the bottleneck, not cpu power and memory. [1] https://github.com/simdjson/simdjson https://github.com/simdjson/simdjson [2] https://pysimdjson.tkte.ch/ https://pysimdjson.tkte.ch/
- isoprophlex 4y agoIf reading from disk is now your bottleneck, next time put it in a (compressed?) ramdisk if you want to feel particularly clever/enjoy sick speedups
- nn3 4y agothe real trick is to do the debugging/exploration on a small subset of the data. Then usually you don't need all these extra measures because the real processing is only done a small number of times.
- UnCommonLisp 4y agoUse ClickHouse, either clickhouse-server or clickhouse-local. No fuss, no muss.
- kashif 4y agoMight be useful for some - https://github.com/kashifrazzaqui/json-streamer https://github.com/kashifrazzaqui/json-streamer
- 2h 4y agoNote the Go standard library has a streaming parser: https://go.dev/play/p/O2WWn0qQrP6 https://go.dev/play/p/O2WWn0qQrP6
- code-faster 4y ago> Also note that this approach generalizes to other text-based formats. If you have 10 gigabyte of CSV, you can use Miller for processing. For binary formats, you could use fq if you can find a workable record separator. You can also generalize it without learning a new minilanguage by using https://github.com/tyleradams/json-toolkit https://github.com/tyleradams/json-toolkit which converts csv/binary/whatever to/from json
- chrisweekly 4y agoLNAV (https://lnav.org https://lnav.org) is ideally suited for this kind of thing, with an embedded sqlite engine and what amounts to a local laptop-scale mini-ETL toolkit w/ a nice CLI. I've been recommending it for the last 7 years since I discovered this awesome little underappreciated util.
- mattewong 4y agoIf it could be tabular in nature, maybe convert to sqlite3 so you can make use of indexing, or CSV to make use of high-performance tools like xsv or zsv (the latter of which I'm an author). https://github.com/liquidaty/zsv/blob/main/docs/csv_json_sqlite.md https://github.com/liquidaty/zsv/blob/main/docs/csv_json_sql... https://github.com/BurntSushi/xsv https://github.com/BurntSushi/xsv
- liammclennan 4y agoFlare’s (https://blog.datalust.co/a-tour-of-seqs-storage-engine/ https://blog.datalust.co/a-tour-of-seqs-storage-engine/) command line tool can query CLEF formatted (new-line delimited) JSON files and is perhaps an order of magnitude faster. Good for searching and aggregating. Probably not great for transformation.
- DeathArrow 4y agoYou can deserialize the JSONs and filter the resulting arrays or lists. For C# the IDE can automatically generate the classes from JSON and I think there are tools for other languages to generate data structures from JSON.
- reegnz 4y agoAllow me to advertise my zsh jq plugin +jq-repl: https://github.com/reegnz/jq-zsh-plugin https://github.com/reegnz/jq-zsh-plugin I find that for big datasets choosing the right format is crucial. Using json-lines format + some shell filtering (eg. head, tail to limit the range, egrep or ripgrep for the more trivial filtering) to reduce the dataset to a couple of megabytes, then use that jq-repl of mine to iterate fast on the final jq expression. I found that the REPL form factor works really well when you don't exactly know what you're digging for.