5 ms·
To analyze and process the pushshift Reddit comment & submission archives we used Rust with simd-json and currently get to around 1 - 2GB/s (that’s including th
by 19h 4y ago
To analyze and process the pushshift Reddit comment & submission archives we used Rust with simd-json and currently get to around 1 - 2GB/s (that’s including the decompression of the zstd stream). Still takes a load of time when the decompressed files are 300GB+.
Weirdly enough we ended up networking a bunch of Apple silicon MacBooks together as the Ryzen 32C servers didn’t even closely match its performance :/
- xk3 4y agozstd decompression should almost always be very fast. It's faster to decompress than DEFLATE or LZ4 in all the benchmarks that I've seen. you might be interested in converting the pushshift data to parquet. Using octosql I'm able to query the submissions data (from the begining of reddit to Sept 2022) in about 10 min https://github.com/chapmanjacobd/reddit_mining#how-was-this-made https://github.com/chapmanjacobd/reddit_mining#how-was-this-... Although if you're sending the data to postgres or BigQuery you can probably get better query performance via indexes or parallelism.
- zX41ZdbW 4y ago[flagged]
- 19h 4y agoUnfortunately we're not just searching for things but extracting word frequencies of every user for stylometric analysis, so we need to do custom crunching. Spreading this task into many sub-slices of the files is annoying because the frequencies per user add up quite a lot, which results in quite a massive amount of data.
- e12e 4y agoFrom https://github.com/ClickHouse/ClickHouse/issues/22482#issuecomment-812283862 https://github.com/ClickHouse/ClickHouse/issues/22482#issuec... it looks like a local load into clickhouse is expected to take 6-7 hours (in 2017?). I wonder how clickhouse-local would fare today (I'm guessing the dataset is so big, that load/store - then analyze would be better....).
- zX41ZdbW 4y agoMost of the time is spent in decompression - the source dataset used to have files in .bz2, which is the main contributor to total time. The dataset itself is just around 10 billion records.