2 ms·
If the files are sorted, merge in Python (`heaps.merge` or `sorted` on chunks plus `itertools.groupby`). If they’re not sorted, a single pass is mostly out of t
by goodside 7y ago
If the files are sorted, merge in Python (`heaps.merge` or `sorted` on chunks plus `itertools.groupby`). If they’re not sorted, a single pass is mostly out of the question — you need out-of-core sorting. GNU sort with intermediate compression of temporary/intermediate files is hard to beat there for performance, but if your join keys are more complicated than can be specified in GNU sort you need a more involved pipeline. If raw IO is actually the bottleneck, that’s usually a sign you need better compression — look into Zstandard and lz4, especially for intermediate files. If IO is still the bottleneck, consider possible sharding keys (Python ‘hash’ plus modulo works well, or CRC32 etc.) that would let you divide the problem up among a large number of workers. Everything is easier if you get the data into S3 storage early.