3 ms·
Missed some basic info about distributed systems: Partitioning, bucketing, sorting, bloom filters, the fact that you normally create specific logic to backfill
by crorella 5y ago
Missed some basic info about distributed systems:
Partitioning, bucketing, sorting, bloom filters, the fact that you normally create specific logic to backfill rather than run the same 'daily' process N times.
If you come from Dataswarm/Airflow then you can use JINJA for the templating, UPSERTS are better when you want to save space by not keeping the same data for each snapshot you take day after day (on multi terabyte tables the space savings are considerable).
More topics that are usually prevalent in distributed systems:
- Join order
- Type optimization for smaller tables (ones you can fit in memory to do broadcast joins)
- Field codification for extra compression/smaller intermediate shuffles when processing data.
- Field sorting for extra compression
- Caching tables that participate in several parts of the process.
- Rolling accumulating pre-aggregates for multi-day processing
- T-Digest for performant aggs / cubes
- HLL for performant aggs
- etc