3 ms·
I'd advocate starting as simple as possible. It's very context specific, but one very simple architecture to consider if seeing if duckdb can handle your data
by RobinL 4y ago
I'd advocate starting as simple as possible. It's very context specific, but one very simple architecture to consider if seeing if duckdb can handle your data (maybe on quite a large machine).
For larger data, i'm a fan of writing everything in SQL, and executing it using AWS Athena (a Presto engine), which is extremely cheap.
If you have even larger data, you can consider AWS Glue (steering clear of any aws specific stuff and literally using it as spark SQL as a service). For many simple cases you don't really need to know any Spark, and you're just submitting SQL to AWS for execution.
But all of this assumes you're storing your data as files on disk (e.g. parquet). I'm less experienced in running everything in more traditional SQL engines like MS SQL server or postgres.