4 ms·
I witness the overengineering regarding "big" data tools and pipelines since many years... For a lot of use cases, data warehouses and data lakes are only in th
by tobilg 2y ago
I witness the overengineering regarding "big" data tools and pipelines since many years... For a lot of use cases, data warehouses and data lakes are only in the gigabytes or single-digit terabytes range, thus their architecture could be much more simplified, e.g. running DuckDB on a decent EC2 instance.
In my experience, doing this will yield the query results faster than some other systems even starting the query execution (yes, I'm looking at you Athena)...
I even think that a lot of queries can be run from a browser nowadays, that's why I created https://sql-workbench.com/ https://sql-workbench.com/ with the help of DuckDB WASM (https://github.com/duckdb/duckdb-wasm https://github.com/duckdb/duckdb-wasm) and perspective.js (https://github.com/finos/perspective https://github.com/finos/perspective).