5 ms·
Most use cases currently served by Apache Spark clusters would run 10x faster on a laptop with fast SSD (Macbook perhaps), and an ad hoc cat/grep/sed pipeline /
by 988747 4y ago
Most use cases currently served by Apache Spark clusters would run 10x faster on a laptop with fast SSD (Macbook perhaps), and an ad hoc cat/grep/sed pipeline /s
- glogla 4y agoThere's a DuckDB and Polars and similar tools now, which can finally outperform the venerable unix tools. Once the data is on the laptop you can get order of magnitude faster execution than Spark. The unsolved problems are 1) what if the data and what you do with it suddenly doesn't fit on a laptop (giving everyone 64 GB RAM laptops for example seems like a waste) and 2) how do you deliver the relevant subset of the data from the petabyte place where you store it to the laptop. If someone could solve that, Spark could finally go to hell.
- 988747 4y agoThe solution to that is called Snowflake