2 ms·
a) You can easily run Spark jobs on a single box. Just set executors = 1. b) The reason centralised clusters exist is because you can't have dozens/hundreds of
by threeseed 1y ago
a) You can easily run Spark jobs on a single box. Just set executors = 1.
b) The reason centralised clusters exist is because you can't have dozens/hundreds of data engineers/scientists all copying company data onto their laptop, causing support headaches because they can't install X library and making productionising impossible. There are bigger concerns than your personal productivity.
- rr808 1y ago> a) You can easily run Spark jobs on a single box. Just set executors = 1. Sure but why would you do this? Just using pandas or duckdb or even bash scripts makes your life is much easier than having to deal with Spark.
- cgio 1y agoFor when you need more executors without rewriting your logic.
- this_user 1y agoUsing a Python solution like Dask might actually be better, because you can work with all of the Python data frameworks and tools, but you can also easily scale it if you need it without having to step into the Spark world.
- threeseed 1y agoBut Dask is orders of magnitude slower to Spark. And you can still use Python data frameworks with Spark so not sure what you're getting.
- rpier001 1y agoRe: b. This is a place where remote standard dev environments are a boon. I'm not going to give each dev a terabyte of RAM, but a terabyte to share with a reservation mechanism understanding that contention for the full resource is low? Yes, please.