3 ms·
Why not just use GNU Parallel (or something similar) instead of Spark?
by doobwa 11y ago
Why not just use GNU Parallel (or something similar) instead of Spark?
- elyase 11y agoI think this could have been done with GNU parallel. One advantage I see with Spark is that is that it is easier to interact with Python, for example these two lines are all is needed to call the relevant Python function: urls = sc.parallelize(batched_data) labelled_images = urls.flatMap(apply_batch) So if you already have a cluster with Spark installed (like Databrick does) then it takes less work to just call your Python code than setting up a GNU Parallel cluster and a writing a small wrapper script. Additionally a Python script would have to load/init the models on every call from Parallel. I agree that this is not a great demonstration of Spark main strengths.
- orm 11y agoI think one reason would fault tolerance. Is there a fault tolerance layer in GNU parallel? last time I checked their homepage ( a few minutes ago), there was no reference to fault tolerance. Another reason is, perhaps, scheduling.
- chimtim 11y agowhat fault tolerance does spark give you in this scheme? It cannot look into TF progress and checkpoint all state. Using Spark with TF, seems like an overkill -- you need to manage and install two framework what should ideally be a 200 line python wrapper or small mesos framework at most.
- ole_tange 11y agoDoes --retries count as fault tolerance?