2 ms·
> as long as setting up a cluster is a bit easier! If you're on AWS, it's already quite easy to set up a cluster today, no? There's EMR of course, and there a
by nchammas 10y ago
> as long as setting up a cluster is a bit easier!
If you're on AWS, it's already quite easy to set up a cluster today, no?
There's EMR of course, and there are tools like spark-ec2 [0] and Flintrock [1].
There are a few more tools listed on Spark Packages that target different cloud providers [2], too.
Disclaimer: I am the primary author of Flintrock and am a contributor to spark-ec2.
[0] https://github.com/amplab/spark-ec2 https://github.com/amplab/spark-ec2
[1] https://github.com/nchammas/flintrock https://github.com/nchammas/flintrock
[2] https://spark-packages.org/?q=tags%3Adeployment https://spark-packages.org/?q=tags%3Adeployment
- minimaxir 10y agoSetting up Spark clusters is easy relative to setting up clusters, but not easy enough yet to set up and configure relative to simply downloading a package in R/Python.
- nchammas 10y agoWell, the absolute easiest way to run Spark is to do it locally (e.g. you can brew install it on a Mac and just go) or to pay for a proprietary service like Databricks, which makes setting up a cluster take a few clicks. That said, I think `flintrock launch my-cluster` is almost as easy as doing `pip install ...`. You do need an AWS account and you do need to set your preferences like region and key name in a config file, but I don't see how you can get out of doing even that without subscribing to some managed service like Databricks that abstracts everything away and replaces it with a nice Web UI.
- mgummelt 10y agoIf you're running a DC/OS cluster, it actually is! > dcos package install spark This gives you a Spark CLI, and installs the Spark Cluster Dispatcher so you can run jobs async. It also has options for installing the history server, and for SSL/Kerberos support. Disclaimer: I work at Mesosphere and contribute to the DC/OS Spark package.