3 ms·
>By being able to construct on-demand clusters programmatically that auto-terminate on completion, we’ve been able to use ephemeral clusters for all our data jo
by bcbrown 11y ago
>By being able to construct on-demand clusters programmatically that auto-terminate on completion, we’ve been able to use ephemeral clusters for all our data jobs. For much of the day, we can have very few data processing clusters running at any given time. But periodically, we spin up many large clusters via EMR that train all of the models that we need to learn. This usage pattern is neither harder to implement nor more expensive than a serial execution of all of our jobs and matches our preferred workflow much better. For our usage pattern, this actually represents a large cost savings over a persistent cluster.
I recently started working at a startup with a similar approach. My previous company had a persistent cluster in a colo, and I found that to be much easier to develop against and use. It takes a significant amount of time to spin up a cluster, which slows down the feedback loop, and makes the development process longer.
I also wonder about cost. It seems to me that if you have enough jobs running every day, at some point you should be able to have them scheduled such that a persistent cluster has fairly high utilization, such that it becomes cheaper than spinning up ad hoc clusters. Have you done any analysis to show that ephemeral clusters are cheaper? Is there a cutoff point where that is no longer true?
- jeffreysmith 11y agoI agree that the spin up time is not 0, but how much that matters is really influenced by the size of your jobs. I would also call out that, when developing against Spark, you can and should be running your jobs locally to develop and debug them. So, I would lean on local Spark execution to keep a tight feedback loop. This is one of the big advantages of using Spark over Hadoop or similar more complex systems. Your jobs scale predictably from a local laptop to a cluster. Sure, there's a cutoff point, but within a range, this is largely a choice. We're a startup, so our traffic is always increasing. We want to worry about sizing our infrastructure as new jobs come on line. There is no perfect static size for our clusters. We've also been scaling down as well as up in the size of our infrastructure as we improve our code's performance and understand our systems better. So ephemeral just makes sense to us for the moment. Certainly we might want a persistent cluster at some point, but dynamically-scaled out clusters will likely still be valuable at that point.
- bcbrown 11y agoThanks for the response. I should note that my experience is with Hadoop, not Spark, which is part of why I was interested in your article.
- jeffreysmith 11y agoSure. The story is somewhat different with Hadoop. But it's really quite feasible to adapt your systems incrementally from Hadoop to Spark. One of the details that I elided in the post is that we're just now decommissioning our last Hadoop machine learning job, nearly a year after moving to Spark. These things take time. That's why it's important that new and useful tools like Spark, Impala, etc. work with the existing Hadoop ecosystem.