3 ms·
I agree that the spin up time is not 0, but how much that matters is really influenced by the size of your jobs. I would also call out that, when developing ag
by jeffreysmith 11y ago
I agree that the spin up time is not 0, but how much that matters is really influenced by the size of your jobs.
I would also call out that, when developing against Spark, you can and should be running your jobs locally to develop and debug them. So, I would lean on local Spark execution to keep a tight feedback loop. This is one of the big advantages of using Spark over Hadoop or similar more complex systems. Your jobs scale predictably from a local laptop to a cluster.
Sure, there's a cutoff point, but within a range, this is largely a choice. We're a startup, so our traffic is always increasing. We want to worry about sizing our infrastructure as new jobs come on line. There is no perfect static size for our clusters. We've also been scaling down as well as up in the size of our infrastructure as we improve our code's performance and understand our systems better. So ephemeral just makes sense to us for the moment. Certainly we might want a persistent cluster at some point, but dynamically-scaled out clusters will likely still be valuable at that point.
- bcbrown 11y agoThanks for the response. I should note that my experience is with Hadoop, not Spark, which is part of why I was interested in your article.
- jeffreysmith 11y agoSure. The story is somewhat different with Hadoop. But it's really quite feasible to adapt your systems incrementally from Hadoop to Spark. One of the details that I elided in the post is that we're just now decommissioning our last Hadoop machine learning job, nearly a year after moving to Spark. These things take time. That's why it's important that new and useful tools like Spark, Impala, etc. work with the existing Hadoop ecosystem.