3 ms·
1) Buy a box with 1TB of RAM. Very doable these days, albeit a bit pricey still 2) Scale out. Spark can easily handle hundreds of nodes in a single cluster. Ag
by tupshin 12y ago
1) Buy a box with 1TB of RAM. Very doable these days, albeit a bit pricey still
2) Scale out. Spark can easily handle hundreds of nodes in a single cluster. Aggregate RAM across all of them can be used.
3) Cache intermediate data sets and/or hot data sets, as opposed to the entire data set.
- sitkack 12y ago256GB machines are the sweet spot right now.
- ironchef 12y agoFor #2 and #3, see here: http://spark.apache.org/docs/latest/programming-guide.html#rdd-persistence http://spark.apache.org/docs/latest/programming-guide.html#r... We typically do #2 at our company and it's been fine so far. The bigger issue isn't a single 1 TB data set. It's multiple large data sets as one then must handle the data shuffles during joins, etc. The ability to keep the RDD in memory through the operations still tends to beat normal long winded hadoop operations anyways...
- rs_atl 12y agoLong winded in more ways than one. I would still use Spark even if it were slower than Hadoop, just to get a sane API.