2 ms·
We currently use Spark for generating reports that scan the same data set multiple times. We got a 20-30x performance gain over hadoop+hive. Spark's ability t
by danjo 15y ago
We currently use Spark for generating reports that scan the same data set multiple times. We got a 20-30x performance gain over hadoop+hive. Spark's ability to keep data sets in memory works very well for our use case.