3 ms·
Except that doing very, very simple stuff, like counting 10 rows, on a laptop is excruciatingly slow.
by Bootvis 3y ago
Except that doing very, very simple stuff, like counting 10 rows, on a laptop is excruciatingly slow.
- qsort 3y agoWell I mean, that's a problem with all those solutions. If your dataset is so small it fits in RAM it's simply ridiculous to reach for those tools, and you don't start seeing real advantages until you have truly massive datasets. Anecdotally using Spark on less than 1TB just means wasting your time, but in Spark's defense, that's not what it's for.
- iamcreasy 3y agoDo you mean 1TB of data? or cluster with total memory capacity less than 1TB?
- qsort 3y agoI mean data size. It's anecdotally the threshold where "normal" ETL methods start having problems. Obviously performance is hard and the only correct answer is "it depends", but I have never seen Spark improve things on datasets that small.