3 ms·
I think his point is that bloated, over-engineered Big Data systems—whether batch or streaming—are overkill for the vast majority of problems.
by edw 10y ago
I think his point is that bloated, over-engineered Big Data systems—whether batch or streaming—are overkill for the vast majority of problems.
- placeybordeaux 10y agoThere are just many points that don't really apply to stuff like spark or tez that runs on YARN: ex: Hadoop << SQL, Python Scripts I completely agree with Mapreduce << SQL, Python Scripts I do a lot of my processing on sparkSQL and through RDD transformations as opposed to Mapreduce limiting, slow KV style processing.