3 ms·
After you shard some of the computation out to different processors to filter before single core aggregation, then you can push down some pre aggregation to lim
by lsb 6y ago
After you shard some of the computation out to different processors to filter before single core aggregation, then you can push down some pre aggregation to limit the communication between parallel threads, and it keeps going!
Super interesting to see parallelism for reading logs this way, like for when you have too many logs for one core and too few for Spark
- timbray 6y agoThat turns out to be hard because in your typical log file you have bursts of requests localized to a particular part of the file that get something into the top-N without even appearing in other segments. So essentially you can end up having to do a complete N-way merge between the segments to make sure you really find the top hits.