3 ms·
We are neither trying to denounce distributed processing nor claim that single node engines can solve all the problems. For example, we still haven't discussed
by grt17 8y ago
We are neither trying to denounce distributed processing nor claim that single node engines can solve all the problems. For example, we still haven't discussed how to ingest this amount of data (80 million tuples/sec) or even more than that to a single node. How realistic is this scenario?
I agree that the benchmark used here is not a good example of a real-world streaming application, as it's not even compute-intensive. However, it is still used as the main benchmark for the evaluation of streaming frameworks.
In this blogpost, we are trying to start a discussion about how modern streaming systems perform and how they are supposed to perform. Towards this direction, we should reconsider what we believe is regular. For example, in cases like this, if you can use a single node instead of 5 for your computations, why shouldn't you consider about it...
- scott_s 8y agoI work in the area, and I don't consider the Yahoo set to be the main benchmarks for evaluating streaming frameworks. As a field, we don't have an agreed-upon set. I think a better place to start from are the DEBS Grand Challenges: http://debs.org/grand-challenges/ http://debs.org/grand-challenges/
- grt17 8y agoI agree that there isn't a widely accepted benchmark for streaming frameworks. However, you can find papers accepted recently in SOSP (Drizzle) or SIGMOD (Structured Streaming) --- I think there will be one even in VLDB--- having this benchmark as their main way of comparing streaming systems with each other.