3 ms·
I think a lot of people are considering it. The ideas behind Apache Beam are supposed to make the transition easier, with the ability to run in either mode. I t
by artwr 8y ago
I think a lot of people are considering it. The ideas behind Apache Beam are supposed to make the transition easier, with the ability to run in either mode. I think in practice it really depends on your use case. For a large enough amount of data, especially on the ingest side, running it streaming makes sense. Now for that daily report from aggregated table, do you really need streaming ETL/aggregation is another question. The operational cost is non trivial with streaming.
- pgwhalen 8y agoCan you expand on that last sentence? Not that I’d disagree, I’m just curious about what specifically you’re referring to. My team is currently underwater supporting a decades old system of batch ETL, where it’s a major challenge to understand dependencies between jobs, and therefore what to rerun when something fails. We’re exploring a move to streaming ETL (Kafka streams being the leading candidate) as a way of making the dependencies explicit: rather than having to rerun stored procs or Talend jobs that we depend on certain data, we run a fleet of 24/7 microservices that transform and move data downstream whenever it is available.