4 ms·
Basically software for batch processing a tonne of data. eg. importing a SAP feed into a database, or loading a bunch of csv files, or like processing a bunch
by shadowmint 8y ago
Basically software for batch processing a tonne of data.
eg. importing a SAP feed into a database, or loading a bunch of csv files, or like processing a bunch of images...
...anything where you have to convert data from some source through a series of steps (typically a DAG) into some useful output.
However, its often misused.
For example, if you have a trivial amount of data, or trivial process an ETL is over engineered cruft where a simple script would do.
So there are many (rubbish) ‘simple ETL frameworks’ which offer zero value and technical complexity for no benefit.
Unless you need multiple servers processing data through multiple steps and you need the auditing and process control... you probably don’t need an ETL, just a simple script.
- cosmie 8y ago> you probably don’t need an ETL, just a simple script. +1 > Unless you need multiple servers processing data through multiple steps and you need the auditing and process control I'll stress the "multiple servers" part. You can add in a substantial amount of multiple, sequential steps and auditing and process control in a simple script. The part that adds orders of magnitude worth of complexity and operational overhead and points of failure is being able to distribute it to multiple servers. Distributed architectures are operationally and architecturally expensive. And far more often than not, completely unnecessary for a given use case.
- bduerst 8y agoSome ETLs these days are simple scripts - e.g. Spark, Dataflow - with their config files.
- cosmie 8y agoBeing able to define an ETL workload within a simple script is not the same as your ETL system itself being a simple script. While I love both Spark and Dataflow, both of them are incredibly complex distributed systems with very high operational costs. Someone, somewhere is paying a lot of money to have an operational resource maintain that complexity. Whether you have an internal devops resource doing so or you're using a managed service, you're paying for that complexity somehow. And, for a lot of workloads, you aren't actually getting any more value than you would from standing up a ~$50/month standard Debian/Ubuntu server and a set of simple scripts on it.
- Mironor 8y agoYou don't have to have a cluster to run spark scripts, setting master to `local` (and running it on one machine) is often enough for small anounts of data.
- bduerst 8y agoThey don't have high operational costs - you can run them as a script on your local machine. You're making them out to be more complex than they really are.
- debacle 8y ago> batch processing Not necessary. ETL these days can be streamed, realtime, etc.