3 ms·
I've worked on such a system before and am a fan of the idea. First I've heard of the term Data Lake though. The main benefit as I see is when data sources are
by phunge 11y ago
I've worked on such a system before and am a fan of the idea. First I've heard of the term Data Lake though.
The main benefit as I see is when data sources are external, with ill-defined or ambiguous schemas. Often when you fit data into an ETL pipeline, you find out issues at the output of the pipeline, but the fixes need to happen way upstream. Often this involves rebooting the entire process entirely and renormalizing all your data somehow.
If you delay interpretation and normalization to later in the processing pipeline (i.e. in the system I worked on, we did it lazily at interpretation time), then doing smarter things with the data is a matter of changing code -- and it's a lot easier to ship fixes to code than to ship fixes to data!