6 ms·
> The biggest shift has been towards data lake (store everything) away from data cubes (store aggregates). I don't think this is any shift. The "store everythi
by altdatathrow 6y ago
> The biggest shift has been towards data lake (store everything) away from data cubes (store aggregates).
I don't think this is any shift. The "store everything" has always existed in my experience, that's how the aggregates were built in the first place. The aggregates were for speed and convenience, and you drill-down as necessary, including to the individual record level.
Maybe the shift is people thinking that it's cheaper to just analyze the entire corpus on-demand because we can throw a spark cluster at it?
- hestefisk 6y agoI agree, data warehouse was what the data lake is today. Data cube is the aggregation of data in the warehouse, and then you can drill down and roll up. Difference between warehouse and lake is the emphasis on correctness (one canonical data model) and deduplication of data (when warehouses were invented, storage was expensive so one tried to normalise it into a star schema with as little duplication as possible — when emergence of cheap storage, this is less important and we can spend less time developing fancy ETL processes to make everything fit into one, conformant data model).
- altdatathrow 6y ago> this is less important and we can spend less time developing fancy ETL processes to make everything fit into one, conformant data model And that's precisely why modern data processes are inferior to 20 years ago. People reinvent the wheel over and over and spend massive budgets on unnecessary tech stacks that would be alleviated if the time was simply taken to model the data. A clean data model is about a whole lot more than simply storage space.
- hestefisk 6y agoAgree if people would model everything to a “Grand Unified Data Model” of everything, it would be a lot more efficient... unfortunately that is very hard and puts a massive bureaucracy around data governance and management. It slows things down. I guess a more modern approach is to relax those constraints a bit and realise that some data can be expressed in different ways, and that duplication isn’t too much of a concern because storage is cheap. That said, it’s not an excuse for reinventing the stack over an over. I think the thirst to reinvent has largely been driven by the shift from expensive proprietary solutions like Teradata and Oracle to open source ones. That’s a positive shift.
- haddr 6y agoAnd this is why big enterprises are moving out of data warehouse model for processing all the data and prefer a data lake concept. It is better to centralize some aspects of the data, but definitely not all (like common relational model for all sources, as it is in data warehouse). The data warehouse model quickly becomes super expensive to maintain and evolve.
- dima_vm 6y agoYou miss the critical difference -- nowadays people don't store aggregates, they just scan sharded data very fast. That simplifies a lot of things, because you don't need to keep two databases in sync (raw and aggregated).
- deleted 6y ago[deleted]
- deleted 6y ago[deleted]
- deleted 6y ago[deleted]