2 ms·
Although I've not yet set them up in production, there's a lot of heat these days using a fan out architecture through a message broker (i.e. Apache Kafka), and
by escanda 11y ago
Although I've not yet set them up in production, there's a lot of heat these days using a fan out architecture through a message broker (i.e. Apache Kafka), and ingest that data and transform it into different data models through a stream processor (i.e. Apache Spark), to some file format which will be later be queried, and processed, into an even higher level data model; a much more layered approach than before, which makes sense from an economical point of view since data acquisition is more expensive than data processing and storage.
Here in Spain some private banks are making heavy use of those technologies to replace their reporting originally based on mainframe technology.
Perhaps a more business analyst oriented concept of the data lake may be the semantic layer [1]. This concept may differ from Fowler's in that is not so data oriented, and augments it, but underneath, some of the goals, as providing self service querying facilities to analysts, and making use of as much of the ingested data as possible, are similar.
[1] https://www.veroanalytics.com/blog/its-time-to-unleash-the-semantic-layer https://www.veroanalytics.com/blog/its-time-to-unleash-the-s...