5 ms·
Having worked in a couple large enterprises that tried to make the Data Lake concept work, I would love to see this concept in practice where it actually does s
by baakss 11y ago
Having worked in a couple large enterprises that tried to make the Data Lake concept work, I would love to see this concept in practice where it actually does something useful. Thus far, both attempts I've seen ended up falling back to traditional reporting structures, like data warehouses. This was in financial services and energy.
Serious question, does anyone know any companies that are employing this successfully? And if so, in what fashion? I'd definitely love to hear about a success story and what value was provided.
Edit: The example in the article seems to be more related to the failure of the data warehouse than the success of the data lake.
- rpedela 11y agoI am using the "data lake" concept without realizing it. It just seemed like the right thing to do. I am working on a "Google for SEC filings". There are about 15 million company filings and other data spanning 20 years available on sec.gov. The data is about 700 GB compressed, but unfortunately you have to download each filing individually from their FTP server. When I first started, I wrote a script that would download the filing and then process it into the format I wanted. However their FTP server is very slow and there are 15 million individual downloads, so it was taking forever. Rather I wrote a script that mirrored the FTP server to S3 as fast as possible while still being respectful to their bandwidth and server capacity. And this still took almost 3 weeks. Now I have a "data lake" of raw SEC filings and other data which I can pull from at any time on S3. And the important part is the performance is significantly better so the processing time is relatively small.
- lukateake 11y agoWhat does the bandwidth and storage cost per month?
- rpedela 11y agostorage: 700 GB * $0.03 = $21 per month one-time upload cost (bandwidth free): (15 million / 1000) POSTs * $0.005 = $75 one-time download cost: ((15 million / 10000) GETs * $0.004) + (700 GB * $0.09) = $69
- pbnjay 11y agoOT but funny enough I worked with an accounting PhD student to extract SEC filings and mine for some keywords and associated numbers and tables, from this same FTP and I remember it being so dog-slow. This was like 8 years ago too, sad its in the same shape.
- stonecupi 11y agoThe most successful one I'm familiar with is the one employed at Netflix (I worked there on BI). However, they did not abandon the concept of a data warehouse, they just enhanced it with the data lake. Probably about 80% of ad hoc analytics come out the data lake while standard reporting needs are covered by a Teradata dwh. When ad hoc queries become regular needs they build aggregates in the data lake and move it to teradata or redshift where it can be sliced and diced along various dimensions. I see this same strategy being attempted at some non-tech companies as well. Too early to say whether it will succeed.
- phunge 11y agoI've worked on such a system before and am a fan of the idea. First I've heard of the term Data Lake though. The main benefit as I see is when data sources are external, with ill-defined or ambiguous schemas. Often when you fit data into an ETL pipeline, you find out issues at the output of the pipeline, but the fixes need to happen way upstream. Often this involves rebooting the entire process entirely and renormalizing all your data somehow. If you delay interpretation and normalization to later in the processing pipeline (i.e. in the system I worked on, we did it lazily at interpretation time), then doing smarter things with the data is a matter of changing code -- and it's a lot easier to ship fixes to code than to ship fixes to data!
- mynegation 11y agoI have seen it implemented in one of the financial services companies. Both data warehouse and and data lake type databases lived on MS SQL servers (with the usual production/testing division), and nightly batch processes performed clean up from data lake to data warehouse. It kind of worked. The usual problems with these are: (1) getting hold of external clean up functionality. Sometimes it was stored with the database itself as stored procedures, but sometimes it was not. Remedy: make sure clean up functionality is readily available. (2) dealing with new kinds of "dirt" (i.e. problematic data in the data lake). Remedy: this is unavoidable. The best you can do is to insert lots and lots safety checks and diagnostics into your clean up programs
- bmh100 11y agoI have successfully personally implemented a data lake, as well as touched other companies' data lakes. In the system I designed, I have dozens of data sources feeding in the same "type" of data, but all in different formats and terminology. My multi-stage system applies transformations, including business logic and data cleansing, to each source individually. Then the data sources are combined and linked to other information sources as a unified data model. Consumer applications [1] can take the unified data model and quickly get up and running with for similar analytic applications that maybe just need different views. Or more customized applications can take the transformed data of a single source, apply additional logic that incorporates the raw data, and then incorporate other information sources. In this way, the system is flexible and open, while providing solid data governance. Overview: (1) Data Source -> (2) Data Source Transformation -> (3) Unified Data Model -> Dashboard A consumer application could acquire data from (1), (2), (3), and/or other data sources. [1]: In this context, "consumer application" means an data transformation process or a data presentation application, not anything to do with the B2C market sector.
- escanda 11y agoAlthough I've not yet set them up in production, there's a lot of heat these days using a fan out architecture through a message broker (i.e. Apache Kafka), and ingest that data and transform it into different data models through a stream processor (i.e. Apache Spark), to some file format which will be later be queried, and processed, into an even higher level data model; a much more layered approach than before, which makes sense from an economical point of view since data acquisition is more expensive than data processing and storage. Here in Spain some private banks are making heavy use of those technologies to replace their reporting originally based on mainframe technology. Perhaps a more business analyst oriented concept of the data lake may be the semantic layer [1]. This concept may differ from Fowler's in that is not so data oriented, and augments it, but underneath, some of the goals, as providing self service querying facilities to analysts, and making use of as much of the ingested data as possible, are similar. [1] https://www.veroanalytics.com/blog/its-time-to-unleash-the-semantic-layer https://www.veroanalytics.com/blog/its-time-to-unleash-the-s...
- bjt 11y agoAt work we built something that's maybe halfway in between the data lake and data warehouse. It's working well for us. The basic setup: - All data is CSV or json-document-per-row text files on Amazon S3. - We have a Django web application that keeps track of metadata (which dataset lives in which S3 bucket/folder, who uploaded it, the names of the columns in that dataset, and their types). - The REST API in the Django web piece can provide temporary signed S3 URLs that can allow anyone in the company to create a new dataset and upload the files to S3. - The REST API also provides Hive and Redshift "CREATE TABLE" commands for all datasets (which it builds from the columns/types data stored in the DB). We've talked several times about open sourcing it but haven't gotten around to making that happen yet.