7 ms·
We need new data books, so we started one: Cloud Data Management
- sarcasmatwork 7y agoYay, free ebook. Thanks!
- DataDaoDe 7y agoI completely agree with the sentiment here. As an engineer tasked with building data driven systems and architectures, I'm well aware of the amount of buzzword / enterprise nonsense floating around in this space - and it is enormous. You can couple that with the fact that its really difficult to find any practical books and resources that you can actually apply to solve real-world problems when working in small-teams and on quick deadlines. Getting business to act in a data driven and analytic way, and building the architecture for it - is a non-trivial task. This looks like it could be one of those rare good resources in the space. Great work by the guys from chartio!
- thingsilearned 7y agoThanks! We're truly looking for this to be community driven as well. So if you see places to contribute, or where you might disagree, or where you could share a story - do let us know or make a pull request on GitHub! Besides the need for a new data book, we realized that it needed to be of a different format, as the space is moving so fast the expertise is very distributed.
- ryantuck 7y agoI've recently dug into the Agile DWH Design and The DWH Toolkit books and the design tips in them all seem really compelling, for good reason. Though as I've actually started modeling, I've found that the creation of "proper" fact/dimension tables has felt at times like overkill, given the technologies we're using (BQ / Postgres / Looker). So, perfect timing! Really looking forward to checking this out.
- momirlan 7y agoRediscovering the wheel much ? You guys are starting again like 30 years ago, with simple queries and think all that was built is overkill. When you mature and get into complex analytics and joining data from 20 sources you will understand all that was built. Just you wait...
- thingsilearned 7y agoWe're definitely not trying to start from scratch or throw out all the old knowledge/practices - just update them for the common data stacks used today. In the book we use much of the old terms and recommendations. Most of the high level organizing is still totally right - but a lot of the optimizing and work done for performance and cost reasons is very different now. For example ELT makes now much more sense than ETL for the reasons Kostas wrote about here: https://dataschool.com/data-governance/etl-vs-elt/ https://dataschool.com/data-governance/etl-vs-elt/ And many things previously done for cost and performance reasons are just not relevant anymore thanks to the big innovations in C-Store warehouses.
- supercanuck 7y agoELT has been around for 15+ years. Inmon refers to it as a Persistent Staging Area (PSA) in an Enterprise Data Warehouse. The difference now is, you have Hadoop and cloud providers that will take credit cards and give you as much space as you can pay for. The concept is not new, it was just cost was a factor back then because capacity was fixed and memory was expensive. the only thing that has changed is the commoditization of hardware has allowed for different behaviors that would have been cost prohibitive.
- thingsilearned 7y agoYup! And that commoditization of hardware has made it really inexpensive to have a Data Lake, where you first put all your data in raw format (so you only need to do EL - and not T all in the same step). And then, because of the way C-Store sources like Redshift are built it makes a ton of sense to just do your T step as a set of Views (materialized or not) onto of that Data Lake. It allows you to not do E & T & L all together. It's really nice (less complex, easier to implement, less costly, and more flexible) to have that T part pulled out and done after.
- rilt 7y agothis is awesome. would have loved to have this book when we were building this at my previous co. very approachable in how it explains each segment on the whole and zooms in on them individually. would have saved our team a lot of time as we came to very similar conclusions but over the course of a few months
- acak 7y agoOn an unrelated note, would anyone be able to tell me which blog engine (and theme) their website is running on?
- thingsilearned 7y agoWe use Jekyll. It's a custom design from our own awesome Steven Lewis.
- acak 7y agoRather well designed. Thanks!
- deleted 7y ago[deleted]
- the_watcher 7y agoCool! The framework makes a lot of sense, and articulates what I've observed very well.
- scruple 7y agoI'm in the data space today. It is indeed a very confusing space to be in. This is a fantastic set of resources you've created here. Well done!
- evandev 7y agoIs this book available in print at all?
- thingsilearned 7y agonot yet. Just PDF. But we'll be working on getting it into print eventually.
- iblaine 7y agoI see Panoply is part of this effort. Panoply creates a lot of spam on hackernews, reddit, twitter, and elsewhere. As an example, doing ETL with drag n drop tools, and not in code, is a dying skill in the industry. [edit] Basically I'm wary of companies commenting on data standards, where those same companies also sell a product in the data industry. You're probably going to find more honesty on guidelines from open source contributors like LinkedIn, Netflix, Sitchfix, etc.
- camel_Snake 7y agoI'm surprised to see dbt make that list of yours - I don't think I've ever seen them spam the online communities I visit.
- timwis 7y agoHas anyone found any similar guidance around master data management? Matching/deduplicating, and feeding back into source systems.
- jameslk 7y agoOne of the experiences that stuck with me most going from doing software engineering stuff to data engineering stuff is the very difficult and sometimes complete lack of tooling around testing and debugging issues in SQL and pipelines. As software engineers, we have debuggers, unit testing, mocking frameworks, e2e tools, BDD languages like Cucumber, code quality tools, etc. But when you're working on a pipeline and you're wondering what will happen if you run it, sometimes your best bet is actually to just run it and wait 20 min for it to complete. Or run some portion of it ad hoc. Or if you want to know what a SQL query is going to do in the black box of your query engine, you might try to parse the esoteric language of a query explainer. The best tools seem to be available after deployment such as data quality tests, dashboards and alerts. I think there's a lot of opportunity to improve the ecosystem.
- thingsilearned 7y agoI totally agree. There has been some progress here recently. Have you checked out DBT and their testing features? https://docs.getdbt.com/docs/testing https://docs.getdbt.com/docs/testing
- hinkley 7y agoOccasionally I ask what the world would look like if we had a web browser that was designed from the first to be testable. Would e2e testing be simple and repeatable? I could just as easily ask the same about databases, or dozens of other tools in our toolchain. Makes me wonder if there's space for something like SQLite and testable.
- closeparen 7y agoProbably Datomic. Databases are hard to test because they're big piles of global mutable state. A database designed around values, pure functions, and other FP constructs should be inherently easier to test.
- heroHACK17 7y agoWriting this comment as I 'actually run it and wait 20 min for it to complete.'
- brucej 7y agoGreat to see more information about data out there for people to learn from. BTW my colleges and I created a declarative (SQL) open source (MIT) framework for Apache Spark to make ETL and ML super easy, if you want to check it out it you can read more here: https://arc.tripl.ai https://arc.tripl.ai . We've recently started combining this with Argo and Delta Lake which is working well for us in the source, lake to warehouse stages.
- neilobremski 7y agoThis is just lovely. I'm the lead of a company's data warehouse project, having inherited it from someone who was a DBA that read a book and then made the whole thing ... I tend to joke that I'm a digital janitor and now I'm a big data janitor. Well, things are going well enough, but cleanup on a live system with users and zillions of reports is incredibly difficult. Part of the issue is isolation and size: too many have too much access to too much data. It's overwhelming and it also leads to a lot of high cost because queries aren't understood and they're run against massive datasets. I've been educating myself as well as I can between fixing errors and this book is just the thing I need to calm my nerves. It totally makes sense how the stages are set and even just this blog post overview has given me ideas of how to carve some things up. Kudos, cheers, and all that. Happy Halloween!
- wenc 7y agoI worry that this promotes the data architecture philosophy that is currently in vogue in tech companies, but is actually a bad fit for many traditional enterprises. Most of the time, data architecture needed depends on use case. Tech/web companies deal with massive amounts of unstructured/semi-structured data ingested at a fast clip, so the architectural thinking here works. However I would argue that many traditional enterprises whose major sources of data are primarily highly structured (SQL databases), a lake is actually not needed. A young data engineer working in a traditional enterprise, enamored with the idea of data lakes say, might try to ETL SQL databases into an object store (add RBAC etc.) only to rebuild it back out into a data warehouse. This will almost always turn out to be the wrong approach. The simpler and more manageable approach is actually to federate existing databases, add cataloging etc. and not even use a data lake at all.
- kfk 7y agoI think by lake they mean s3 and similar alternatives (like hdf is for hadoop). This is not crazy as processing s3 files is quite easy. What you say is true but the problem is that no analyst deals only with one db, they have to deal with an increasing number of data sources, that’s what the lake is for. I also don’t believe you need a warehouse after the lake and that one source of truth is pipe dreams, so I’m with you this architecture is not 100% what companies probably need.
- wenc 7y ago> no analyst deals only with one db, they have to deal with an increasing number of data sources Which is why I mentioned federation. Most data sources in an enterprise are SQL native or accessible and it does not make sense to dump a SQL database into S3 just to be able to combine it with other forms of data. Federation means you can operate across multiple databases (eg do Cross database joins etc.)