3 ms·
I fully agree with your experience here. We have a super hard mission at Chartio - it's not just about the interface but also how the data is setup. The inter
by thingsilearned 7y ago
I fully agree with your experience here. We have a super hard mission at Chartio - it's not just about the interface but also how the data is setup. The interface, being as flexible as it is and also enabling full schema (instead of dataset) browsing is a pretty big part though in also allowing a more agile version of data modeling. It had that very much in mind and we've written a book (soon to be published with Wiley) on proper modern data governance techniques.
https://dataschool.com/data-governance/ https://dataschool.com/data-governance/
Our next phase is to help people get to that cleaner source of truth much more quickly than traditional dimensional modeling approaches. Tools like Visual SQL and DBT (https://www.getdbt.com https://www.getdbt.com) are really changing the complexities here.
- jeanloolz 7y agoI would love to be involved with that book. I'm a data engineer myself and I have built SQLBucket, a python library I have been told is similar to getdbt.com (although I'm not familiar with DBT, the similarities have been mentioned to me on various occasions). Shameless plug: http://github.com/socialpoint-labs/sqlbucket http://github.com/socialpoint-labs/sqlbucket
- thingsilearned 7y agoOh awesome! I haven't heard of SQL bucket but I'll check that out as I love anything that encourages SQL based modeling. I was writing my own and then DBT came out and we push people there primarily. Send me a note - dave-at-chartio and I'd love to chat sometime.
- mulmen 7y agoI'd love to see the data modeling book. I spend a lot of energy shouting the virtues of The Data Warehouse Toolkit into the void. You are right it is outdated but it isn't entirely (or even mostly) wrong. My coworkers seem much more interested in making a bigger EMR or adding nodes to Redshift than designing a reasonable data mart because "star schemas don't scale". I'm interested to see what you come up with, it is a huge gap in the current literature.
- thingsilearned 7y agoIt's free to download on dataschool.com. I agree on the gap and that's why we wrote this book. We launched that on HN here 5 months ago https://chartio.com/blog/cloud-data-management-book-launch/ https://chartio.com/blog/cloud-data-management-book-launch/
- mulmen 7y agoOh wow, I missed this entirely, thank you for pointing it out and I will give it a read tonight.
- gpu_explorer 7y agoVery interesting what you wrote in your article. Most interesting is how you seem to realize while designing your product that the spreadsheet surface is the most intuitive to users. They like also the baked results you present quickly. So you can see really the problem of your customer then. What is really good is to assemble a library of visual queries for the customer. This is a good idea for the reason that many users have the same fundamental types of queries on their data. When finally you have enough of the basic queries that the user can do useful work without programming then you can find a way to customize this yes. Have you data on how many similar queries customers use? Then you should know how to create the basic set of important operations.
- thingsilearned 7y agoI wish I could say that we had that insight from the beginning and that it worked right away, but we ended up failing into that realization after a lot of designs and prototypes. So what you describe is somewhat built in to what we have now. Users still have to choose what columns they want to look at (there's no real way for us to guess that) and then we do apply some knowledge on what type of data they're looking at to help them get to what they're likely looking for. We've also tried at times to make default dashboards for data sources when people connect. We can do this to some extent with known Schemas like connecting GA, SalesForce, Hubspot, etc, but for databases - that's proven to be a largely impossible task so far. Everyone's data is so different, and have such odd conditions to consider filtering by, that the auto dashboards end up being quite useless.
- iagovar 7y agoIm one of those users, and I tend to use HeidiSQL or DBeaver for exploring.
- hobs 7y agoYeah, doing ETL everyone wants me to use their data and perform magic with it, but I really want to see their stored procedures and queries running on top of it before I can understand the best way to "auto dashboard" anything. Finding out the distinct set of values in any column helps a lot, referential integrity helps a lot, but without those queries its pretty dang hard.