14 ms·
Sorry nothing positive to say here. I’ve been using dbt and MDS for nearly 3.5 years and I believe the entire approach is profoundly broken. There’s really not
by dm03514 3y ago
Sorry nothing positive to say here.
I’ve been using dbt and MDS for nearly 3.5 years and I believe the entire approach is profoundly broken. There’s really nothing “modern” about it, especially compared to software engineering.
https://on-systems.tech/blog/135-draining-the-data-swamp/ https://on-systems.tech/blog/135-draining-the-data-swamp/
Building on Extracted operational data is hard at best and a business altering security liability at worse.
I believe The MDS trails at least a decade behind modern software engineering practices, lacking industry guidance and generic tooling to support: CI/CD, versioned deployment artifacts, zero downtime deployments, unit testing, observability, monitoring and alerting.
MDS Data engineering is a meme in the industry, the expectation of “100%” combined with the lack of modern tooling makes success really hard to achieve.
- civilized 3y agoI was a fan of dbt for a while, but the shine wore off when I saw one of my smartest coworkers try to use it. His development speed was about an order of magnitude slower than what I expect from data scientists using dataframe packages like R's dplyr, Python's polars, or Spark DataFrames. In my own experience, dbt is significantly better than writing raw SQL, but still nowhere near a normal software development experience. It needs an IDE badly, but the current IDE is cloud-only. I am currently building my analytics pipelines on top of R's dbplyr package, which allows me to run dplyr operations on database tables just as if they were local data frames. Since it's R, it's not without warts, but at least I get to use a real programming language and the free and fantastic RStudio IDE.
- tomnipotent 3y ago> allows me to run dplyr operations on database tables Which require that data take a round trip between the R process and the database. It's not uncommon for these jobs to spend more time in the read/write step than doing meaningful work, and why I prefer dbt when possible to keep transformations as close to the data as possible.
- leledavid 3y agoThe dbplyr package takes your R code and executes it in your database, returning a remote dataframe. This is quite useful because there is no mvoving data back and forth.
- nograpes 3y agoThe second point the grandparent made was that they were using dbplyr, which allows you to avoid having the data take a round trip between the R process and the database.
- golergka 3y agoWe built the perfect IDE for dbt at Deep Chanel, with autocomplete automatic real time type checking for the whole project. Sadly, the company closed this June, and even the website is already down.
- neighbour 3y agoYou need to find a way to release this product. I would pay to use it.
- golergka 3y agoIt's already been released and publicly available for free for almost a year.
- CRConrad 3y agoSo where is it, if "even the website is already down"? ETA: Seems it isn't. https://www.deepchannel.com/ https://www.deepchannel.com/ That what you meant?
- golergka 3y agoHmm, may be I've had network issues last time I checked. Anyway, that's the only place you can get it — it's a completely free desktop app, but it's not open source.
- civilized 3y agoThis is extremely interesting. You might be onto something.
- jochem9 3y agoImo dbt sits in the analysts space. SQL is the language they know and dbt is the best solution to scale that. In that sense it should only be used for last mile data transformations. Not for wrangling raw data into neat tables.
- fifilura 3y agoI am curious, what do you perceive as the difference between data transformations and "wrangling raw data into tables"? Edit: I am also curious why you consider SQL to be inferior for the latter?
- jochem9 3y agoThe difference is that raw data can come in many shapes (e.g. impossibly nested jsons) and unexpected quality (changing field names...). I cannot easily work with this in SQL and definitely not write tests to cover the ever increasing complexity of dealing with messy data. After that step the data is a lot more uniform, so then it's easy to use SQL.
- camgunz 3y agoThe core divide in this space is SQL vs. $OTHER_LANGUAGE. If you're fluent in SQL you'll be great at dbt. If you're only fluent in not-SQL you won't. A lot of these discussions are basically "can we please just use the language/tooling we know", and sometimes the answer's yes! A lot of people breathe a big sigh of relief when they discover you can use like, Spark through Python and what-not. Regardless of whether or not it's a good idea or good use of resources, this space is advancing because the market demands it. But like, I don't think you're ever gonna use dbt with not-SQL; it's all in on SQL. Feels like maybe it was the wrong fit for your team. I should say I'm not sure if your coworker was an SQL person so dunno if this directly applies. Just saying what my experience has been.
- itsoktocry 3y agoDbt isn't a substitute for pandas/R to a data scientist. They are complements. Dbt is for the data transformation pipeline that prepares the data so the DS can write simple queries using their favourite tools on curated data.
- gigatexal 3y agoY’all don’t have CI/CD? Maybe it’s hyped but we do simple stuff. Snowflake schema DWH. Jinja2 templated SQL or dbt. Airflow with a monolith tool that does transforms and such. Is it perfect? No. Is it understandable? Very much so. Testing is such an interesting concept in data engineering. One needs consistent test data. We aim to implement that with snapshots eventually but now we have sanity checks at each layer. > lacking industry guidance and generic tooling to support: CI/CD, versioned deployment artifacts, zero downtime deployments, unit testing, observability, monitoring and alerting. Ci cd is doable and easy: changes are deployed via GitHub actions, the CI part is a bit missing I guess without tests. Versioned artifacts also, add a tag to the airflow job to know what tagged version of the code is running. Zero downtime deployments we do that all day — I mean using views we can a/b deploy changes to the underlying tables and then do a simple schema change to the view and nobody knows the difference. Unit testing still yet to be done. Observability and alerting we use internal dashboards and sentry. What am I missing?
- Exoristos 3y agoRelevance and actual value, I'd assume.
- antupis 3y agoTesting and monitoring in every aspect are the only things that are greatly lacking, you always need to glue something together to get nice tests or some kind of data contract monitoring.
- tianzhou 3y agoFWIW, we are building a CI/CD solution snowflake https://www.bytebase.com/docs/tutorials/database-change-management-with-snowflake-and-github/ https://www.bytebase.com/docs/tutorials/database-change-mana...
- xyzzy_plugh 3y agoThe mistake everyone makes is treating the entire space as if it's somehow different than the rest of software engineering, when it is precisely exactly the same.
- datavirtue 3y agoThis. Nearly everyone is confused.
- esafak 3y agoMachine learning people have cottoned on to this, so MLOps was invented.
- tomrod 3y agoThe training part of MLOps is important to ensure replicability and other desirable properties or the ML artifact, but the rest is clearly good CI/CD and observability.
- wodenokoto 3y agoWe had an application engine come in and run an ML project. The first thing he did was remove real data from training, as he was shocked to see that development wasn’t done against dummy data. And once the dev environment is running on real data, the changes in how you develop and operationalize just seems to cascade in my experience.
- fifilura 3y agoI am slightly confused by your reply. Can you elaborate a bit, was it good or bad?
- wodenokoto 3y agoYou can develop a database schema for your application on dummy data, and test your business logic on dummy data. But you can't develop a machine learning model on dummy data. It is best practice to keep live data out of your development environment for normal software development, but it is impossible for machine learning projects.
- kgdiem 3y ago> CI/CD, versioned deployment artifacts, zero downtime deployments, unit testing, observability, monitoring and alerting. This definitely stood out to me when I was working with a lot of Snowflake’s new products, especially their “native apps”. I started on some bespoke tooling given the lack of everything that you mentioned and at the time was thinking about how to productize it / turn it into some open source tools but I’ve not gotten back around to it since leaving that job. I had a really great experience working with sqlglot from Tobiko Data https://tobikodata.com/ https://tobikodata.com/ — haven’t had a chance to check out their sqlmesh project yet but I have some faith in their work.
- camgunz 3y agoYeah seconding sqlglot, has saved me _tons_ of time. I think yeah, they're smart and onto something here.
- benjaminwootton 3y agodbt was a huge leap forward in this regard. It enables source code control, modularity, testing, CI/CD, environments, documentation, git branching, local development experience. Not a fanboy, but it's who reason for being is basically to enable "software engineering practices for data".
- antupis 3y agoYup, it is not perfect but I will take dbt every day against the old way as some random DDL files in git.
- CRConrad 3y agoSo basically its huge advantage is that it saves its code as text, which makes it gittable? (Dang, and here the tool I've been designing in my head for several years was going to claim that niche... So unique. ;-)
- camgunz 3y agoNah, the main thing is the dependency resolution. So you make pipelines like: - employees - tall_employees - tall_employees_by_salary These tables (dbt would call them models because they can also be views depending on how they're configured) depend on each other, so you have to build them in a specific order. Without something like this, you're manually running SQL to build your pipelines, and that's very error prone. With something like this you can run tests (they're also SQL in dbt, which is pretty common in data engineering teams), run your pipeline in CI, etc. There's a bunch of other features, but they all hang off this basic idea.
- riku_iki 3y ago> you're manually running SQL to build your pipelines you can have some sh file which runs SQL in desired order..
- 3y ago
- itsoktocry 3y agoA surprising number of companies (even large, well known ones) don't even have basic dashboards and metrics, most of what you mention is irrelevant to most companies.