11 ms·
Is the "modern data stack" still a useful idea?
- tomrod 3y agoIntriguing that the CEO of DBT declares MDS ("modern data stack") to be a meaningless term (i.e. no longer useful). Maybe a good demarcation that we are entering the post-modern phase of cloud tech generally? Cloud still there, but not the only option. I look forward to this happening with GenAI -- the tools we come up with will be pretty cool in the long run. I hope we can find ways around platform enshittification because that really sucks. Similar for blockchain tech. I see a common thread in these techs too -- the same type of LinkedIn influencers shop these during their hype phases. Ultimately, useful technology and best practices come out of it!
- actionfromafar 3y agoWell put! Nothing's modern anymore, the tech cycles are closing in on one another so much the old one hasn't gone out of phase until it's overcome by not one but several new ones. In a nice touch of self-refentiality, your message itself is very post-modern.
- tomrod 3y agoI hadn't considered that. Thanks for pointing that out to me :)
- Williams77 3y ago[flagged]
- karakanb 3y agoDisclaimer: I am the co-founder of a competitor, Bruin (https://getbruin.com https://getbruin.com). We are exactly the kind of integrated platform Tristan is talking about. The article resonates with me a lot, and it is because we called this out months ago. I find the idea of Tristan walking back on the premise of MDS and claiming it to be not useful anymore funny because they were one of the main drivers of the term and the whole hype around it. He even acknowledges that they played ball with other companies in the space: > There was a lot of valuable co-marketing, partnership deals, co-sponsored events, and co-selling. This had real value for everyone involved—customers and vendors alike. Companies voluntarily integrated their products together, cross-promoted each other publicly, and built partnerships that made owning and operating these technologies far easier for customers. Sorry, but no, this didn't have value for the customers, only for the vendors. The customers were left alone by themselves to deal with all of this complexity, and the vendors made a lot of money off of that. They convinced companies that they needed a bunch of different tools to build a simple pipeline, and rode on the wave of huge valuations based on these same ideas that they are walking back on. The companies prefer integrated solutions now because they woke up. Instead of paying 150k/y each to Fivetran, dbt and whatnot, they realize that they are better of just hiring an engineer or two in the worst case. It is 2024, and none of these tools still properly talk to each other. Do you want to get an end-to-end lineage of your data? Good luck with that. How about quality? How about governance? Companies are left alone with this hype cycle. I'd claim that a significant part of the blame lies on the executives and leaders in the companies, who just jumped on the ship for the sake of building their CVs and skipped the critical thinking step. None of them seriously asked themselves the question of whether or not it makes sense. To be honest, I feel sorry for all the money spent on building solutions around all of these hype-driven products.
- neeleshs 3y agoDisclaimer : I'm the cofounder of Syncari. I agree 100%. The whole notion of modern data stack is a myopic , neither here nor there philosophy to address data in my opinion. Just take a look at the sheer number of tools one has to run - ETL(or more commonly, data dumpers), Wearhouse, transform tool,"reverse ETL", data quality tools, BI. And if you want some ML, that's whole another game. We also believe strongly in an integrated approach to this space, with a data model at the center of it all.
- gigatexal 3y agoDeleted
- tomnipotent 3y agoHe's not "walking back on the premise of MDS", but acknowledges that the term "MDS" had been co-opted and lost its descriptive fidelity. > Sorry, but no, this didn't have value for the customers, only for the vendors Absolutely untrue. I've discovered multiple vendors over the last two decades because of co-marketing and partnership deals, and I know many others that have. My success rate with these vendors is no different from vendors I found myself or were referred to me by others. > The customers were left alone by themselves This is true with every vendor integration with every company on the globe. You either 1) have internal expertise to push through the pain, 2) hire support from the vendor itself (often X hours come baked into the deal), or 3) hire a consultant. This is as true for the MDS vendors as it is for Salesforce, AWS, or an ERP. It's true for your company, too. > They convinced companies that they needed a bunch of different tools Except you do? You need a database, you need something to extract data and move it locally, you need something to transform that data into something useful, and you need the ability to deliver that data to end users. This was true twenty years ago, true ten years ago, and still true today. The all-in-one vendors prior to MDS were horrible, and I'd rather cut off my arm then use something like Pentaho again. I'm also not convinced the current all-in-one vendors are any better. One of the advantages of the pick-your-stack MDS was that you could tailor the tools to you organizations specific needs, rather than have to fight against whatever your mega-vendor offers. I loved the fact that I could use open source for 90% of my stack, and then use PowerBI for the last 10%. Or I could be 100% FOSS and deliver data with Metabase. Or I could use Tableau for 50%, and FOSS for the other 50%. > they realize that they are better of just hiring an engineer or two in the worst case This simply isn't true for every size or class of company. Many businesses are better off with vendor-supported platforms than trying to build something bespoke, which can quickly become a liability with staff turnover or lack of strong internal technical leadership. > It is 2024, and none of these tools still properly talk to each other What does this even mean? How does a data extract tool "talk" to a transformation tool? Or "talk" to the database? I can see from the screenshot on your website that it offers a GUI with drag-and-drop - is that your definition of "talking"? Or are you referring to a tight integration that doesn't require additional work to glue them together? > How about quality? How about governance? These are management and organization issues, not tooling. I've seen amazing quality and governance with teams working in MSSQL/SSIS, and I've seen horrible quality and governance from teams using Hadoop/Spark. A tool can make a competent practitioners job easier, but it's not going to make an incompetent competent.
- extr 3y agoWhat would HN commenters define as a truly modern data stack right now? Mostly asking what this chart https://a16z.com/emerging-architectures-for-modern-data-infrastructure/ https://a16z.com/emerging-architectures-for-modern-data-infr... would look like in 2024.
- mr_toad 3y ago> What would HN commenters define as a truly modern data stack right now? Just don’t spend too much time debating what it is, or the analysts who need numbers yesterday will develop their own solutions before you’ve even published your A3 architecture diagram. A bird in the hand, as they say.
- rch 3y ago- Postgres and Kafka - Airflow and Flink - Kudu+Impala or Clickhouse - Iceberg/Parquet or Delta - Ozone(HDFS) or Ceph - Spark Rapids and/or Ray+Metaflow - Open Metadata or Atlas+Ranger
- te_chris 3y agoIT DEPENDS. But for standard ecom/ops stuff (and this is heavily biased by the fact this is what I know and have deployed with low overhead): Big query. This is fed by an ETL pipeline of the stuff you need. dbt to ELT into useful, reportable tables. Something to visualise - I still like Mode, but there's lots of new BI toys around. KISS. This is extremely low maintenance at the actual requirement levels of most businesses with < 50 employees.
- 3y ago
- dm03514 3y agoSorry nothing positive to say here. I’ve been using dbt and MDS for nearly 3.5 years and I believe the entire approach is profoundly broken. There’s really nothing “modern” about it, especially compared to software engineering. https://on-systems.tech/blog/135-draining-the-data-swamp/ https://on-systems.tech/blog/135-draining-the-data-swamp/ Building on Extracted operational data is hard at best and a business altering security liability at worse. I believe The MDS trails at least a decade behind modern software engineering practices, lacking industry guidance and generic tooling to support: CI/CD, versioned deployment artifacts, zero downtime deployments, unit testing, observability, monitoring and alerting. MDS Data engineering is a meme in the industry, the expectation of “100%” combined with the lack of modern tooling makes success really hard to achieve.
- civilized 3y agoI was a fan of dbt for a while, but the shine wore off when I saw one of my smartest coworkers try to use it. His development speed was about an order of magnitude slower than what I expect from data scientists using dataframe packages like R's dplyr, Python's polars, or Spark DataFrames. In my own experience, dbt is significantly better than writing raw SQL, but still nowhere near a normal software development experience. It needs an IDE badly, but the current IDE is cloud-only. I am currently building my analytics pipelines on top of R's dbplyr package, which allows me to run dplyr operations on database tables just as if they were local data frames. Since it's R, it's not without warts, but at least I get to use a real programming language and the free and fantastic RStudio IDE.
- tomnipotent 3y ago> allows me to run dplyr operations on database tables Which require that data take a round trip between the R process and the database. It's not uncommon for these jobs to spend more time in the read/write step than doing meaningful work, and why I prefer dbt when possible to keep transformations as close to the data as possible.
- leledavid 3y ago
- williamcotton 3y agoIt seems that there’s a philosophy with these kinds od cloud data services: measure everything then figure out what to measure after. It seems that the better approach is to make a hypothesis first, collect only the data you need, and then analyze. Otherwise it seems the indirect costs of retooling around “measure all the things indiscriminately” and the direct costs of paying for such services don’t seem worth it. Can someone offer a reasonable rebuttal?
- staticautomatic 3y agoIt doesn't always make sense to "collect only the data you need." For example, someone in our org will ostensibly have a good reason for setting up an event in Google Analytics. As the head of analytics I have no use for their event but I'm going to end up ingesting it into BigQuery anyway because the built-in GA4->BQ transfer sends ALL the data, and I'm fine with that because it's damn near free for me to ingest and store. I could instead "collect only what I need" by running the transfer through an ELT tool and filtering out that event, but why would I bother doing that work and paying the ELT vendor money when Google will do it for free and charge me like $1K/year to store 100M rows of everything collected in GA4?
- williamcotton 3y agoThat cost analysis seems incomplete. How much did it cost to implement in the first place? That seems like millions of dollars in capitalization before service expenses are even considered!
- staticautomatic 3y agoSorry but I can't tell if you're being sarcastic or not. Google Analytics and BigQuery are both nominally free, the service expense is $1K/year, and the implementation cost literally a few minutes of my time.
- williamcotton 3y ago
- Animats 3y agoOh, it's something for the ad industry. Ads are watched only by people who don't have good ad blockers.
- hn72774 3y agoDBT in simplest terms is just a way to orchestrate transformations in dependency order. Table "A" needs to be updated before table "B," because "B" selects from "A." It works well for that. I've seen it used to get wild west SQL logic into version control. To replace scheduled SQL workbooks running entire data pipelines. I've also seen over-engineered integrations with other orchestration tools in the "stack." Once a DBT project grows to a certain size, roughly 500-1000 .SQL files, it gets hard to manage. Not impossible, though it takes intentionality about how to group things together and scale them out operationally. "Slim CI" is a recent buzzword that has some nice ideas about build automations. Cosmos is supposed to solve a lot of the automation gripes with dbt. I haven't tried it yet but would like to. https://www.astronomer.io/cosmos/ https://www.astronomer.io/cosmos/
- gms 3y agoI've been in data and analytics for over a decade and co-founded a consolidated-ETL company (https://www.polytomic.com https://www.polytomic.com). Nice to see Tristan realising this (people should be commended for changing their minds). There were three problems with this term: 1. It mostly resonated with VCs and industry observers whose jobs are to peddle in buzzwords, rather than users who simply want their problems solved. 2. It was (and is) ill-defined. Ask a group of people in the industry to define it and you're guaranteed to get different answers. 3. It committed the cardinal sin of using the adjective 'modern' in a noun. At some point, everything today stops being modern. Couple this with (2) and your term is now even more meaningless. At a mundane level there's no drama here: just another example in the long list of noise contributors to the free-money party that was going on during the Covid years (as the essay acknowledges).
- Lyngbakr 3y agoI've worked in data for a while now as an Data Engineer, Data Scientist, and Director and working with experienced software engineers has highlighted to me that most of the data stack is fluff. All the layers/tools that are heaped upon one another just lead to complexity and dependence on paid services. I'm not advocating for reinventing the wheel, but rather that a bespoke solutions seem to be worth the cost. In short, I'd prefer to spend the budget on experienced developers who can build maintainable systems than on a plethora of MDS tools. YMMV, though.
- Waterluvian 3y agoIf you wouldn’t mind: in your experience how often does the problem seem to stem from “we might need that later…” or “we don’t have a concrete set of questions we want to answer?” I have considerably less experience but so far I have experienced a complicated data science stack being used as a way to handle a ton of data without clear questions and desires on how to act on possible answers. As if they’re searching a haystack for… something?
- Lyngbakr 3y agoIME, it has usually stemmed from "that's how everyone else is doing it, so we should too" without considering specific use/business cases. Senior management were simply not willing to deviate from the norm.
- tharkun__ 3y agoAnd also: I need a job. My job/workplace has these tools. They are usable for my job and if I were to question the tools that would come on top of my workload, so I use them even if I know that cheaper tools were used at my last job. Do I care if they pay my salary and the tools are usable?
- tadfisher 3y agoThat's perfectly reasonable, but it sounds like you're saying the tools work for your use case. So long as the executive/management level is evaluating the cost-benefit of other tools, and is receptive to voices from ICs, then there's no problem to solve. It's just not what the GP is describing. In some orgs you end up with these expensive, standard-ish "suites" that are way too generalized, don't solve business needs as efficiently as other tools, are painful for ICs to use to meet requirements, and it seems like no one can really explain why these tools are being used over something else other than rumors (the previous CTO came from this vendor, we got a sweetheart deal 10 billing cycles ago, etc.). So really this an issue of ownership; someone needs to be accountable for infrastructure decisions and outcomes, and it's way healthier for an org if ICs are aligned with management on the same. It sounds like you're not in that sort of org, so good for you.
- deleted 3y ago[deleted]
- datadrivenangel 3y agoThe Modern Data Stack is Dead! Long Live the Analytics Stack!\ Interesting to see one of the strongest popularizers of the term acknowledging this shift. The free money is over, so we need to get back to work and dbt is the least bad way to organize lots of SQL for data management.
- staticautomatic 3y agoI'm a newly minted head of analytics who transitioned from a different domain, so I never had to muck my way through the MDS but attentively watched others from the sidelines over the last few years. Best I can tell, "the modern data stack" is just a marketing phrase invented by a cadre of vampire vendors. The lessons I learned watching others translated into a few simple requirements for our nascent "stack" that most importantly include transparent pricing I can reason about and divvy up, as many integrations as possible so I can minimize rolling my own, and a straightforward framework for ETL code. These three requirements plainly disqualify most of the MDS universe. With the benefit of starting basically from scratch and not having to mess around with real-time analytics, it's pretty easy to ignore the MDS vendors. So far I've landed on BigQuery, AirByte, GitHub, BI Engine, Looker Studio, and Pandas 2.x or DuckDB for local stuff. I send as many things as possible straight to BQ, lock junior analysts out of gigantic tables, archive periodically to partitioned parquet files in cold storage, use mostly turnkey integrations, and ruthlessly prioritize custom ETL jobs. Putting GitHub in the mix isn't super ergonomic and we may be in the market for new tools once we cross the "big data" frontier, but that'll be a while from now. I'll probably never know or care what the MDS vendors think I'm missing.
- iamacyborg 3y agoAirByte was definitely one of the MDS vendors. They literally got a huge investment because their deck bragged about having a few thousand Github stars and a few thousand folks in a Slack channel.
- cookie_monsta 3y agoI can't tell if this is envy or scorn
- benjaminwootton 3y agoI don't think it's either. Modern Data Stack is a set of characteristics - cloud based, SaaS, consumption based pricing, ELT bias, open source bias, component based etc. AirByte is all of them. It's just a fact rather than a criticism.
- jonmoore 3y agoThe Modern Data Stack / MLOps product space was succinctly described by one actually-technical CEO as "vending into ignorance"; the author corroborates this with a commendably candid take: >Imagine it’s 2021, peak MDS, and you meet the CDO of a large bank. “Oh cool,” she says, “you’re the CEO of a tech company. What does your product do?” What do you say? >“We build a tool that leverages the power of the cloud to apply standard SQL and software engineering best practices to the historically mundane (but critical!) job of data transformation.” >“We’re the standard for data transformation in the modern data stack.” >I will tell you that, empirically, option #2 is more effective. This tallies with what I've seen from a lot of enterprise CxOs and their teams as technology hype moved from big data and block chain and onto data science/machine learning. There is so much to write about this, but I'll just recommend "Life Cycle of a Silver Bullet" http://freyr.websages.com/Life_Cycle_of_a_Silver_Bullet.pdf http://freyr.websages.com/Life_Cycle_of_a_Silver_Bullet.pdf, which deserves more attention than it's had on HN.
- lulznews 3y agoIs MLOps more or less of a thing than prompt engineer?
- SheddingPattern 3y agoMore of a thing, but it's mostly DevOps.
- jochem9 3y agoMLOps is deploying, monitoring and (re)training ML models. Sits in the DevOps and data engineering space. Prompt engineering is making generative AI do what you want by crafting the right context. I would put it somewhere in the software and data engineering space, given they will most likely integrate applications with it. MLOps comes into play if you have your own trained or tuned model.
- Annatar 3y ago[dead]
- m0llusk 3y agocompiled a quick list of terms to help understand this piece: MDS - modern data stack BI - business intelligence Redshift - amazon redshift data warehouse service product handles large scale data sets and database migrations can handle analytic workoads on big data sets with column oriented approach built on top of massive parallel processing (MPP) from data warehouse company ParAccel (later acquired by Actian) ETL - extract, transform, load ELT - extract, load, transform -- an alternative to ETL that stores raw data Looker - does BI, started with looker data sciences, acquired by google Tableau - data visualization focused on business intelligence clickstream data - user website navigation records snowflake - an MDS data platform solution mongo - nosql database datadog - monitoring, analytics for devops confluent - real time data streams databricks - web based cluster management and data lakes for machine learning meme-ification - reduction of an idea to cartoons peak - highest point preceeding drop off fivetran - data ingestion dbt - dbt labs, data transformation CDO - chief data officer co-marketing - integrated marketing such as linked brands co-sponsored - cooperative sponsorship co-selling - shared sales assets including data records ARR - annual recurring revenue ZIRP - zero interest rate policy private multiples - private company valuation calculations, metrics forward revenue - revenue expected in the future PowerBI - microsoft business analytics
- bigger_cheese 3y agoI think you need to go a step further and define a few more terms like what you mean when you say "Analytics" and "business intelligence" For example >PowerBI - microsoft business analytics I don't think I would classify Power BI as an analytics tool - In my limited experience with using it I have found that it has some nice ways to build reports and visualize data graphically but only very rudimentary statistical functions and is not really suited for performing data analysis. My organization is travelling along the "modern data" pipeline so I'm grappling with some of the same confusion myself about what everything is in the new world... In the "old data" world I'm familiar with - I would use stuff like SAS or R to perform ad-hoc analytics i.e. analyze data using statistics, perform regressions, Timeseries forecasting, distribution analysis etc. For reporting and visualizing data: that would fall to things like excel, Cognos, ggplot2, Proc report and Some javascript plotting packages like d3js etc. To Extract data would involve running SQL query against databases (Oracle, DB2 or Microsoft SQL Server etc directly) directly then using sas or R to transform the data as needed. Modern approach seems to be Databricks to ingest data into datalake, Python via Juypter notebooks to perform analytics hosted on Azure ML studio running in cloud on a "compute instance" for another level of Indirection...Then Power BI to visualize and build reports.