4 ms·
One thing that’s not clear to me is why is there so much competition and crowding in “data massage” space. There is Snowflake, there are all kinds of ETL tools.
by fireeyed 6y ago
One thing that’s not clear to me is why is there so much competition and crowding in “data massage” space. There is Snowflake, there are all kinds of ETL tools. The customer lists these startup posts have overlaps. Is it just Marketing departments inside these companies playing around with these tools or the CIOs cycling through the hottest startup on TechCrunch list ?
- lmeyerov 6y agoin my experience, most ai projects die before the ai part people hate hiring data engineering (plumbing people feels like cost), and data eng like tools that work but most are.too niche/happypath-oriented, so even w trifacta etc, a lot of open territory. SW can solve a lot of that, in theory, so everyone wins. And I agree that until there is an oss winner, the proprietary stuff will keep getting churned through. So ultimately whatever your data platform does (aws, databrick, whatever) or oss you're bringing. A lot of room for vendors to carve out niches b/c of connectors x use cases, until platforms/oss eats them all. VC's will see some ARR and name brands and thus be happy to fund: a lot of gaps any startup can fill. (I am impressed by airbyte for a few non-technical reasons even without having used it, so not a knock on them, so just some clues for the continuing froth in their market.)
- ABeeSea 6y agoThey all promise to reduce your data engineering budget. The problem is that building a data connector is a one-time platform problem per data source. Once it’s solved; it’s solved. None of them solve the problem of ETL design and data warehousing design.
- edmundsauto 6y agoIt sounds like you don't think solving for data connectors + necessary maintenance has a lot of value. I would agree, not FTE levels of value, but most companies I've seen in the SMB space would do well to pay $1-3k per month to have their data all housed in one spot. That lets their 1-2 DS/DE/SWE spend their time actually analyzing the data. Maintaining connectors is also a good way to demotivate high achievers - better to have them further down the value funnel.
- linkjuice4all 6y agoComing from an ad agency background I’ve seen a lot of attempts at “unifying” various data sources from client’s analytics and sales data, agency tools, and third party data sets that are all in different formats, date ranges, and scopes. Warehousing that data might also require firewalling clients or teams for privacy or “competitive/conflict” reasons. These aren’t difficult problems to solve with a few knowledgeable devs but that is nothing but added cost and some agencies just aren’t good at hiring the right devs - especially if their previous exposure has been basic front end web developers from their clients. “Data warehouse” has also become a selling term even if “really big database” is a more accurate term. Hopefully more of these companies start to distinguish themselves in this space but their competition isn’t each other - it’s entry-level data people blasting through Excel.
- spennant 6y agoAd agencies have neither the stomach nor the business model to support hiring devs at market rate.
- neumann 6y agoIt is pretty crazy. I worked for a large organisation where management was far closer to 'technology leaders' and 'technology strategists' than engineering and data science principles and leads. They would endlessly swoop in to our division asking us to assess another product they have bought to fix the legacy problems of multiple data sources. All of them were brittle af. They all anticipated a very idealistic data source and the absence of non-technical people curating data in excel ten different ways. Even though we were the data science team, we usually ended up providing far more value to the organisation because we could do data engineering and cleaning and ended up being the source of truth for a lot of data required by the wider organisation. We got pitched dozens of sexy solutions to fix all our ETL problems, but when we started asking questions it was always seemed like a well designed custom pipeline couldn't be beaten for both data quality assurance, reliability and speed.
- mtricot 6y agoThat's exactly why we are approaching the problem with open source. It changes the dynamic of how it gets adopted. we've been in your shoes where a tool is being pushed Top-Down and now you have to deal with a super complex, super expensive, rigid & half working product. Instead Airbyte gets adopted by engineers, data scientist... to solve one problem and then the usage expands from there. We can improve the product based on the feedback we get from the real users. And if a feature, a connector is not there, anyone can actually add it!
- marcinzm 6y agoThere is value to having all your data in one place and managing data connectors is generally not very fun. They also tend to be annoyingly brittle and it's very visible when they go down. That all makes for a perfect recipe to cause burnout in a small data team. Competent data engineers are also fairly difficult to hire right now.