11 ms·
Building a data team at a mid-stage startup
- waynesonfire 5y agoTLDR, refine your thoughts.
- oliv__ 5y agoRefine your mind
- simonw 5y ago"This is basically a (somewhat cynical) depiction of things that may happen at a lot of companies early in the data maturity stage" I don't think this is very cynical at all! Feels pretty accurate to me.
- nerdponx 5y agoThis is an incredibly valuable writeup. Great job.
- ttz 5y ago> MBA types I chuckled. Then cried, because at least his MBA types can use SQL. My MBA types use Excel. OT: Good article. Like and agree with the push for centralizing data first, then building outwards so external teams can move towards self-service.
- munk-a 5y agoBuilding a good process into your company to receive a query, execute it against a read-only database, and shovel the results back to the user as a CSV file will pay dividends and is, honestly, pretty trivial in most cases.
- jaggederest 5y agoBlazer is my go-to for this kind of thing: https://github.com/ankane/blazer https://github.com/ankane/blazer Pretty easy to set up and share queries, dashboards, whatever
- ttz 5y agoFunnily enough, this is what I did, except I built an app where I write the queries as "pre-built" parameterized ones (sanitized, of course). People still do a bunch of stuff in Excel, though, and every once in a while, it breaks, and I have to dig through the mess. Excel is great when it's just for yourself and you can manage it... it's a pain when others have to figure out someone else's.
- herodoturtle 5y agoI'm an MBA type that studied math and computer science, and for a living programs distributed database solutions. I chuckled too.
- div3rs3 5y agoDone well (like here), The Goal like storytelling, is both educational and interesting.
- gumby 5y agoGreat article. The confusion about what team does what is priceless...yet so common! To provide some sympathy for the folks already working there: you always replace systems well after you've overrun them. When the ad hoc system works (consider that google spreadsheet at a time when there were three support people and perhaps a dozen customers) you're not going to decide to replace it with something more complicated. Then you're busy growing so you just keep the system going through sheer force of will. You only replace it when the effort is unbearable; at that point you say, frustratedly, "I wish we'd done this sooner."
- IMTDb 5y agoWhat would be the name of the position/profile of someone in charge of building the data warehousing architecture/ETL pipelines? I my view, they need make sure the warehouse model is a correct representation of the business and that it can be leveraged to answer basic or not-so-basic questions using SQL. They also need to promote it's usage internally by ensuring it is accessible and easy to use and guide other team to a more data oriented mindset. I feel that this is a specialised position not exactly similar to a developer, but every time I look for "data scientist" I get guys that want to do machine learning prediction models, which is not exactly the same stuff either.
- sischoel 5y agoWhat about "data engineer"? There seem to be a lot of jobs for that title nowadays.
- skrtskrt 5y agoYeah we would call this Data Engineer (likely Senior level or up for someone that has had experience building multiple data warehouses) plus the DevOps/SRE work required to stitch all the architecture together
- sjg007 5y agoThe bigger issue is adaptability.. can you migrate schemas preserving older clients, typically that’s by providing a decent middleware…. SQL views are one way, APIs are another etc… All of that while improving performance.
- teej 5y agoA new role has arisen in the last few years that captures much of this responsibility - Analytics Engineer. This article by Claire Carroll describes the role and motivation for it https://www.getdbt.com/what-is-analytics-engineering/ https://www.getdbt.com/what-is-analytics-engineering/
- tmp_anon_22 5y agoMost common would be a DevOps or SRE on an observability team.
- herodoturtle 5y agoFor the last 15 years I've been building (what I consider to be) accessible database solutions, for a bunch of different industries. This sentence from the article resonated with me: > You're starting to lay the most basic foundation of what is most critically needed: all the important data, in the same place, easily queryable.
- correlator 5y agoThank you for writing this. I personally just walked into a very similar role and this rang really true. This article made me realize how much more effort I need to put into the data culture side of the role.
- mindvirus 5y agoThis is a wonderful article, thank you for sharing. I really like the narrative of bringing people with you on the journey, and celebrating the small wins that lead to a good long term outcome.
- czep 5y agoThis is so eerily familiar I swear I've had many of these exact conversations word for word. The only way this doesn't turn into a complete nightmare of a cluster is if the exec team "gets it". If so, you just might stand a chance at building a data team that gels with the rest of the org. But if the exec team simply hired you for window-dressing, expect to be treated like a scapegoat and a punching bag. Any mistakes will be your fault. Any wins will be to the credit of the business. The Director of Product will ask to "embed" dedicated DS headcount and you won't have any real power to shape the roadmap. If the exec team doesn't give you equal footingf with Product (or Marketing, Finance, and Eng for that matter) then this will rapidly become a soul-sucking job. However, if E-team does give you the authority to call Product's bullshit, and tell Finance to stuff it, and not take direction from Eng leads, then you actually might be able to accomplish something really cool.
- PragmaticPulp 5y agoThis applies to most specialties. Companies tend to have a few teams that lead the charge and expect everyone else to follow. Knowing which teams get the authority and which teams are along for the ride at a company is important for knowing what your job experience will look like. > However, if E-team does give you the authority to call Product's bullshit, and tell Finance to stuff it, and not take direction from Eng leads I know this was meant partially in jest, but if you reach the point where you're at odds with all of the teams and departments in the company you may get a lot done in the short term, but long term it's going to be difficult if you don't have some allies in each of those departments. Obviously no one should roll over and take orders from other departments, but some times it's necessary to do some give and take to build rapport. It's a balance, not a war.
- czep 5y agoThanks for the tips! One mantra I've tried when starting at a new job is "for the first 3 months say yes to everything, for the next 3 months say no to everything." The idea is you first immerse yourself in everything, to find out what works and what doesn't. Then you dedicate time to fix the broken processes so that hopefully when you hit 6 months your team is better positioned to be more efficient. Obviously you can't be too rigid, but it seemed to work for me when I had buy in. Curious if you think that approach sounds good.
- oliv__ 5y agoNo snark implied but what a great ad for the author! This was very fun to read, and an interesting window into the processes and inner workings of a startup that size.
- cobertos 5y agoPart of me wonders what the long term of a transition like this looks like. Would this company be able to keep its data consumption healthy, or would it drive product changes that might harm it's users or lead to dark patterns?
- tsrez 5y agoIt's such an interesting and valuable article on building a data team, esp. insightful for organisation starting out. Guess the challenges in traditional/larger companies starting out a data team might look slightly different.
- soumyadeb 5y agoSuch a great read. Have been in this position in a large public org. Over a year was spent just creating a catalog of what all data the company has and figuring out how to pull them into a data-warehouse
- plank_time 5y agoThis is probably the singly best written and most realistic article I’ve read on HN ever and I’ve been on HN for a long long time. It’s so realistic I wonder if the author took it from his diary or something. Everything about it is supersaturated with authenticity and teaches better than any other article I’ve read. Kudos to the author, and I would love to see this style of article take off.
- maileslin 5y agoErik is a legend in the modern data world. Wrote Luigi and built Spotify's first recommendation engine. He has the ground-level experience to lean on
- alexpetralia 5y agoHis post on Berkson's Paradox is excellent!
- neighbour 5y agoExcellent article. For me, the timing couldn't be better as I am about to step into a role not too dissimilar to the one described in the piece. It will be interesting to see if I run into many of the situations the author describes.
- civilized 5y agoWow, a story where things start out a mess and end up a lot better! Can we write one of these for society too?
- xpe 5y agoThere are bright spots. You might enjoy this book: Radical Equations: Civil Rights from Mississippi to the Algebra Project by Bob Moses [1] See also: The Algebra Project https://algebra.org/wp/ https://algebra.org/wp/ [1]: https://en.wikipedia.org/wiki/Bob_Moses_(activist) https://en.wikipedia.org/wiki/Bob_Moses_(activist)
- plaidfuji 5y agoSo many gems in this article… > You notice a a lot of the code starts with very complicated preprocessing steps, where data has to be fetched from many different systems. There appears to be several scripts that have to be run manually in the right order to run some of these things. > “We need to focus on delivering business value as quickly as possible”, you say, but you add that “we might get back to the machine learning stuff soon… let's see”. So so relatable. But the key insight is a really really key insight. > What I think makes most sense to push for is a centralization the reporting structure, but keeping the work management decentralized. Why? Primarily because it creates a much tighter feedback loop between data and decisions. If every question has to go through a central bottleneck, transaction costs will be high. On the other hand, you don't want to decentralize the management. Strong data people want to report into a manager who understands data, not into a business person. I have the same role at a non-software company, and to me this is nothing short of a complete reimagining of IT. It’s not just, “make sure everyone’s computer works and help them install software,” it’s, “build a model of the business, determine what information flows and metrics are crucial to success, and build an IT and analysis infrastructure around that model.” The CIO will soon be better thought of as the Chief Optimization Officer.
- jabagonuts 5y agoReally enjoyed this narrative, but what about the next phase? Going from mid-stage to mature startup? > Note that you took on a lot of “tech debt” earlier when you started dumping the production database tables straight into the data warehouse. How do you manage expectations when the year-long honeymoon is over, the business grows tremendously, and the centralized data warehouse reaches a breaking point?
- Artgor 5y agoWhen I had started reading this article, I had thought that it would be a sad story about another startup failure. The blogpost turned out to be a fascinating story of the success. I really liked it. But after I had finished reading it, I have realized that it is a sad story, if we look from the eyes of data scientists in the team. People were hired to do cool machine learning projects, but it turned out there is no infrastructure for them. After the new boss had arrived, they had to work as analysts for months. What is more sad - the new boss dangled a carrot before them several times, but each time the carrot disappeared.
- machinelearning 5y agoVery interesting perspective. As a early-mid stage startup, you definitely want to invest in generalists who are able to build infrastructure before hiring specialized ICs. I honestly had flashbacks when the author mentioned the carrot dangling thing. I’ve personally experienced this and as a naive early career swe, I gave the manager the benefit of doubt for a year even though I knew there was no way they could guarantee it. This is just pure manipulation. The worst part is that he wrote the job description himself and resorted to manipulation to cover up his mistake of hiring for the wrong job role.
- spicyramen 5y agoCan correlate, author is a truly a genius. We had a company mandate to be ML first, we went through a lot of phases and so many conversations happened as described in this amazing piece. Thanks Erik
- zippy5 5y agoThis was wonderfully written and if your gonna start a data team, this is how you do it. But I can see that I’m the only one who thought it was crazy to start a data team in the first place. This company makes 10M and spends 3M on the team and infrastructure to make data a core competency? A vast majority of wins discussed were lowly differentiated web / mobile / supply chain analytics which they could have gotten and setup with 3rd party software for an order of magnitude cheaper. I can only imagine what this hypothetical startup could have learned if they spent that money actually talking to customers, and running more experiments. I’ve heard people talk about data as the new oil but for most companies it’s a lot closer uranium. Hard to find people who can to handle / process it correctly, nontrivial security/liabilities if PII is involved, expensive to store and a generally underwhelming return on effort relative to the anticipated utility. My take away was that startups benefit tremendously from a data advisor role to get the data competency, as well as the educational and cultural benefits, but realistically the data infrastructure and analytics at that scale should have been bought not built. Obviously there are a couple of exceptions such regulatory reasons like hippa compliance for which building in-house can be the right choice if no vendor fits your use case.
- chupchap 5y ago> it’s a lot closer uranium Love this analogy!
- lifeisstillgood 5y agoAs someone who reaches for code if they need to blow their nose, what is a 3rd party vendor going to supply that a “English-to-SQL translators” wont do? (I have not finished the article, but the idea that devs / data scientists can be replaced by some vendors makes me wonder what I have missed) Edit: Also love the Uranium quote :-)
- zippy5 5y agoSo my assumption is that for a given business model, like e-commerce or Saas business much of the highest value analysis is fairly standardized and can be templated. For example breaking down conversion rate by weekly cohort is something that can be pretty easily be done in google analytics. The problem with English to sql translators or most coders in general are the assumptions we make, in particular about the underlying data. For example, say we want a join two tables, so we write a query to join on two columns and often call it correct which it is from a logical or schema perspective it is. However, null values, defaults like 0, many to one relationships vs one to one relationships, issues with instrumentation such as networking timeouts or bot detection, etc all can impact the down stream metrics. My point is that when there are 500 lines of sql in a query such as those mentioned the article, there’s a lot of ways to be mostly correct but to cumulatively be wrong. Like many popular enough open source tools, 3rd party vendors get battle tested, issues get found before you, and they can justify devoting more resources to rigorously ensure correctness than the average analyst has the time or energy todo because their business depend on you trusting the outputs. I’m not saying you couldn’t do all this yourself. But given the sheer number of analytics tools that are reasonably priced, you might have chosen to spend your time on something more specialized like a recommendation system.
- te_chris 5y agoThis is a good write-up, but for the sort of insights they’re getting they’re over staffed and overpaying. A combination of a cloud dw (big query, e.g), cloud etl (stitch, fivetran) and dbt for the T in ELT to build useful reporting tables, along with some sort of sql based BI (mode, in our case), could deliver the same insights for a fraction of the price. Throw in a sub to Heap or similar for ad-hoc product analytics as a cherry on top. I concede, of course, that they’re rescuing a bad situation, not starting from scratch, but still.
- GlennS 5y agoI liked this article, but I have two questions: 1. Is it definitely a good idea to build a separate data team, rather than embedding people with analytics knowledge in feature teams? Is it possible to do the latter, but still have end up with a well-curated source-of-truth for your data? 2. Is A/B testing and driving your business by metrics really a good idea? My (uninformed) impression is that data-driven is responsible for rather a lot of rot: - Extremely irritating websites. - Businesses ignoring important things because they can't measure them. (Financialisation, hand-in-hand with the MBA types the author decries.)
- alzaeem 5y agoI share the frustration with how many A/B testing driven development processes end up. Leads to a very iterative process with lots of small changes, rather than big bets. Also, trying to get statistical significance from iterative changes when you don’t have a ton of data is problematic.
- iamacyborg 5y agoI think that’s just down to a lot of folks who think ab testing is the answer to every problem not necessarily having a background in maths or stats. I see it all the time in marketing teams where people’s are so conditioned to think of testing as the default that they don’t understand what they’re doing or why.
- dijksterhuis 5y ago> Is it possible to do the latter, but still have end up with a well-curated source-of-truth for your data? It's important to get the core centralised data infrastructure up and running (even if it's dirty af) as that helps with the bulk of the data work. The oft quoted not completely true but kinda true statistic is that 70% of data work is finding, cleaning and storing the data. Analysis and modelling is the easy bit. You could do it the other way around. Hire some data people in each team and get them to meet up every once in a while. But I'd wager the central data stuff that makes everyone's life easier will get pushed back behind the "urgent" team work every time. #ConwaysLaw Edit: it's possible to do both btw. E.g. Have a bunch of centralised data engineers that do the heavy lifting stuff. With data scientist/analysts embedded in teams doing the fine grained modelling stuff. It's not a binary choice (once things are up and running). > My (uninformed) impression is that data-driven is responsible for rather a lot of rot. I agree! I was talking to someone else (not a tech head) the other week and realised why they hate tech so much... User interfaces that just... Don't work. Showed him a terminal cli and he went nuts over it. Then again, we're two kinda weird ye olde "back in my day" kinda people... So...
- roystonvassey 5y agoThis is a perfect encapsulation of my career as a data-guy square peg in a round hole, filled with jargon and misplaced understanding of data in general. Despite all that you read and hear about data science advancing, you’ll be surprised to see how poorly leveraged, or worse, billions of dollars are sought to implement the latest tool that promises to change the world. Tech and data as we imagine it be in the FAANG kind of companies is far different than how it is in older industries. It’s not just systems that need upgrading, company cultures do and that’s never an easy or fast process. I’ve been in the data Analytics space for 16 years now and I still feel, more often than not, I’m part of the minority, working to demonstrate true data use-cases
- AtNightWeCode 5y agoI really enjoyed reading this. Very well written. At companies I worked teams can never read data from the DW btw. My experience with A/B tests is that they are way overrated. On the poor data quality. You sit on a product like a call center. Frontend developers thinks it is an excellent idea to store all data in some doc db blob. Then business wants stats about number of calls based on users... Be careful when putting tabular data into doc dbs.
- babublacksheep 5y agoExtremely relatable content throughout. Especially around teams beating their own drums while CEO questions around metrics. ;) Will wait for a follow up post on how decentralised data team created data silos and how we solve it using data discovery and data standardisation. :P Disclaimer: I have built decentralised data teams and it scales well.