26 ms·
Engineers Shouldn’t Write ETL
- corporateguy5 8y agoThis article matches my experience exactly. Some companies will hire a “data scientist” on pedigree. They will be low on skills and high on charisma. The engineers are burdened with implementing the ideas as well as shoulder the failure of the algorithms. “You spent the last few months implementing algorithms and none worked?”. Very little blame will go to the data scientist. In tons of cases data scientists are more like product managers.
- Plough_Jogger 8y ago(2016) tag.
- protomyth 8y agoReport Developers, on the other hand, are folks who have made a career around designing reports in a specific tool (e.g. Microstrategy, et al). They are specialists. Is this the common perception, because it really doesn't line up with my experience?
- kevinpet 8y agoSeems like the author is talking about the pre-big data version of business intelligence with star schemas and attempts at drag and drop tools, which has been supplanted somewhere around 2010-2015 by open source big data tools. I wasn't at a big enough company to have a proper BI department pre the data science renaming, so I can't really opine on whether it's true.
- bigger_cheese 8y agoAt least in my Org Reports are pretty much an after thought left to the data engineers (like me) to "take this metric I've developed" and display it on the morning report. Writing/updating a report is easiest part of my job it's the data that goes into building it that is hard translating the "simple metric I've developed" and getting it to run in a robust automated and sane fashion is the difficult part. The complexities in my org are two fold. Firstly the infrastructure people don't get data - at all. They speak PLC's and HMI's to them it's all OPC and magic A2A messaging takes care of everything. All data is time series to them and it all goes into an historian (which is basically a giant ring buffer i.e it gets flushed periodically) anything beyond that is past their level of expertise. The data needs to be batched together the time series information has to be processed into "event frames" - this data was all part of this sequence of conveyor belt movements for example. Then you need to link it to related events etc and archive it in some kind of sane fashion so that in six months time if there is a product defect or something like that you can trace the entire series of event frames for that particular production batch. Secondly the people the article calls "data scientists" (in my org these are Engineers - real ones of the Chem and Mech variety) don't know anything about databases or handling data they prototype their metrics in Matlab, Fortran, Excel and the like. You really need someone to translate their code into something sane that can be automated. Engineers are not taught to code at all. I know I studied engineering at university Fortran is the lingua franca. Code is just a way of representing mathematics. Asking these people to do all the data processing pipeline is just not going to happen. It's not their job. They write the simulations and models they have the domain knowledge thats whats important for them to be worrying about.
- protomyth 8y agoOk, I think I'm getting the specifics of this situation. So, we are talking about internal reports, not something that could actually get in the external customer's hands.
- bigger_cheese 8y agoYes this is internal stuff. I work at a large industrial manufacturing plant. Reports that go externally are done by certified people. (Laboratory technicians for product specifications and finance analysts for stock market stuff).
- protomyth 8y agoI’ve done external reports for clinical trials and agriculture, and I guess they weren’t as up on getting certifications. Thanks for the very detailed replies.
- iblaine 8y ago> March 16, 2016 It's 2018. A lot has changed since 2016. The line between sw engineering & data engineering is much thinner.
- meritt 8y agoI'm curious what innovative tools have emerged in the past two years that changed the dynamic?
- haney 8y agoNot sure what the author was referring to but, Airflow has gotten better / gained wider adoption during that time and my team started using DBT which saved a bunch time and was new during that period.
- closeparen 8y agoAirflow makes the distinction wider, as creating data pipelines requires even less software engineering.
- iblaine 8y agoAirflow and tools like it are probably the biggest reason for the shift. Another issue is the need to integrate different technologies requires having the skills of a software engineer. When the landscape for tech in DE was oracle, mysql and Cognos, DE's didn't need to know about OOP or consensus algorithms. Because the landscape now includes hadoop, redshift, kafka, spark, airflow, notebooks, TiDB and lord knows what else, DE's need to have most of the skills of a software engineer to be successful.
- EToS 8y agoLots of people want their key discipline to be the centre of the universe, you see it across designers, content creators, engineers, testers etc.. The key to any team in my experience is to have a healthy mixture of specialists (narrow scope, high resolution) and polyglots (wide scope, lower resolution), and to promote collaboration as much as possible..
- whack 8y agoI've worked in a hedge fund in the past, where my role sounded a lot like what the author describes as a "Data Engineer". I would have thinkers, ie people with a lot of financial experience, come up with ideas on which datasets we want to import from which vendors, and how we should handle the 80 different types of corporate actions that are contained within this dataset. I sometimes gave my own suggestions on how to improve upon their ideas, but for the most part, I was happy to focus on implementing their ideas, in the most clean, elegant, robust and testable manner possible. I was happy to do the "plumbing" work of improving upon our tech stack and architecture, in order to make the entire system better functioning and easier to maintain. According to the author, I'm supposed to resent the fact that I'm a "doer/plumber", and not a "thinker". In reality, it was the opposite. Do I really want to spend my entire day reading the Bloomberg manual and figuring out which tables/columns will give us the data we want, and the nuances of what this dataset does and does not cover? Sorry, I have zero interest in doing that. I enjoy programming. I enjoy system design. I enjoy building stuff. I have zero interest in becoming an expert on how to interpret the Bloomberg symbology file. Besides, if I ever left the financial industry and joined a tech company, that knowledge will become completely useless. Did I or anyone consider myself to be a "menial" plumber? I don't think so. I was getting paid hundreds of thousands of dollars, because the "thinkers" recognized the value that I brought to the table. They appreciated that I could quickly and robustly implement the ideas that they had, and keep the system running smoothly without hiccups. They recognized anyone can do a "good enough" job, but it's much much harder to find someone who can do a great job. And for my part, I was perfectly happy to be that guy. If you're someone who wants to expand your breadth and take on more "thinker" responsibilities, more power to you. But just don't forget that there are people like me out there too. There's no shame in being an excellent "doer".
- internet555 8y agoWhat’s funny to me is how many incompetent “thinkers” appear in meetings. Obviously, thought (even removed from implementation entirely) often has immense value. Eg, many people spent a lot of time thinking about arithmetic, linear algebra, floating point, compilers, and now I can go run whatever cool algorithm on my computer. But I continually seem to run into these people who seem borderline incompetent at anything but spewing out whatever pops into their head. Half is nonsense, one-quarter would be actively destructive if you tried to implement it, they always seem to know everything about everything but whenever it’s something you know really well you can tell that they are very confused, etc. when I meet these people now I just think “oh, you’re one of those guys who is good at saying a lot of things” and then move on. Oh well
- amarshall 8y ago> There is nothing more soul sucking than writing, maintaining, modifying, and supporting ETL to produce data that you yourself never get to use or consume. Instead, give people end-to-end ownership of the work they produce (autonomy). I think this is more the point than “engineers shouldn’t write ETL”: the engineering-related department consuming the ETL’s output should likely be the ones writing/maintaining it. Or, perhaps more generally: don’t delegate entirely to another team if the team that cares about the result is capable of meeting their own needs.
- liamconnell 8y ago> the engineering-related department consuming the ETL’s output should likely be the ones writing/maintaining it. This is exactly the author's point. The Data Scientists are consuming the ETL's output, so they should learn how to write and maintain ETL since it isnt very hard or time consuming with modern tools.
- WorldMaker 8y agoThough this backfires sometimes for the engineering department in that then the engineers get forced into an "inner-platform effect" problem that they instead have to build an ETL platform abstracting enough ETL abilities for their company's data scientists' skill level, yet generic enough for their company's data scientists' arbitrary questions/needs. That is its own soul-sucking experience. "Can't we just hire people that can learn Power BI better? Why are we still writing data tools for people that think they know Access but barely know Excel?"
- humbleMouse 8y agoThis quote is ridiculous. Some people enjoy plumbing high speed reliable/transparent data pipes. I could care less what goes thru the pipes I make.
- m_ke 8y agoThe "T" in ETL is important here. You might not care but the end users of that data should. It's very easy to take raw data and strip it of a lot of useful information by transforming and normalizing it. The person doing modeling or data analysis should ideally be dealing with the raw data, know how it was collected and understand what each field really means.
- msencenb 8y agoDoes anyone have experience with ETL as a service like StitchData (not related to stitchfix)? The startup I'm employed at needs some data analysis, but it is not big data, simply a way to unify analytics into a queryable database. I'm not looking forward to writing any ETL code, and was hoping someone here had a tool to help.
- ben_jones 8y agoNo tool but I've had success writing code generators to better interface between various data systems, autogeneratoring various accessors and utility functions.
- fraserharris 8y agoI would suggest checking out my employer, Fivetran (fivetran.com). Many tech firms use us to centralize their data for analysis. We have startup pricing for sub-50 employees.
- NovaX 8y agoDo you know when DB/2 support will be added (Windows, AS/400)? I've been wanting to switch over from AWS DMS and our in-house support.
- fraserharris 8y agoWe are planning on adding support for DB/2 in Q1/Q2 2019. We are getting SAP HANA out first
- jakestein 8y agoWe (Stitch) have DB2 in production today https://www.stitchdata.com/integrations/db2/ https://www.stitchdata.com/integrations/db2/
- NovaX 8y agoIs it log-based cdc? The documentation says only 2 dbs are supported that way. We want to replicate customer databases which have schemas outside of our control, with the goal of migrating them to our products. The other vendor schemas might not have your required columns. There are also many flavors of DB2. So very interested but the details are sparse at best.
- bacon_waffle 8y agoETL means "Extract, Transform, Load" https://en.wikipedia.org/wiki/Extract,_transform,_load https://en.wikipedia.org/wiki/Extract,_transform,_load
- kartan 8y agoThank you. I think that is good practice to introduce abbreviations correctly. Even that it is easy to forget when you work with them all the time. "How do I introduce an abbreviation in the text? The first time you use an abbreviation in the text, present both the spelled-out version and the short form." https://blog.apastyle.org/apastyle/abbreviations/ https://blog.apastyle.org/apastyle/abbreviations/
- dmh2000 8y agoYes, its poor writing/editing of articles that don't define their acronyms. or at least the important ones in the title.
- deleted 8y ago[deleted]
- unholyguy001 8y agoGuess that didn’t last cause look they are hiring data engineers https://www.stitchfix.com/careers?gh_jid=1252958&gh_jid=1252958 https://www.stitchfix.com/careers?gh_jid=1252958&gh_jid=1252...
- deleted 8y ago[deleted]
- finnley 8y agoThere's a huge difference between writing ETL to apply business logic vs. grab the data from a common API like Google Analytics. There's no tool in the world that can write all the logic you need to transform data the way your organization uses it unless you have an extremely simple, common use case. What this article is really saying is that replicating your data from source apps shouldn't be manually coded. The harder part still needs someone to write code so business users don't need to.
- kevincennis 8y ago> Autonomy means the data scientists own that code as well. All the way into production. This does not strike me as a great idea.
- EdwardDiego 8y agoYep, data science and software engineering are two very different disciplines.
- deleted 8y ago[deleted]
- evrydayhustling 8y agoThis is a great read, and this is a critical sentence: > We are not optimizing the organization for efficiency, we are optimizing for autonomy. Efficiency is for production pipelines where the product is thoroughly defined and production costs eat deeply into profit margin. Most software organizations have massive margins - but only if they get to the right product. Organizing people for ownership and autonomy engages their creativity, but also ensures that the org can move forward even when one side or the other falls behind.
- closeparen 8y agoThe author completely lost me. Analysts produce reports. Data scientists produce models. We don’t ask a data scientist to produce a model unless we have a serious intention to put it in production. There are significant engineering challenges in taking a model from the data scientist’s batch-mode workbooks and Hadoop queries to a reliable near-real-time online service, and the relationship can get dysfunctional, but it has nothing to do with data scientists being BI in disguise.
- polm23 8y agoPreviously. https://news.ycombinator.com/item?id=11312243 https://news.ycombinator.com/item?id=11312243
- jimbokun 8y agoI think this diagnoses the problem well, but ignores an obvious solution. A team of one data scientist and one engineer, completely responsible for building a model, and seeing it through into production, meeting all applicable SLAs and performance metrics. Or maybe it's two data scientists and one engineer, or one scientist and two engineers, whatever is required. The point is to have a small team you can hold completely accountable for their output. They sink or swim together, so there is no debating whether the scientists or engineers get the credit or take the blame. They are assessed by the effectiveness of the end product they produce.
- disposedtrolley 8y agoSmall teams are awesome in so many scenarios. I recently wrapped up a 4-week proof of concept for a client on knowledge management and discovery using NLP. I was able to work with someone apt at machine learning while I focused on building out the UI and backend. We delivered a first release about 3 days after we started, giving ample time to seek feedback and let the users shape the direction.
- tkyjonathan 8y agoIn 4 weeks, I was able to create a data mart that had self-healing (we had issues with Python/events missing data which should have reached the existing data warehouse) and the physical data models in it sped up an existing 6 hour ETL task down to 0.15 seconds AND speed up a production query that took 5 seconds per click down to 0.07 seconds. No team needed or proof of concepts. Actual working data models, up to date tables + ETL code in production. Background is DBA.
- technofiend 8y agoIt's not in my experience performant but Pentaho is definitely ETL-for-dummies easy to use. Similar to your average user pivoting data in Excel rather than learning Python or R, sometimes having a tool with suboptimal performance is better than optimizing an adhoc or short term process.
- raihansaputra 8y agoYES. I just knew about Pentaho recently and it’s amazing. Sad that they just scrubbed off info about the free community version on their webpage, and to automate the jobs on the community version you need to do cron/Task Scheduler stuff outside of the app. I know it’s a limitation to make people jump ship to the paid version, but I just hope it’s integrated so I don’t have to think about setting up cron jobs to have automated ETLs and just have people responsible to create the jobs do the scheduling too.
- dgudkov 8y agoIf you're looking for a real ETL-for-dummies, take a look at my EasyMorph (https://easymorph.com https://easymorph.com). We've made a number of simplifications that specifically target "dummy" users, e.g. columns may mix values of different types (text, numbers, etc.).
- v4n4d1s 8y agoThanks for developing easymorph! Free version helped me through my bachelors degree. It's my go-to tool to introduce people to ETL and similar concepts.
- dgudkov 8y agoYou're welcome! Great to hear it happened to be of help :)
- technofiend 8y agoAnd it's a partner with Tableau? Yeah that sounds pretty good! Thanks for the tip.
- marcell 8y agoArticle from 2016
- datademon 8y agoAs an undergraduate who is about to graduate with a degree in "Data Science" this post encapsulates a lot of my worries as I move into the work world. Should I focus on being a "thinker" a "doer" or a "plumber"? For the first three years I was planning on being a CS major until I was denied from the department: now the data science major is my only hope to graduate. I feel as though my programming skills are solid: but not good enough to be on any sort of fast paced infrastructure/devops team. On the flipside: I feel as though I am so far behind on stats/math knowledge that it's pointless to try and become a data scientist/analyst. I've thought about data engineering (the 'doer') as a happy compromise between the two. However there are barely any intern or entry level data engineering positions that I can find. The ones I do find require knowledge of so many frameworks that I don't know where to start. Additionally, I'm not even sure if data engineering even is a happy compromise, especially after reading the post. Time is ticking, and sooner or later I'm going to have to figure out what route to take, and how I want to specialize. I go to a hyper competitive university in a hyper competitive region of the country and I'm starting to feel like I'm falling behind and getting lost. If any of you older/more experienced engineers and scientist have advice or wisdom for me, I would very much appreciate it.
- nitrogen 8y agoA bit OT, but as a more experienced engineer who dropped out of school to start a company, I'm curious: why weren't you able to get into your school's CS program? Don't worry too much about "falling behind". There will always be time to learn more math or a new framework. Worry more about finding that first job, any job, then you can branch out once inside the industry. Networking beats recruiters beats sending a resume, so try to find a friend who already works where you want to be.
- datademon 8y agoI did poorly on a math class that was required to declare the major. It's ironic since now that I'm in the data science major, I have to do even more math classes and less programming classes. I would love to do my own startup. I have a few ideas floating around. But I feel like I lack the discipline to sit down every day and force myself to work on them without external deadlines/pressure. In terms of jumping into the tech industry: I understand the advice about looking for any job when starting out. It just seems that even a lot of the entry level jobs are very specialized.
- solatic 8y agoCan't this whole thing be boiled down to "DevOps for Data Science/Engineering"? Different parts of the org with different skillsets and cultures practicing empathy for each other by communicating interests in version-controlled code, allowing for guard-railed autonomy, which leads to business agility. Yep. Sounds about right. > Optimize for autonomy not efficiency Optimizing for efficiency without considering the cost of work in progress (WIP) (irrelevant ETL models), rework (unscalable models), or unplanned work (unscalable models that make it to production) results in company silos (data engineering, infrastructure engineering) cheering local maxima while covering their ass in the face of a business that's suffering from a long lead time. Two teams with two backlogs will accomplish work exponentially faster compared to three teams with three backlogs. It boggles my mind how books like The Phoenix Project are not required reading.
- Annatar 8y agoI see this every day in my job. He so nailed the problems. Data scientists must be made responsible and accountable end-to-end for their solutions. And they must be grilled on operational deployability and maintanability before, during and after deployment. They have to become accountable.
- Roritharr 8y agoMy father in law is such a "thinker". He has been since the 70s and worked on all kinds of projects from IBM Mainframes for up to Hadoop and Kafka for Insurance Companies and Telkos. It's ridiculous to me how hard it is for him to find a new job at 60. He financially doesn't have to, but he wants to train younger guys on how to deal with all the weirdness one encounters in ETL Jobs.
- dagw 8y agoI often really enjoy it when I get a chance to do ETL work. The 'T' in ETL can many times involve some pretty fun and creative challenges. And even in the general case there is something really satisfying about putting together a clever and well constructed ETL pipeline.
- craig_asp 8y agoI've worked in BI (end-to-end - data modelling, reporting, ETL, etc.) for more than 10 years now across various organisations and since "data science" became all the rage, I had the pleasure to work with a few data scientists. From what I've seen so far, they are very good as statisticians (some of them university lecturers) but when it comes to building ETL pipelines, I don't think any of them could actually do it properly. Properly as in an ETL process which connects to various data sources, writes to logs, is repeatable, restartable and so on. It is not easy to get to know how to build a proper ETL process and it is not easy to learn how to "do data science" correctly as well. I see it as more productive (from my personal experience) to let the "data engineers" do the "data engineering" work - build data models, ETLs, etc. and let the "data scientists" do the "data science" work - build and fiddle with statistical models. Just like with a "full stack" developer, and the separation of work between "back end" and "front end" developers, it might be better to let each do what they do best unless you have people who can do both properly (but often it's hard to find them and they would actually be better in one area or the other). The frustration between the two camps - data "engineers" and "scientists" is usually due to mismanagement (distinct teams doing each bit separately, coordinated by one to many management layers) rather than suboptimal division and allocation of labour. Small teams of two to four people which contain the correct mix of experts would benefit from the strengths of both data professional types, and would avoid the problems around syncing the effort.
- yannis7 8y agoarrogant and pompous uses of terms "mediocre, soul-sucking, etc etc", while the fundamental ideas of the article range from trivial to minimal-value
- didibus 8y ago> We strive to lead the business with our output rather than to inform it I think the business hires data scientist to be informed. Not to make business decisions on their behalf. > Data scientists love working on problems that are vertically aligned with the business and make a big impact on the success of projects/organization through their efforts. They set out to optimize a certain thing or process or create something from scratch. These are point-oriented problems and their solutions tend to be as well. They usually involve a heavy mix of business logic, reimagining of how things are done, and a healthy dose of creativity Again, I'm confused? That sounds like the data scientists should have majored in business then. If data scientists start doing that, what will all the other business folk do then? Data scientists should just build out reports that provide valuable insights and potential patterns that can help make business decisions. The difference with prior reports engineer or data analysts or wtv, is that a data scientist is assumed to be able to generate statistical analysis or/and pattern analysis over the data. While prior, a data analyst only needed to perform basic versions of that which did not go beyond what SQL could do. The data engineer should enable the data scientist to perform this analysis by both working with the software engineers to acquire it safely, securely, reliably and at scale. And working witj the data scientist in order to apply his statistical analysis efficiently and at scale to a possibly very large data set. Finally, he might need to work with both software engineer and data scientist to setup real time or close to real time versions of the analysis. All result from the analysis should be presented (aka reported) to the business. The data scientist can suggest interpretations or ideas to address findings, but it's the business role to make tactical and strategic decisions about business processes and products. And if you're doing ML as part of a process, then you need a ML scientists. Say you need to build out voice recognition, or the likes. Basically comp sci or math majors with ML masters or PHDs.
- tkyjonathan 8y agoETLs, physical data modelling and data marts/warehouses used to be handled within database admin's task in small to medium sized companies and largely with ETL tools or just SQL.
- just_myles 8y agoYup. That's been my experience. The DBA used to handle all these tasks and as of I don't know a 5 years ago, it's been segmented into data engineering. I think in this case it's a good thing. I always considered that a non-administrative tasks.
- ghc 8y agoIt’s a bad situation in your typical enterprise, but it’s even worse where I’ve spent my career: working with realtime industrial data. I became convinced that building time series data pipipelines was a bad idea after many late nights in the office fixing fragile systems that couldn’t handle real-world complexity. As fun as it is to build with and learn new technologies, it’s a bad idea to build data pipelines unless you have a lot of resources and good leadership that can make peace between all the different people who touch the data. Unfortunately in the world of sensors and equipment there aren’t many solutions, so I started a company (at https://sentenai.com https://sentenai.com ) to save others from my years of struggle. It turns out it’s even harder to build a general time series data pipeline solution, but we’re making progress.
- jgalt212 8y agoThe counterpoint to this, and the economic value therein is PG's Schlep Blindness piece. http://www.paulgraham.com/schlep.html http://www.paulgraham.com/schlep.html
- dwaltrip 8y ago> Enable Everyone to be Best in the World This particular line really rubs me the wrong way. Do the best you can... You won't be the best in the world, but you can still have a positive impact.
- motymichaely 8y agoDifferent parts of engineering require different skill set.. Someone has to do the data engineering part (be it the data scientist, data engineer, ops, whatever..). This requirement hasn't changed since 2016: 50 to 90 percent of time is spent "Cleaning" Data for Analytics. You just need engineers with the right skills and tools to help reducing this time and get things done. https://s3.amazonaws.com/xplenty-assets/infographics/raw_data_cleaning_is_killing_bi.pdf https://s3.amazonaws.com/xplenty-assets/infographics/raw_dat...
- ArchTypical 8y agoI wrote plenty of ETL. Maintained high throughput using whatever I could find. Then I got another job and had to write ETL for AdTech, where the volume is unlimited. Nothing about it is surprising or hard. Engineers are great at handling known data and transforms, then adapting to unknown data.