5 ms·
First, this is a nice resource, so good job! As someone who has worked for the past several years in this space, I'd say the biggest problems in data engineeri
by DataDaoDe 6y ago
First, this is a nice resource, so good job!
As someone who has worked for the past several years in this space, I'd say the biggest problems in data engineering are wholistic in nature. Sure, you need to know Python, SQL, Data Warehouses, Data Modeling, etc., but to me by far the biggest problems have to do with the entire architecture i.e. How do you extract data from potentially unreliable data sources, pull that data into some staging area, build further workflows that base off of this raw data to reliably update or create data warehouses/marts or deploy ml models. How do you allow everyone in your company to access and work with the data in a compliant and secure way? How do you test any of this? How can distributed teams, sometimes technical, sometimes more business oriented interact with the architecture and add/control data and release it into the overall company data stream? Has anyone found a reliable and maintainable way to setup CI/CD for company data architecture/pipelines/projects?
To me these are the big problems. And if anyone has any resources for any of these topics I would be super interested, since I deal with these problems daily :)
- asicsp 6y agoSee if the table of contents of these books address some of your requirements: * https://nostarch.com/seriouspython https://nostarch.com/seriouspython * https://www.manning.com/books/practices-of-the-python-pro https://www.manning.com/books/practices-of-the-python-pro
- fractionalhare 6y agoThose resources won't really help OP. What they're talking about is better handled by bespoke ETL architecture alongside workflow orchestration tooling (like Airflow or Prefect) to handle versioning and deployment of modeling and ingestion services in production. The orchestration part handles the workflows that comprise your ingestion and ETL processes. These are like managed cron jobs specific to data engineering lifecycles. The bespoke part of the architecture is what you'd compose together to handle all of the other requirements; for example, what applications do you build, and how do you design your data warehouse, such that the architecture can be used by both data science and marketing teams?
- snird 6y agoThank you for the kind feedback! I absolutely agree. The hard problems are either organisational (how to communicate to analyst and have agreed work method with business?) as well as dealing with third party unreliable resources. I feel like these things you can learn only through experience. No written resource can reliably transfer this knowledge. I think data engineering especially is something that requires at least apprenticeship to get into. Both for juniors and for senior developers transferring to data engineering position.
- elevenoh 6y ago>but to me by far the biggest problems have to do with the entire architecture data engineers are dependent on software engineers. and the software engineering is the more difficult part IME.
- jskdvsksnb 6y agoYou seem to badly misunderstand what "data engineer" means. IME the role is basically infrastructure up to custom ETL jobs. It is a specialized strain of software engineering.
- crorella 6y agoNot only that, - How do you make sure the users are not writing badly optimized tables/pipelines that end up consuming too many resources. - How do you facilitate data discoverability so they don't end up creating a new table where 90%+ of the data is already present in another already existing one? - How do you make sure they are mindful with the way the model the tables such there are not too many files due to bad partition/bucketing, compression is leveraged and good datatypes are picked?
- atomicnumber3 6y agoI work in this space and the questions outlined in you and your parent's comment are absolutely spot-on. The worst part is, these things tend to grow organically where the original ancestor of everything is engineer #3 of the 5-person startup who decided to write a cronjob to dump the prod db every night to feather files on a samba share. Then one cron job becomes three. Then ten. Then you're using S3. Then there's dependencies. Then you're using Luigi/Airflow. Then you're using Spark. Then you're constantly messing with partitioning and performance. Then you're using Hadoop and YARN and configuring queues and capacity scaling. At this point the "infra" team is 10 people and there's 20 data scientists. Then you triple in size a couple dozen times and now it's time to figure out how to retroactively slap security controls on top of everything. That turned into a bit of a rant. Honestly I love this space but I agree, as hard as the software bits are, the hard part is the entire picture as a whole.
- somurzakov 6y agothese are largely solved by data lake vendors: Snowflake, Databricks and alike. an experienced databricks/snowflake architect (or a couple of them) can easily set up and maintain data lake that supports most of what was mentioned. Overall I absolutely agree, that rather than learning python or SQL, it is much better use of one's time to learn/get certified as Data Lake Architect and be able to create a large data lake from scratch and set up pipelines and maintain them.