3 ms·
I'm interested in hearing opinions about this principle in the context of data engineering: > This also means leaning heavily into all the service offerings an
by jackgill 6y ago
I'm interested in hearing opinions about this principle in the context of data engineering:
> This also means leaning heavily into all the service offerings and orchestration tooling that is afforded to you by your platform.
I've built a data lake and several ETL pipelines using AWS native services (Kinesis, Lambda, Athena). It works but it's a bit...fiddly. I spend a lot of time configuring these services and handling various failure modes. I've been wondering if I should be looking at third party vendors like Fivetran or Matillion for ETL.
Does anyone who's worked with AWS data engineering services have thoughts on the trade-off between AWS native services and third party vendors in this area?
- ElFitz 6y agoRegarding the configuration and failure modes, I think something like CDK could be a great way to set all this up in a more familiar and readable way https://aws.amazon.com/cdk/ https://aws.amazon.com/cdk/
- ramraj07 6y agoI can strongly attest to snowflake. I regret AWS doesn't offer the same features without making us jump through the maze of services with which we can emulate the same concept.
- jackgill 6y agoThanks for sharing, I've heard many good things about Snowflake. In the past I've seen them as more of a Redshift competitor (data warehouse, as opposed to a data lake) but if they can simplify data ingest then I am definitely interested.
- ramraj07 6y agoThey're fundamentally different only in the model of decoupling storage and compute completely, but in a far more simpler way than redshift spectrum I feel like. Some of their features like zero-copy-clone are just not possible in AWS and make it extremely simple to do pipeline management in way that (at least to me) makes the most sense. It's also the most democratizable model I have seen - anyone who knows the slightest amount of SQL can be set up to explore the data in minutes. The elephant in the room is that you need to use SQL. Their spark connecters are as of now useless, so you either have to go with DBT, some homebrew SQL stringing mess or something like sqlalchemy. We're currently developing some wrappers around sqlalchemy to make this a bit less painful, but it's still so worth it.