8 ms·
I am a big fan of AWS and am happily running our entire tech stack with their services for a very reasonable price. That said, Glue is an absolute dumpster fir
by aketchum 6y ago
I am a big fan of AWS and am happily running our entire tech stack with their services for a very reasonable price. That said, Glue is an absolute dumpster fire of a product. My team and I have wasted countless hours trying to wrangle a DynamoDB -> Glue -> Athena -> Quicksight pipeline and Glue refused to cooperate (we ended up building our own DDB to SQL pipeline after finally giving up on Glue). Hopefully this will increase the usability of the Glue product and actually enable out of the box ETL.
- greggyb 6y agoI'd love to hear more about this - do you have any write-up you've done on the process you and your team went through? If not, would you be willing to share where some of the big pain points were?
- aketchum 6y agoI don't have any write ups but off the top of my head I remember issues with data type casting from ddb -> athena via glue. If a ddb number type was an int in one item and a float in a different Item, glue transformed it to a struct ( something like struct(long:null,double:50.50) and struct(long:20,double:null)). The suggested fix of a cast function didn't work.
- mobjack 6y agoThat brings back bad memories from working with Glue. If the source data isn't 100% clean and compatible with the destination, it is such a pain to get it working. I eventually got it set up, but for the amount of effort involved, I could have just wrote my own custom ETL solution in less time. The scheduling jobs and triggers is nice once set up. I do hope AWS makes improvements because Glue has potential, but it doesn't feel like it is ready for prime time yet.
- MSM 6y agoI'll add to this that because AWS is simply piecing different technologies together under the hood, there are a lot of data type issues. Another example is that some date/time columns got brought in and crawled as a string. That's a bummer because obviously you want to do native operations of these, datediff, datepart, etc. without having to cast all over the place. We manually set them to timestamp and they work perfect (awesome!), even in Athena, so we thought the problem was solved. However, once we did anything with those columns in Glue ETL, those fields got set as nulls. The problem can be fixed of course, but these types of issues happen fairly often and they quietly fail (no errors, just values set to null).
- dumbfounder 6y agoWe also had issues with Glue/Athena, never got parquet to work right, have schema issues with schemaless data coming from MongoDB, and in general it is obtuse and hard to work with. We did a data lab with AWS using Glue and it was an exercise in pain. Then we did a POC with Snowflake and we cried with joy at how easy it was to work with large amounts of data. But our ELT was very light, not true transformation of data, more rearranging. I still think we will need Glue for some workloads and I pray this makes our lives easier.
- chrisjc 6y agoSounds very similar to what we're dealing with, although never turned to Glue to try and resolve our challenge. Ended up going with an ETL as a service (Alooma and now transitioning to ETLWorks) to extract and load our data into Snowflake.
- jjoonathan 6y agoI haven't used it myself but my coworkers who gave it a spin unilaterally agree that Glue is a dumpster fire, even though they're otherwise huge AWS fans.
- RobinL 6y agoWe use Glue extensively, but we have a rule of thumb not to use any of the 'special sauce'. That means is using it purely for 'Spark as a service', so we're pretty much always reading/writing data from/to S3 using spark using a script that would work on any Spark cluster (i.e. not using the GlueContext type stuff, i don't even know what it does tbh). For this purpose I think it's fantastic. Write a PySpark script and press go (we have a package on pypi called etl_manager to facilitate this). It 'just works' for this use case, and there's a huge amount of value for us in not having to think at all about managing or configuring a Spark cluster. Our biggest bugbear was slow job startup times and a lack of pip installs, but both of those are fixed with glue 2.0 which was released recently. We don't use any of the visual/GUI based tools for our jobs, we just write our own Spark code and version control in Github. That's unlikely to change any time soon with products like Databrew. That said, the data profiling tool in Databrew does look like it could be useful as something to refer to when writing code. (I realise this doesn't help with your specific issue, but i thought it was helpful to offer an example of a good experience)
- orf 6y agoGlueContext mostly manages bookmarks as far as I can tell, which are an insanely useful feature for us.
- RobinL 6y agoInteresting - I was vaguely aware of the existence of bookmarks. I'd be interested to know about what you're using them for - they definitely _sound_ useful. I guess it probably depends what sort of workloads you're doing. At the moment we use Airflow to manage DAGs/retries etc. I like it as a user, but from what I understand from our ops people it's a pain to manage.
- orf 6y agoThe use case is pretty simple. You’ve got a bucket that you want to load data from and shuffle it away somewhere else (redshift, s3, whatever). This could be populated by a Firehose, another system, etc etc. Bookmarks just store the greatest “created time” for the files you’re loading from s3. So when you trigger a job it will only load files created since the last successful run. It does some funky stuff to handle s3’s eventual consistency with LIST operations. Super simple incremental loading. This also works when loading data from a relational database, by storing the greatest primary key value.
- drchopchop 6y agoAgreed, Glue turned out to be very underwhelming. We ended up moving ETL flows to self-hosted Prefect (https://www.prefect.io/ https://www.prefect.io/) instead. In general, AWS excels at a lot of core features (EC2, S3, the databases) but the higher-level services feel very thrown-together, with awful documentation. They are launching a large amount of half-baked services these days, (intending to capitalize on vendor lock-in?), but it's making the ecosystem start to look like a confusing mess. A lot of these wouldn't survive if they were marketed as standalone products.
- tobilg 6y agoYou might be interested in the new DynamoDB to S3 export feature announced this week: https://aws.amazon.com/blogs/aws/new-export-amazon-dynamodb-table-data-to-data-lake-amazon-s3/ https://aws.amazon.com/blogs/aws/new-export-amazon-dynamodb-...