15 ms·
AWS releases Glue Databrew, a visual ETL tool
- ManWith2Plans 6y agoHaven't used this yet, but this looks like a really good user experience from their demo video. Haven't used competitors like Alteryx myself, but just having this integrate so well into the AWS ecosystem makes this seem really useful.
- manigandham 6y agoLooks very similar to GCP's Cloud Dataprep (which itself is powered by Trifacta).
- ctvo 6y agoThe thing folks don't mention regarding AWS is the inherent competitive advantage their micro-startups have. We focus on AWS launching managed ElasticSearch or managed Kafka, and talk about them (legally) using open source contributions to make money, but I think those are minor compared to things like this. What AWS has is a culture and institutional knowledge on how to launch new products that take foundational AWS services (S3, Lambda, EC2, DDB, etc.) and glues (!) them together better than what a competing non-AWS company can do. This is a bold claim (since AWS launches some very crappy products), but imagine being able to use AWS infrastructure at cost, having internal knowledge on how to best optimize that infrastructure and access to the engineers that own those services while you build abstractions and better user experiences on top of them. I don't know how cos that compete in any related space can survive. When AWS is willing to throw whatever against a wall (launching 50+ services a year) to see what sticks, sooner or later they're going to land in your space. Become more locked into AWS's foundational services -> these abstractions on top of them start to make more sense in engineering complexity / delivery time / possible cost dimensions -> Use more of these -> Become more locked into AWS's foundational services. This feels very different from Azure or GCP.
- shepardrtc 6y agoWe were using Alooma for ETL for years until Google bought it and started to deprecate AWS connections. It was a massive PITA, but it mostly worked. We switched over to AWS DMS and it was easy. Honestly it didn't take much effort. It has worked flawlessly - literally zero errors - from the day we started it up, and best of all, it's free. All you pay for is the instance it's using for you. That sort of thing can save startups much needed money. Yes, you're tied to the ecosystem - and that's what they want - but it's worth it. Once I talk to people and basically say the same thing you're saying, they start to look at AWS a bit differently.
- gk1 6y ago> sooner or later they're going to land in your space. Absolutely. I've seen this a handful of times with companies I consult, where they suddenly find themselves competing with AWS. I call it the November surprise because it happens around Re:Invent. There are several reasons this is a tough thing to compete against, and AWS's vertical integration is just one of them. I've already written about them and also how to come out ahead if you find yourself in this situation: https://www.gkogan.co/blog/big-cloud/ https://www.gkogan.co/blog/big-cloud/
- taylorwc 6y ago> I don't know how cos that compete in any related space can survive. When AWS is willing to throw whatever against a wall (launching 50+ services a year) to see what sticks, sooner or later they're going to land in your space. This is true for a subset of products, but not uniformly. To the extent you're building an infrastructure product, you get to choose what axis to compete on. If you're going up against AWS, then trying to compete with them on things like cost and reliability are likely poor choices. But something like user/dev experience isn't. DynamoDB has a mongo compatible API and yet Mongo's Atlas hosted service is responsible for most of the company's growth over the past year. Why? Because it provides a unique offering, not just a 'good-enough' offering, which is what a lot of higher-level AWS services are.
- freeone3000 6y agoWhy does this feel different from Azure, who are also expanding services and have the same advantages?
- ctvo 6y agoI don't see the same broad amount of services launched out of Azure as I do from AWS, and definitely not from GCP. I don't know if this is a strategic difference or an execution / cultural difference (AWS ships products faster, but they're almost barely usable in v1)
- VectorLock 6y ago>AWS ships products faster, but they're almost barely usable in v1 Sounds like the widely accepted "minimum viable product" approach.
- FridgeSeal 6y agoThe difference is that AWS services become usable, whereas the Azure services stay broken.
- kthejoker2 6y agoHow are you judging the "broad amount of services" launched out of Azure? They release something on the order of 10-25 updates a week, their services feed runs nonstop. I claim no special knowledge of AWS, but Azure is full apace, certainly faster than even us global SIs can keep up with in terms of providing support and capabilities.
- jjoonathan 6y ago> What AWS has is a culture and institutional knowledge... Their execution is routinely abysmal but it never matters because they have two trump cards: 1. Backdoor through the purchase process bureaucracy 2. Network effects of existing services
- soamv 6y agoAnd yet, snowflake
- kbanman 6y agoSpeaking as a former AWS Engineer, I disagree with the sentiment that they are able to glue AWS services together better than what a competing non-AWS company can do. Internally the use of AWS is subject to the same constraints and APIs you and I have. Their competitive advantage is their captive customer base, which will much rather pay a premium to use an AWS-managed service than use another vendor.
- justicezyx 6y agoThe ability to stop by the desk of S3 team member and ask whatever technical questions and get authoritative answers is enough to defeat any competitors who want to build products on top of S3. Note mentioning accessing to road-map, strategic investment, genuine appreciation of product strength and weakness etc.
- protomikron 6y agoYou underestimate how hard it is in large companies to actually speak to the people in charge that can help you concerning your problem.
- justicezyx 6y agoI did not, I am comparing that to a random guy from some random startup, who is ranking even behind the poor customers who cannot get hold any devs for their confusing issues of using AWS...
- devcpp 6y agoDon't forget being able to get high priority in the backlog if you need a feature from another service in order to launch. Former AWS engineer who launched a service here. That, access to source code and being able to setup an hour-long meeting with any engineer are the big points. Not that I think that lacking these is insurmountable, but they're very nice to have.
- threeseed 6y agoAWS biggest advantage is in the enterprise space. Which is that companies do not procure individual AWS services but rather AWS itself. Meaning that whenever AWS releases a new tool it is instantly approved and available for use across the company (baring internal processes e.g. security hardening). Compare this with a startup which has to go through a 6 month long procurement process complete with vendor bake-offs in order to sell their similar tool. If AWS continues to move into the application space they will surely dominate the enterprise because of this.
- outworlder 6y ago> This feels very different from Azure or GCP. Yes. Compared to those, newly AWS services are more likely to work with, and integrate with existing services. However, the further you stray from 'Compute' the less likely this is to be the case. More 'esoteric' services tend to be their own microcosm and sometimes feel like they could have come from another company entirely (Quicksight? etc) This is still light years ahead of Azure (and to a less extent GCP), where even compute services will not necessarily work with one another. You need to make sure the "SKU"s are compatible. Want to use some fancy storage? Oh no you need to use SKUs XYZ and premium this premium that. Whereas if AWS releases a new storage type (such as IO2), you can pretty much assume you can attach that to any of your existing instances (even if some particular types could be recommened). Not to mention surprising behavior when you try to mix and match features. GCP and AWS, you have instances working perfectly fine, but you have discovered that they provide the ability to create 'internal' load balancers? Cool! Create one, point to the instances, or point to their respective automatically managed groups (ASGs or instance groups). It will be there in case you need it, your workloads are unaffected. Do that on Azure, and now your instances have no internet connectivity whatsoever, as all traffic is now routed through it. There are footguns everywhere. Technically, GCP tends to be the most advanced of the bunch (their automatic instance migration is brilliant, meanwhile AWS keeps sending emails to us saying that some instance is degraded and it's our problem now). Their networking capabilities are impressive as well (first to have global anycast load balancers, Google's premium network, subnets spanning AZs, etc). However, they do seem to be too opinionated. Want proxy protocol on your NLBs, even though NLBs preserve source IP so in theory you don't need this(but with K8s ingress you might). AWS says sure, we have the feature, enable it, we don't care. Google says: why do you need proxy protocol, the source IP is there. These are not the headers you are looking for. Azure says: proxy protocol wat?
- jiggawatts 6y ago> This is still light years ahead of Azure (and to a less extent GCP), where even compute services will not necessarily work with one another. You can't use the SQL Server Virtual Machine extension on an Azure VM to extend the disks if the VM size is one of the AMD EPYC CPU types. During the support call, the Microsoft tech shared a screenshot of the source code for the SQL VM extension, and it had a switch statement that decides if each feature is "supported" or not. Let that sink in: Microsoft literally hard-codes their VM-size-to-feature lookups in probably thousands and thousands of places with huge switch statements full of code like this: case "Standard_M416ms_v2": return false; case "Standard_M416s_v2": return false; case "Standard_M64ls": return true; case "Standard_M64ms": return true; This is their standard coding practice. So next time you try a new VM size or type, don't be surprised if things randomly don't work or "aren't supported" for mysterious reasons...
- tuna-piano 6y agoI think their key advantage is sales. Imagine a product that adds a small amount of value to a company but requires a long drawn out sales process including research on available vendors, pricing, security, use cases, determining requirements, etc vs a developer going to the AWS console and clicking "create databrew". It's no competition. And with sales taking up such a huge percentage of a lot of these SAAS companies revenue, Amazon can pass the lack of sales to the customer as cost savings. Skip the sales process and the sales cost. win win.
- fs111 6y agoDon't buy the copy, buy the original: https://www.trifacta.com/ https://www.trifacta.com/
- wills_forward 6y agoYES. AWS' UI looks like a wholesale ripoff. Sad. Product manager: "Hey guys, see Trifacta? Go make that."
- richardowright 6y agoOr use the opensource take of it - https://cdap.io/ https://cdap.io/
- aketchum 6y agoI am a big fan of AWS and am happily running our entire tech stack with their services for a very reasonable price. That said, Glue is an absolute dumpster fire of a product. My team and I have wasted countless hours trying to wrangle a DynamoDB -> Glue -> Athena -> Quicksight pipeline and Glue refused to cooperate (we ended up building our own DDB to SQL pipeline after finally giving up on Glue). Hopefully this will increase the usability of the Glue product and actually enable out of the box ETL.
- greggyb 6y agoI'd love to hear more about this - do you have any write-up you've done on the process you and your team went through? If not, would you be willing to share where some of the big pain points were?
- aketchum 6y agoI don't have any write ups but off the top of my head I remember issues with data type casting from ddb -> athena via glue. If a ddb number type was an int in one item and a float in a different Item, glue transformed it to a struct ( something like struct(long:null,double:50.50) and struct(long:20,double:null)). The suggested fix of a cast function didn't work.
- mobjack 6y agoThat brings back bad memories from working with Glue. If the source data isn't 100% clean and compatible with the destination, it is such a pain to get it working. I eventually got it set up, but for the amount of effort involved, I could have just wrote my own custom ETL solution in less time. The scheduling jobs and triggers is nice once set up. I do hope AWS makes improvements because Glue has potential, but it doesn't feel like it is ready for prime time yet.
- MSM 6y agoI'll add to this that because AWS is simply piecing different technologies together under the hood, there are a lot of data type issues. Another example is that some date/time columns got brought in and crawled as a string. That's a bummer because obviously you want to do native operations of these, datediff, datepart, etc. without having to cast all over the place. We manually set them to timestamp and they work perfect (awesome!), even in Athena, so we thought the problem was solved. However, once we did anything with those columns in Glue ETL, those fields got set as nulls. The problem can be fixed of course, but these types of issues happen fairly often and they quietly fail (no errors, just values set to null).
- 2wrist 6y agoHave to say as slick as stuff like this looks I do find myself gravtatiing towards ETL in code. (It feels easier to read/understand) How would you change control something like this?
- rhombocombus 6y agoit looks like the jobs are exportable as JSON, which would make it far more maintainable than some of the existing proprietary ETL tools (lookin' at you IBM).
- reasonabl_human 6y agoAzure Data Factory accomplished what you’re looking for by generating and version controlling JSON pipeline definitions... you can just edit the pipeline JSONs too if you’d rather build in code than in the visual canvas
- ghc 6y agoRunning any sort of innovative data infrastructure startup (whether data prep, database, data pipeline, etc.) is now an exercise in futility. The big three cloud providers will embrace your innovation, extend your product, and extinguish your business. Given the market power of cloud providers, every infrastructure innovation now a "sustaining innovation" in Christensen's terminology. The key to success seems to be building a product for a niche the cloud providers think is too small, and then either maximizing your value within that niche so that if your market grows large enough for a cloud provider like AWS to come after you, you can pivot to providing customizations to your highest margin customers. MongoDB is a good example of this. On the other hand, none of the major cloud providers seem capable of moving up the value chain to the application level, so if I were starting a company today I would focus on leveraging my infrastructure-level innovation to create a vertical opportunity in a high margin market instead of seeking to build a horizontal platform (IoT platform for example).
- streetcat1 6y agoRight. However, you should do the opposite. Embrace their innovation and offer it on-perm (for example, on Kubernetes).
- typpo 6y agoI'm glad to see a competitor to Trifacta/Google Cloud Dataprep. My company relies on it heavily, but we constantly run into bugs, UI glitches, and crashes that can sometimes block people for hours or days. It's the sort of software that you hate to use, but the benefits are too good to ignore. The benefit to visual ETL is that non-engineers can do a lot of basic data engineering. We tie this into our more complex code-based ETL pipelines. It was a game-changer for us and helps us get a lot more done.
- deleted 6y ago[deleted]
- blakeburch 6y agoCurious to know - how do you tie the no-code transformation with your code-based transformations? Usually these processes end up siloed from each other.
- somurzakov 6y agolooks pretty basic limited copy-cat of trifacta/tableau prep/alteryx. this tool requires ready mostly clean-ish data to work with. but the #1 problem in data engineering is lack of such data
- crb002 6y agohttps://conexus.com/ https://conexus.com/ should get more love. Based off of https://www.categoricaldata.net/ https://www.categoricaldata.net/ . Ensures that complex transforms are provably correct.
- orf 6y agoGlue is an absolute horrendous mishmash mess that seemed to suffer from a serious lack of investment or vision. The managed spark component is a good product buried under an all-round terrible developer and console UX, and the data catalog/schema crawling is really useful. But I’m glad the lack of investment is turning around with this, the recently released Glue Studio and the fantastic “glue 2 fast startup” job types.
- seddonm1 6y agoThese GUI driven/Visual ETL tools certainly have their place but are firmly at one end of the ease of use vs engineering discipline based ETL continuum. As other posters commented Visual ETL often suffer from source control or limited extension ability but do provide the rapid development environment that users (generally more business oriented) seek. They also tend to trivialize the value of experience/discipline - for example I go to an accountant for my tax because they apply learned-experience relating to tax that I do not have (even though the math is easy) whereas in data engineering seemingly simple tasks such as correctly applying data typing to money or dealing with timezones seems to be glossed over in the pursuit of DIY - and wondering why your money columns don't reconcile or you lose data in failure scenarios. At the other end of the continuum large teams writing bespoke ETL code for every job does not scale well for many reasons (https://reorchestrate.com/posts/code-doesnt-scale-for-etl/ https://reorchestrate.com/posts/code-doesnt-scale-for-etl/). I think the positive reaction to ideas like Data Mesh comes from the failures of these large, centralized teams which coincided with the Hadoop era. Our solution has been to develop an open source (MIT) declarative framework (https://arc.tripl.ai/ https://arc.tripl.ai/) that allows configuration driven ETL - mostly developed via a Jupyter Notebook environment (to allow rapid development and appeal to a larger audience) - whilst making most of the difficult tasks mentioned above easier. This has been in development for a few years now and continues to evolve. We value your feedback.
- mrmonkeyman 6y ago"Clean up data up to 80% faster" 80 huh? So 85% is out the question. Seems oddly specific. Really inspiring, these truthful, helpful numbers.
- georgewfraser 6y agoVisual ETL is not as good of an idea as it seems at first. You end up putting a ton of business logic into the menus of these tools, and it’s not version controlled, and it’s not searchable. You’re better off doing this kind of work in SQL.
- breck 6y agoMy guess (without having tried it) is Databrew is backed by a solid DSL, and you are really generating good clean code when you are using this "Low Code" tool. And that works great—a great DSL + a visual GUI that edits that DSL. Because like you said you absolutely need VC.
- kevinsundar 6y agoYup there is a JSON based DSL that the UI generates. You can export, import it as well.
- VectorLock 6y agoThe $1 per 30 minute session pricing really jumped out at me.
- ineedasername 6y agoSeems like this would be a good fit to expand to cover SageMaker & AirFlow for a really powerful GUI workflow editor that includes ML directly.
- shmoogy 6y agoAre there any alternatives to this style of application? I was going to try to make something similar to this for my team to use that would be able to give them access to map columns and simple transforms, then push the resulting flow to me to move it into an airflow dag. I would really like a visual editor I can adjust with Regex functions and mappings that I can self host and iterate on.
- richardowright 6y agoTry CDAP - https://cdap.io/ https://cdap.io/ . Open source with Google backing it.
- shmoogy 6y agoThank you, this looks incredibly close to what I want! Definitely testing this tomorrow, thanks again.
- awinter-py 6y agomy brain keeps re-parsing this to 'grue datablew'
- QuinnyPig 6y agoThis service name makes me viscerally angry.
- iblaine 6y agoHaving used GUI ETL tools for years (SSIS, Informatca, Talend, Appworx) and now using Airflow(Prefect is an excellent alternative btw), I hope to never go back. Great to see Glue improving and for the industry’s sake I hope it doesn’t catch on. Most ETL should be treated as code. As code, ETLs are easier to write, maintain, and manage complexity.
- kfk 6y agoI started my team thanks to Alteryx but happily moved all to Python a year later. UIs are terrible for change management, versioning and documentation. Once you have a big-ish team doing data work UIs will become a significant drag on collaboration and productivity