8 ms·
Hey guys, I posted this maybe a year ago. It was originally an electron based desktop app which was cool but pretty hard to maintain / get people to download a
by gthompson1 4y ago
Hey guys,
I posted this maybe a year ago. It was originally an electron based desktop app which was cool but pretty hard to maintain / get people to download and test it out. Given that there would likely always need to be a cloud component regardless I decided that it would be easier to port it over into a web app and maintain it that way.
Since doing that I took a job and am pretty burned out on the project. I was thinking about open sourcing it and seeing if it gets more interest that way. Potentially as a self contained docker image that folks could pull down and run.
Glue is pretty similar to these commercial projects https://www.alteryx.com/ https://www.alteryx.com/, https://www.dataiku.com/ https://www.dataiku.com/, https://parabola.io/ https://parabola.io/ and other DAG based data pipeline generators. I also saw this posted on hackernews within the last year or so https://datablocks.pro/ https://datablocks.pro/ which is pretty cool too (also more of of an indie project).
Here is the current API. Probably needs some de-scoping so that its only the important endpoints but its a look under the hood if anyone is curious. - https://app.gluedata.io/docs#/ https://app.gluedata.io/docs#/
Anyway if anyone has any comments, suggestions / feedback please let me know. Happy to share any details or thoughts.
- 0xbadcafebee 4y agoIf you think you were burned out before, becoming an open source maintainer probably won't help :) If I were you I would put out a call for maintainers, and maybe consider a license that keeps modifications open source. That way a company that modifies it will have to contribute back to the project. If you do open source, I recommend moving your roadmap to a GitHub Project, put together a contributors guide/agreement. Companies will definitely be averse to using a project due to trademark or licensing issues. Apparently somebody is using the name (https://alter.com/trademarks/glue-85698439 https://alter.com/trademarks/glue-85698439). You could probably find a way to signify your use is distinctive from that other one, and you may want to register it to protect it. IANAL. I think it has a lot of potential, but finding people who want to build it out/run the project will be the challenge.
- gthompson1 4y ago"If you think you were burned out before, becoming an open source maintainer probably won't help" - Ha yeh I believe that. "That way a company that modifies it will have to contribute back to the project. If you do open source, I recommend moving your roadmap to a GitHub Project, put together a contributors guide/agreement." - Thanks for the advice
- austinjp 4y agoJust wanted to say: nice job! :) I can fully understand you burning out on something like this. How about declaring a time-out period where development is paused while you take a break? Maybe triage bug reports and feature requests in the meantime, but take some time to yourself to recover and reset your perspectives. Just a thought.
- gthompson1 4y agoThanks! Thats kinds of you to say. Appreciate that.
- NoImmatureAdHom 4y agoJust a note: the first thing I did was check to see if the project was open-source, and that was before reading your comment indicating you were considering publishing the source. I did that because I'm not willing to get locked in to another walled garden. That does not, of course, imply that publishing it with a FOSS license is the right move for you! I just thought it might be helpful to report my behavior.
- gthompson1 4y agoYeh I get it. I appreciate the comment. Part of the reason I am considering open sourcing it would be to give the project more life than I can currently give it myself.
- echiuran 4y agoVery cool. A GUI pandas. As a seasoned user of pandas, the way it works seems clear. I get the potential appeal of a no-code solution like this: "Glue is to pandas as Airtable is to relational databases." The question in my mind: how many users are there that want to build these kinds of data transformation pipelines that are also unwilling to learn Python/pandas (or R, or some other equivalent)? Because once you need to build a custom lambda, or do something downstream of the transformation, like a custom data visualization, you need to know how to write code. Which I think gets to the heart of the matter: this replaces the easy part of data science. The hard parts are things like understanding what the columns really are, cleaning up messy data, and doing statistical modeling that yields valid insights. I think open-sourcing it is a good idea. I can see how this could be added to a more complete software ecosystem that is designed to be used by non-coders, like an in-house ELN for a biotech. But as a paid project, I'm almost surely not going to ever give it a try.