12 ms·
GitLab is working on a tool just for data teams
- ageofwant 8y agohttps://quiltdata.com/ https://quiltdata.com/ ticks a lot of boxes in this space for me.
- sytse 8y agoIs that project more for versioning data like https://docs.dotmesh.com/tutorials/subdots/ https://docs.dotmesh.com/tutorials/subdots/ or http://www.pachyderm.io/ http://www.pachyderm.io/ ?
- veritas3241 8y agoThis could potentially become part of the Meltano stack. At GitLab, we're not at the phase yet where we're in need of data versioning. But I could imagine a data registry that's integrated with the workflow of data analysts/scientists to easily link versions of code and data. Thanks for the link - we'll definitely keep an eye on it.
- hn_throwaway_99 8y agoBe interested to know all the competitors in this space. https://data.world/ https://data.world/ is one I am most familiar with.
- jakecodes 8y agoOne major difference will be the complete data life cycle vs providing just one part of it. Just like we do in GitLab except for data teams instead of software development teams.
- sytse 8y agoSome of the alternatives are listed on in this table in the readme: https://gitlab.com/meltano/meltano/blob/master/README.md#data-science-lifecycle https://gitlab.com/meltano/meltano/blob/master/README.md#dat...
- slap_shot 8y agoThis projects competes with too many industries to really give a succinct answer, but here's just Extraction/Loading and Analyze: Extraction/Loading Dell Boomi SAP SAS Pentaho Domo Oracle IBM Microsoft Informatica Talend JitterBit SnapLogic Mulesoft SyncSort Information Builders Actian Attunity Datameer Alteryx Striim Treasure Data Cask StreamSets Snowplow DataTorrent Astronomer Panoply Apache Nifi Stitch Data FlyData Bedrock Data Alooma ETLeap Fivetran Xplenty MethodMill Celigo TerraSky DBSync Youredi Scribe Civis Analytics DataScience Dataloader.io datorama Astera Analyze Microsostrategy GoodData Sisense Looker Power BI Wagon Birst Tableau Qlik Domo Hue Mode Chartio Periscope Pentaho The amount of hype and BS in the Notebook space would require me to spend some time combing through that again.
- chasewright 8y agoslap_shot, I agree and as I disclaimer I also work at GitLab. There is no shortage of data tools in the space today. A majority of my career has been spent in the data & analytics space and I've talked / worked with at least 60% of the companies you mentioned. At the end of the day, these are the questions I've asked over and over again. 1. Do we have enough money / budget for a tool like this? 2. Can we derive enough insights from this product fast enough to make a good ROI? 3. Does this tool use a proprietary language that no one wants to learn or can I code in a language that is relevant? 4. In all honesty, can I get insights faster in a spreadsheet than these tools? 5. What is the learning curve? 6. Can I answer the business question that was originally asked? Open to more discussions around the topic as it is a lot harder to answer than a few philosophical questions, but it certainly resonates with many data & analytics professionals. A nice goal would be to have project where you can stand up a business, turn your data pipelines on, ingest the data, and view the insights needed to make a business decision all within a short timeframe of when a business goes live.
- veritas3241 8y agoTaylor from GitLab here! Happy to answer any questions about what we're doing.
- thebiglebrewski 8y agoKudos to you for trying something new!
- veritas3241 8y agoThanks so much!
- slap_shot 8y agoThis looks like an amalgamation of 8+ open source projects or industries with products put forth by companies that have dozens of employees and worked on their products for years. It also doesn't even categorize the products they compete with correctly[0]. Why not contribute some of your resources to one of the many active open source libraries already trying to solve some of these problems, and focus your engineering efforts on your core product? [0] Fivetran is only considered "Orchestrate" but is actually competes directly with Alooma in the Extract and Load. Also, there are DOZENS of company in that space. https://gitlab.com/meltano/meltano/blob/master/README.md#data-science-lifecycle https://gitlab.com/meltano/meltano/blob/master/README.md#dat...
- sytse 8y agoWhat we're doing different is making one product that does the whole lifecycle instead of having to string tools together. It took us many months to string our toolset together and we felt there had to be a better way. Just like GitLab we try to leverage existing open source projects wherever possible. I agree Fivetran also belongs in extract and load and updated it https://gitlab.com/meltano/meltano/commit/1df9813f5ab42c4479120f4d7e9f6f8e8a06a1ae https://gitlab.com/meltano/meltano/commit/1df9813f5ab42c4479... Do you think it should be removed from Orchestrate? Any other suggestions for proprietary products in that category?
- slap_shot 8y agoAs someone who works very, very closely in this industry, I would just be very careful how much of this you think you want to bite off. Consider how you trust using dbt more than rolling your own transformation tool. Why wouldn't this apply to the rest of your stack? The 10+ companies that offer data extraction and loading are likely a better choice. Again with Analytics - the dozens of companies that offer BI tools are probably going to be the better choice. Maybe you can build all these tools better than the hundreds of companies with thousands of employees and millions of dollars. It just seems like the odds that you build the best of each is so unlikely. I would have been more impressed if your team had designed some API that other tools/platforms could plug in to coordinate a lot of the above jobs with your CI system. There is a SERIOUS need for that and I've had a lot of conversations with companies about what that would look like. To answer your quest, no, Fivetran does not currently belong in the orchestration area, IMO. I've heard they are soon to release some sort of orchestration tooling to compete with dbt, but it isn't the type of orchestration you get with Airflow.
- cheghook 8y agoI can't understand why GitLab thinks they have to embark on a new project every so often instead of focusing on their current product and features. There is just a lot to work on, so many of the current features/products are half assed. At my place we moved to GitLab 2.5 years ago and updates where smoother back then but the past few months we had to hire a new sys admin for our build machines and GitLab server to follow on new issues created on GitLab.com and decide if it's safe release and even then he still reports 4-5 issues to GitLab support after every update. We were expecting it to be an easy `yum update` like a normal package but it's just getting worse update after update. It's so bad that my manager asked me to look into GitHub + another CI/CD solution.
- sytse 8y agoI'm sorry to heard your experience with GitLab hasn't been smooth. We have more people then ever working on the core of GitLab. And the number of reported issues per customer are going down. But every problem is one too many. Please email me at sytse@gitlab.com if you're open to a call about your situation.
- wmccullough 8y agoI appreciate how dedicated GitLab is to continuously improving the product. Thinking about moving my projects from Bitbucket to GitLab for that reason.
- sytse 8y agoYay! Thanks for that.
- cheghook 8y ago> And the number of reported issues per customer are going down. This doesn't mean anything, maybe the customers are simply tired of reporting issues. For example last year we didn't do any updates for 6 months because we were afraid it'd break something and we were too busy to be willing to spend the time reporting problems. We also don't report issues that are already open on gitlab.com, reporting the issue means your customer is willing to spend time reporting, following up and testing your bug. This is your job, not the customer's. At the moment we are only reporting issues that are either blocking us from work or slowing down our development. The majority of issues we are facing are performance problems. I just wrote a script to plot the number of issues on gitlab-ce over time and percentage of open/close issues, and the overall period they have been open for, you are accumulating issues with: `backend`, `UX`, `technical debt`, `performance`, `CI/CD`, ... labels, a lot of them don't have a Milestone and have been open for a long time. I am not sure how emailing you would help us, it's not like the problems are not reported or you don't already know about them. It just appears that the priority of GitLab, as a company, is not shipping a quality product anymore. EDIT: I work in the aerospace industry and one of the stages of our pipelines is to run stress test on our product. I would suggest you to run a stress test on an instance of GitLab, this would be an amazing place to start looking for performance problems.
- deleted 8y ago[deleted]
- n42 8y agoIs there any example of an open source software company that has taken on so many products at once, so early in its life, and succeeded?
- sytse 8y agoWe did https://about.gitlab.com/2017/10/11/from-dev-to-devops/ https://about.gitlab.com/2017/10/11/from-dev-to-devops/ when we where at 50% of our current number of engineers. So far so good.
- gregoriol 8y agoReally no, look at all the comments here (and this is only from techies): you have lost us, we don't know anymore what you are doing, or even trying to do.
- sytse 8y agoWe are trying to make a single application that covers the whole DevOps lifecycle, from planning your change up to monitoring its effect. We're doing it because we believe there are emergent benefits to having the lifecycle in a single application https://about.gitlab.com/handbook/product/single-application/#emergent-benefits-of-a-single-application https://about.gitlab.com/handbook/product/single-application...
- gregoriol 8y agoThe idea seems great, but it's not working: there is no single application that can fit all uses, and you are loosing most of the users on the way. I'm using Gitlab, btw, but only for the self-hosted git and it's user interface (ie. your core). All the other parts (bug tracking, CI, chat, ...) are in different and more appropriate tools for each of our use-cases... because most of yours are not complete enough, or sometimes it's not even clear how they actually could work for us (mattermost for example).
- n42 8y ago
- tamersalama 8y agoIs there some resemblance with Floydhub http://floydhub.com/ http://floydhub.com/ ?
- veritas3241 8y agoPersonally, I quite like the approach FloydHub has for deep learning projects. At GitLab, we currently don't have any deep learning projects happening - we're still further down the AI hierarchy of needs - i.e. focusing on solid data infrastructure and descriptive analytics. I fully expect we'll have a use case for the "cool" machine learning stuff, but there's a lot of groundwork to cover with the basics first. Meltano is focusing on those basics for right now.
- NegatioN 8y agoDoes anyone have a comprehensive list of similar offerings to floydhub? or OSS alternatives? I think this market is not being served properly, most of them seem to still require most of the heavy lifting to be done by the ML practitioner. I suppose I would even be okay with a service that just saves all my graphs from tensorboard for later reviewing.
- houqp 8y agoI am interested in knowing more about how you think FloydHub can better serve the market. FloydHub does have metrics support for later reviewing: https://docs.floydhub.com/guides/jobs/metrics https://docs.floydhub.com/guides/jobs/metrics. Are you only interested in using tensorboard for graph viewing?
- georgewfraser 8y agoData pipelines are not a great subject for an open-source project. We've been building these for the last 3+ years at Fivetran, and I can tell you that the challenge is: - Studying each source to figure out the right data model - Chasing down a million weird corner cases - Working around dumb bugs in the data sources This is the kind of problem where paying for software really works better. When people build data pipelines in-house, they tend to hack at it until it works for their use case and then stop. When we build data pipelines, we map out every feature of the data source, implement the whole thing at once, and then put it through a beta period with multiple real users. This is easy to do when you have a tight-knit dev team; much harder for a group of part-time open-source contributors.
- MechanicalTwerk 8y agoI kind of agree with this. To take an example outside of ETL/DW/BI, when I first saw Zapier I was skeptical of how many APIs they could support because I'd seen a decent amount of open source ESBs like Mulesoft run out of steam after a certain number of connectors. Zapier, being proprietary from day one (albeit less featureful than a full blown enterprise ESB) has done better than I expected. Still, they only support 100 or so datasources and the types of data/objects/triggers/whatever they support is limited at times. IMO at some point both open source and proprietary models fall apart in the face of the long tail. Amazon has tackled the long tail of ecommerce but that's an enormous market that allows them to employ hundreds of thousands of people to tackle that long tail. Tackling the long tail of connectors (whether it's for ESBs/SaaS integration or ETL/DW/BI) is just too expensive compared to the size of the markets that are willing to take a shot at it.
- js8 8y agoCan't agree more. IMHO, if you want to make a dent in the space, figure out better debugging tools! In particular, tools that explain how a certain (specific) value was calculated in the system, tools that let you bisect the source data in some way and let you focus on the source data that are likely to have a problem, tools that help you figure out that certain intermediate value in calculations is an outlier, tools that let you test certain assumptions about data over the whole pipeline..
- gandutraveler 8y agoLooks like gitlab just wants to be in news since Microsoft's aquisition of GitHub.
- _pmf_ 8y agoGitLab's usage of team members in marketing material is creeping me out (as does the whole team page[0]). [0] https://about.gitlab.com/team/ https://about.gitlab.com/team/
- sytse 8y agoWe say team members instead of employees because some are contractors. Why does it freak you out? BTW We don't call it a family https://about.gitlab.com/handbook/leadership/#management-team https://about.gitlab.com/handbook/leadership/#management-tea...
- _pmf_ 8y agoI wouldn't want to have that level of public affiliation with my employer (no matter who that employer might be).
- danpalmer 8y agoReading this I was concerned that it would be written in Ruby. While Ruby is a reasonable language for server development, it has almost no data science community when compared with some other ecosystems. I was very glad to see this is Python! Python has some of the best data tools out there, and a mature ecosystem for solving all the engineering problems that go along with a great data stack.
- ksec 8y agoI am on the opposite side, Given Gitlab is a Ruby house I was secretly hoping some innovation coming from Ruby Data Science.
- sbr464 8y agoAre you releasing/sharing any of the extractors you built for various services?
- jakecodes 8y agoAll of our extractors are available in our source code, which is open source. http://gitlab.com/meltano/meltano/ http://gitlab.com/meltano/meltano/. Right now we are working towards an MVP, so things might be in flux, but we value any feedback you have.
- Luuseens 8y agoThe page talks mentions MVC, and the issue page[0] keeps mentioning MVC as well. Was this supposed to be MVP, or something else? Model-view-controller doesn't make sense in the context. [0] https://gitlab.com/meltano/meltano/issues/10 https://gitlab.com/meltano/meltano/issues/10
- jakecodes 8y agoWe use the term mvc here, as "minimal valuable change", in a recognition that it may not be a product yet.
- tbrock 8y agoI wish they would focus on making a fast, stable, GitHub alternative.
- parasubvert 8y agoThis is Gitlab taking stuff they were doing already internally and making it available to a broader audience. Once you take VC funding, you gotta go where the money is. Everyone wants/expects "fast, stable, like Github" for free unless you have special needs. So, you do analytics on what people are doing with your free site, you offer enterprisey features, you get into the "platform" business etc. I think Gitlab distracts itself, spreads itself thin, and isn't great at partnering, its ambition to do-it-all knows no bounds, which is both commendable and a smh moment. It's not likely sustainable or scalable. They're definitely trying to "go big or go home" as a company, which is not how most originally felt about Gitlab (a fast, stable OSS alternative to Github). At the same time, I can't blame them. I think it comes down to: Don't hate the player, hate the game.
- sytse 8y agoWe are building a fast, stable, GitHub alternative. We have hired 3 times as many people in our security team for GitLab.com (not our product team for security) as are working on Meltano. We have hired 3 times as many people in our SRE teams as are working on Meltano. And we still have a lot of vacancies for both https://about.gitlab.com/jobs/ https://about.gitlab.com/jobs/
- parasubvert 8y agoI meant no offense. But you’re also dabbling in building your own k8s distro / platform, you have your own CI/CD , Jira storyboard , and data science stuff, etc. My point is that you’re aiming a lot broader than Github ever did - you are competing more as a suite than as a focused product. And I’ve seen personally this impact the support side with customers, partnership side, etc. I help maintain a medium-large Gitlab for one of your bigger customers. Anyway this isn’t the place for me to get specific, I am just saying that you are taking a risky path in terms of sustainability IMO as a rando on the internet.
- ajbosco 8y agoDo you see this as a (future) competitor of Airflow/Luigi type workflow tools?
- sytse 8y agoYes, the orchestrate part (working on GitLab CI) is an alternative for Airflow. Also see https://gitlab.com/meltano/meltano/blob/master/README.md#data-science-lifecycle https://gitlab.com/meltano/meltano/blob/master/README.md#dat...