9 ms·
Launch HN: Dataform (YC W18) – Build Reliable SQL Data Pipelines as a Team
Hi HN!
We’re Guillaume and Lewis, founders of Dataform, and we're excited (and nervous) to be posting this on HN.
Dataform is a platform for data analysts to manage data workflows in cloud data warehouses such as Google BigQuery, Amazon Redshift or Snowflake. With our open source framework and our web app, analysts can develop and schedule reliable pipelines to turn raw data into reliable datasets they need for analytics.
Before starting Dataform, we managed engineering teams in AdSense and led product analytics for publisher ads. We heavily relied on data (and data pipelines!) to generate insights, drive better decisions and build better products. Companies like Google invest a lot to build internal data tools for analysts to manage data and build data pipelines. In 5 minutes I could define a new dataset in SQL that would be updated every day and then use it in my reports.
Most businesses today are centralising their raw data into cloud data warehouses but lack the tools to manage it efficiently. Pipelines run manually or via custom scripts that break often. Or the company decides to invest engineering resources to set up, maintain and debug a framework like Airflow. But that’s just for scheduling and the technical bar is often too high for analysts to contribute.
We saw a need for a self-service solution for data teams to manage data efficiently, so that analysts can own the entire workflow from raw data to analytics. We built Dataform with two core principles in mind:
1. Bring engineering best practices to data management.
In Dataform, you build data pipelines in SQL, and our open source framework lets you seamlessly define dependencies, build incremental tables and reuse code across scripts. You can write tests against your raw and transformed data to ensure data quality across your analytics. Lastly, our development environment also facilitates the adoption of best practices, where analysts can develop with version control, code review or sandboxed environments.
2. Let data teams focus on data, not infrastructure.
We want to bring a better, faster and cheaper alternative to what businesses have to build and maintain in-house today. Our web app comes with a collaborative SQL editor, where teams develop and push their changes to GitHub. You can then orchestrate your data pipelines without having to maintain any infrastructure.
Here's is a short video demo where we develop two new datasets, push the code to GitHub and schedule their execution, in under 5 minutes.
https://www.youtube.com/watch?v=axDKf0_FhYU https://www.youtube.com/watch?v=axDKf0_FhYU
You can sign up at https://dataform.co https://dataform.co. If you're curious how it works - here are the docs: https://docs.dataform.co https://docs.dataform.co and the link to our open framework: https://github.com/dataform-co/dataform https://github.com/dataform-co/dataform
We would love to hear your feedback and answer any questions you might have!
Lewis and Guillaume
- buremba 7y agoCongrats on your launch! How is this different from DBT?
- G2H 7y agoOur framework (and the CLI interface) is pretty similar to be honest, except it’s in JS. We also spent quite a bit of time focusing on performance. The main difference is that we provide an all-in-one web platform for teams to develop, test and schedule their SQL pipelines. With one web interface and in 5 min, you can write a query, test it, add it to a schedule, push the changes to GitHub (or submit a PR) and monitor your jobs. We make it easy to enforce software engineering best practices like version control and isolated deployments. For example, our version control works similarly to Looker. Analysts can work simultaneously from different branches while not having to use different tools nor use the command line.
- buremba 7y agoWhat’s the rationale behind using JS instead of Python/Jinja? Any insights on that?
- 1ewish 7y agoGreat question :) 1. Speed. We wanted compilation to really, really fast. 2. In all honesty we just weren't big fans of Jinja, having used it for quite a while. JS templating is OK out the box, and we are considering React like syntax in the future. 3. Love it or hate it, NPM packages are pretty easy to work with and we are working on a number of packages at the moment. 4. When you start to look at things like UDFs and Cloud functions which enables some really cool use cases, JS seems to be prevailing (in Snowflake and BigQuery at least). I will admit though that we do usually get a bit of a shocked reply when people hear it's not Python!
- buremba 7y agoI see your point but I think that data analysts are usually familiar with Python, not JS. The UI is cool and intuitive, it looks like you copied most of the concepts except Jinja from DBT. That's perfectly OK though, let's see where it goes!