8 ms·
Definitely a fair point! The primary reason we haven't provided pricing is that we have just launched and wanted to collect more data points before setting mak
by hichkaker 6y ago
Definitely a fair point!
The primary reason we haven't provided pricing is that we have just launched and wanted to collect more data points before setting making the pricing public.
Our current offering is:
1) Free for diffing datasets < 1M rows
2) $90 / mo / user for diffing datasets of unlimited size
3) Cross-database diff, on-prem (AWS/GCP/data center) deploy, Single-sign-on are in custom-priced enterprise bucket.
We would love to hear your thoughts on this.
- lukevp 6y agoCould you use the free tier for regression testing a subset? Like up to 1M either first, last or random sample? Or do the datasets themselves have to be prefiltered down to 1M results?
- hichkaker 6y agoYes, you could. The limit is for the source dataset size but you can prefilter it (there is an option to pass in free-form SQL query instead of table name when creating a diff). For the majority of diffs we see with sampling applied, sample sizes are <1M rows (more is often impractical in terms of information gain for higher compute costs) especially if your goal is to assess the magnitude of the difference as opposed to get every single diverging row.
- GordonS 6y agoDoes your solution diff database schema too, or is it purely for diffing data in identical schemas? I'm also really keen to hear why you built this tool - what use cases you expect. I've used free diffing tools a few times before in the past, but I think every time it was to make sure I hadn't messed up "manual" data migrations (which obviously aren't a good idea).
- hichkaker 6y agoYes, it does diff the schema too. The main use cases we've seen: 1) You made a change to some code that transforms data (SQL/Python/Spark) and want to make sure the changes in the data output are as expected. 2) Same as (1) but there is also some code review process. In addition to checking someone's source code diff, you can see the data diff. 3) You copy datasets between databases, e.g. PostgreSQL to Redshift and want to validate the correctness of the copy (either ad-hoc or on a regular basis). We have a signup-free sandbox where you can see the exact views we provide, including schema diff: https://app.datafold.com/hackernews https://app.datafold.com/hackernews
- ironchef 6y agoMost folks I know are doing this (1 and 2) by testing against a replica (or in the case of snowflake just copying the DB or schema) ... then running data tests locally and downstream (great expectations, DBT tests, or some airflow driven tests). Is the value prop “you don’t need all they grunt work” as opposed to above direction?
- hichkaker 6y agoYou raised a great point. Data testing methods can perhaps be broken down to two main categories: 1. "Unit testing" – validating assumptions about the data that you define explicitly and upfront (e.g. "x <= value < Y", "COUNT(*) = COUNT(DISTINCT X)" etc.) – what dbt and great_expectations helps you do. This is a great approach for testing data against your business expectations. However, it has several problems: (1) You need to define all tests upfront and maintain them going forward. This can be daunting if your table has 50-100+ columns and you likely have 50+ important tables. (2) This testing approach is only as good as the effort you put to define the tests, back to #1. (3) the more tests you have, the more test failures you'll be encountering, as the data is highly dynamic, and the value of such test suites diminishes with alert fatigue. 2. Diff – identifies differences between datasets (e.g. prod vs. dev or source DB vs. destination DB). Specifically for code regression testing, a diff tool shows how the data has changed without requiring manual work from the user. A good diff tool also scales well: it doesn't matter how wide/long the table is – it'll highlight all differences. The downside of this approach is the lack of business context: e.g. is the difference in 0.6% of rows in column X acceptable or not? So it requires triaging. Ideally, you have both at your disposal: unit tests to check your most important assumptions about the data and use diff to detect anomalies and regressions during code changes.
- sails 6y agoI think doing a deeper analysis into why this is a good tool in addition to dbt would be useful for me to understand. Locally Optimistic [] has a slack channel and do vendor demos, with a _very_ competent data analytics/engineering membership. I think you'd do well to join and do a demo! [] https://locallyoptimistic.com/community/ https://locallyoptimistic.com/community/
- simonebrunozzi 6y agoI have a quick suggestion for you: two options you can mitigate this "issue". Option 1: make it free, up to a certain dataset size. You can harvest interested leads like the gentleman above. Option 2: (if you don't want to deal with a huge volume) offer it for $50 one-time fee, up to X size, for Y months (e.g. $50, up to 1 GB, valid for 3 months). Nice way to filter qualified leads. There are variations from the two options above, but I think you can easily get the general idea. Thoughts?
- hichkaker 6y agoThank you for the suggestion! We're leaning towards Option 1: free diffing for datasets < 1M rows. Option 2 seems a bit tricker since we are in a way creating a new tool category and it can be harder to convince someone to pay before they try and understand the value (unlike, say, a BI tool – everyone knows they need some kind).
- GordonS 6y ago$90/m/user is a lot, especially when you consider the several other SaaS services you could get together for that price. That's definitely enterprise pricing, which fits with the "call me pricing" I guess. It's a niche, and a small one at that, but I don't doubt you'll find some enterprises willing to pay what you ask. But outside of enterprise I just can't see anyone paying that, especially when free tools exist (albeit not nearly as polished and features as yours). I've got to wonder about YC backing for what seems like such a small niche though - very possible I'm not seeing something you have planned for further down the line.
- hichkaker 6y agoThank you for the feedback! Agree with you about the niche. Diff is our first tool that helps test changes in the ETL code, and the impact is correlated with the size and complexity of the codebase. Diff also provides us a wedge into the workflow and a technical foundation to build the next set of features to track and alert on changes in data: monitoring both metrics that you explicitly care about and finding anomalies in datasets. We've learned that this is something a much larger number of companies can benefit from.
- sterlinm 6y agoI don't agree that $90/user/month is unreasonable in every context. Yeah it's probably too much for consumers but honestly the consumer need for this tool seems pretty niche to me. It's also probably too much for large enterprises where you'd have a lot of people who want the tool, but they're probably going to either build it themselves or pay $$$$ for big lame ETL tools. For mid-sized companies though that could be a bargain. I worked on a data migration project at a company with <100 people and <5 engineers where we had to hack together our own data-diff tools and this would have been a bargain.