15 ms·
Databricks response to Snowflake's accusation of lacking integrity
- michaelhartm 5y agoData Wars: Snowflake vs Databricks (0 - 2)?
- drawturkey 5y agoSnowflake has way more revenue, is worth 3 times more than Databricks and is growing faster. I'd say Snowflake is still in the lead. Plus, just look at Snowflake's customer list. It's a "who's who", Databricks is a "Who's that?".
- thrtlvlmidnight 5y agoI took a look at Databricks public customer case studies[1] and haven't a clue who any of these companies are: Atlassian? Adobe? ExxonMobil? PagerDuty? McAfee? HSBC? Starbucks? AstraZeneca? GlaxoSmithKline? Comcast? FINRA? Regeneron? Riot Games? Nielsen? HP? Conde Nast? Viacom? McGraw-Hill? Cisco? NBCUniversal? Hopefully they can scale to the enterprise soon. [1]https://databricks.com/customers https://databricks.com/customers
- glogla 5y agoLists of "references" like these are worthless. Because larger companies tend to be fragmented, especially companies that have more complicated business lines and are used to departments and divisions acting independently. You know what, our company uses both Snowflake and Databricks. For Databricks, there's one or two projects that someone built on it running in production. For Snowflake, there's a sizeable use because we bought a smaller company that used it for reporting and warehousing. Neither of them are "the chosen tool" and will see any growth unless wind changes. But we could be (F50 company) counted as reference by both I guess.
- falaki 5y agotl;dr: The data warehouse company used a pre-baked TPC-DS dataset and claimed they have similar performance to Databricks. Turns out if you use the official TPC-DS data generation scripts, you get much worse performance.
- slownews45 5y agoEven worse, they claimed to have similar performance to Databricks AND claimed databricks "lacked integrity". WOW, talk about chutzpah!
- arnon 5y agoThat's altering the methods - and generally considered a violation of the validity of the results.
- tyingq 5y agoI read the original post, the Snowflake response, and this. From that I gather that both of them aren't being completely honest or fair when making comparisons. A fair amount of truth, but also some clever wording and omission on both their parts. Which is not surprising or particularly new in this space :)
- slownews45 5y agoDatabricks results are available at tpc.org [1] Snowflake has shown NOTHING close to this. [1] http://tpc.org/results/fdr/tpcds/databricks~tpcds~100000~databricks_sql_8.3~fdr~2021-11-02~v01.pdf http://tpc.org/results/fdr/tpcds/databricks~tpcds~100000~dat...
- tyingq 5y agoYes, I wasn't saying they were lying about their tpc.org posted results. I'm saying both companies made use of clever indirection, wording, presentations of stats, etc. Like price/performance, and which of your competitor's tier's to select when doing that, and which of your own. Or over-provisioning the competition's setup, for example.
- naattee 5y agosnowflake should just pony up and do a TPC-DS audited benchmark
- maslam 5y agoEveryone win when data platforms submit audited benchmarks...
- avip 5y agoI've used both products in production. Both are good++. The blog wars seem extremely ridiculous to me. I don't recall ever choosing one over another based on how fast it runs on some imaginary arbitrary dataset.
- paxys 5y agoManufactured rivalries can be a great thing for business. We have been debating Coke vs Pepsi, Nike vs Reebok, McDonald's vs Burger King for decades now while these companies laugh all the way to the bank.
- javajosh 5y agoLike the post but I would add "Ford v Ferrari" there. A synthetic 100T test is much like an F1 course - not something you deal with during your commute, but it's nice to know what the limit is, and that there are people pushing that limit.
- kartoonhero 5y agoIts not ridiculous at all. This is the coming of age for a brand new data architecture. One of the biggest FUDs for a data lake architecture is performance - and this benchmark should put that concern to rest.
- syntaxfree 5y agoI don’t know, “coming of age” seems to imply that there’s some pre-maturity period out of which something is emerging.
- buttaphingas 5y agoI actually see them as variations on the same architecture. Databricks keeps their metadata in files, Snowflake keeps theirs in a database, but they both, ultimately, are querying data stored in a columnar format on blob store (and, to be fair, Snowflake have been doing that with ACID-compliant SQL for a lot longer than Databricks). So using SQL over blob at high performance has been around for a while. Databricks say their solution is better because it's open (though keep the optimizations you need to run this at scale to themselves, i.e. is ultimately proprietary). Snowflake says theirs is better because it's a fully managed service, meaning no infrastructure to procure or manage, is fully HA across multiple data centers by default etc. Databricks push 'open' but really still want you to use their proprietary tech for first transforming into something usable (Parquet/Delta) and then querying with Photon/SQL, though you can also use other tech. With Snowflake you can just ingest and query, but it has to be through their engine. Customers should do their own valudation and see which one fits their needs best.
- 1cvmask 5y agoThis reminds me of the old performance ads of Oracle where they would show you how everything ran better on Oracle. They used to put those ads at airports, business lounges and the back cover of newspapers and magazines read by non-technical executives like the FT and Economist. Everyone technical knew they would game every environment to come out with superior results. I suppose it worked. As the top executives buy big system software and ignore the IT crowd who could easily point out the flaws in the methodology of the"studies". Breakdown of one of those example ads: https://db2news.wordpress.com/2011/06/08/a-closer-examination-of-oracles-database-performance-advertisement/ https://db2news.wordpress.com/2011/06/08/a-closer-examinatio...
- supercanuck 5y agosimiliar as to how SAP is still showing growth even thought their core product (ERP Financials) hasn't changed much.
- initplus 5y agoA key part of the Oracle strategy is making it a breach of license to publish any benchmarking data. No performance data about Oracle's database is allowed to be published without their approval, which means no negative results are published.
- doppelganger1 5y agoOracle Exadata is very fast but expensive. I bet it would beat a similarly sized cluster from these 2 vendors. The problem is price to performance and elasticity. Because DB and SF are in the cloud, they have a lot more options that Oracle doesn’t have. This is why Kurian left Oracle to go to Google, because LE would not allow Oracle to make cloud native products that would run in other clouds. The SF cofounders are ex Oracle engineers and LE was not interested in creating a cloud native DB from scratch. If he did, we wouldn’t have a SF computing right now.
- glogla 5y agoYeah, the biggest benefit something like Snowflake or Databricks or whatever AWS tools has over the more traditional technologies is the pay-as-you-go pricing. We're are now trying to scale unnamed technology running on EC2 from 100 nodes to 200 cores and the process to buy larger license is pretty painful. If we were using Snowflake or Databricks, we could just scale it up and update our opex estimate.
- Normal_gaussian 5y agoso, alternatives? Aside from the Azure/GCP/AWS internal offeringa I know about Snowflake and Firebolt, Databricks is new to me.
- ethbr0 5y agohttps://en.m.wikipedia.org/wiki/Databricks https://en.m.wikipedia.org/wiki/Databricks "Databricks is an enterprise software company founded by the creators of Apache Spark. [...] Databricks develops a web-based platform for working with Spark, that provides automated cluster management and IPython-style notebooks."
- kofejnik 5y agomaybe clickhouse?
- glogla 5y agoClickhouse is good if you're building application. It has lot of great features and incredible performance, but there's an expectancy that people using it know what they're doing and can work around its limitations (like limited support for joins and sql in general). Something like Snowflake works much better when you're building a platform that you can give to two hundred data analysts or various skills spread over fifty teams, so they can build their own stuff. The nice UI, broad feature set (materialized views, time travel, automatic backups, superfast scaling up and down, ...) and general just-work-iness makes it nice for that, but you're going to pay for the privilege. Databricks is somewhere in the middle - things are way less polished, features don't always work and you still have to figure out things like backups and partitions on S3 on your own, but some people like that. Expect to also pay a pretty penny for hundreds of Spark clusters nobody knows who uses.
- solidangle 5y agoWhen was the last time you used Databricks? You should definitely try it again. Their product offering has improved a lot in the past few years. > broad feature set My experience is that the feature sets of Snowflake and Databricks are very similar. Both have time travel support. Snowflake has materialized views, but Databricks has Delta Live Tables. Databricks has a distributed Pandas API, but Snowflake recently introduced Snowpark. Databricks also has autoscaling and they recently launched a serverless offering that makes autoscaling super fast aswell.
- xiaodai 5y agoLol
- __MatrixMan__ 5y agoInstead of blog posts written but experts in app A based on their experience with app B, I wish there were a platform for this kind of comparison. Some objective third party sets the goal and then each company submits automation (selenium?) that configures their own app to achieve the goal. Entrants are scored by: - time - storage - compute - config complexity No need to waste time making your opponent look bad, just focus on making your self look good, and do it on a level playing field.
- rxin 5y agoIsn’t that what the official TPC does?
- falaki 5y agoThat is exactly the role of tpc.org.
- renewiltord 5y agoIf you want some information like this quick, you're gonna have to pay to run it.
- dreyfan 5y agoDatabricks is a rapidly approaching IPO. Trying to justify their valuation with their overpriced in-memory hadoop.
- kartoonhero 5y agoDatabricks is way more than hadoop or spark. A great analogy - Spark is a great engine but you need to design and build all of the other subsystems. Databricks is an F1 car - everything is built out. You get in and drive - FAST.
- glogla 5y ago> Databricks is an F1 car F1 cars really unreliable and need a lot of engineers to keep running, are very expensive, and completely impractical in normal use. They are fast but only on very specific roads, they couldn't survive on normal roads. What do you know, you might be right! :D
- drawturkey 5y agoYou nailed it. Meanwhile the rest of the world just needs a camry.
- dreyfan 5y agoDatabricks is a shit platform that encourages terrible data practices and accretion of technical debt.
- exsmelliarmus 5y agoSeems pretty good to us! Can you give more information?
- 0x500x79 5y agoAs people noted elsewhere, you have to be VERY careful with using databricks for a full data warehouse due to the fact that it drives you to notebook driven development and scheduling of those notebooks when data pipelines should follow similar development practices as other software projects. Great for proof of concepts, but when you start to build out complete pipelines please look into how to make the pipelines more sustainable and maintainable.
- gnabgib 5y agoRelated post (2 days ago, 95 comments): [Snowflake’s response to Databricks’ TPC-DS post](https://news.ycombinator.com/item?id=29206959 https://news.ycombinator.com/item?id=29206959)
- redwood 5y agoAs much as I love seeing competition in the space and am enjoying my popcorn, I really don't understand what Databricks is doing here: this feels like a childish foodfight rather than an obsession with the customer...
- kf6nux 5y agoI'd say helping customers spot fraud* is serving the customers' interests. * I haven't executed the test suite, but fraud seems likely.
- cai22r 5y agoinsert gif: he started it
- jjoonathan 5y agoAll publicity is good publicity. Both participants in a fight can win by implicitly excluding their real competitors.
- s_barrow1 5y agoDatabricks is not known for the SQL/DW space. The original blog was focused on breaking the TPC-DS performance record and provide validation of the Lakehouse architecture. DB didn't ask for a war of words with Snowflake - SF dedicated a whole response stating DB lacked integrity and filled it with false and misleading information. I commend DB for responding back (only because of the integrity accusations). Snowflake has asked for this response by acting petty from the outset
- saj1th 5y ago:) That is a good question. Why spend eng cycles to submit results to the TPC council - why not just focus on customers? I believe the co-founders have addressed this in the blog. > Our goal was to dispel the myth that Data Lakehouse cannot have best-in-class price and performance. Rather than making our own benchmarks, we sought the truth and participated in the official TPC benchmark. I'm sure anybody seriously looking at evaluating data platforms would want to look at things holistically. There are different dimensions like open ecosystem, support for machine learning, performance etc. And different teams evaluating these platforms would stack rank them in different orders. These blogs, I believe, show that Databricks is a viable choice for customers when performance is a top priority (along with other dimensions). That IMO is customer obsession.
- hello_moto 5y agoSerious question: Databricks, Snowflake, Dremio. All these "Data" platform companies => which one do you have for your Data Lake and Data Warehouse solution? I'm sick and tired of these companies Snake Oiling the Data industry by offering "the easiest" platform to satisfy your Data Lake + Warehouse solution only to fall hard whenever you hook it up with your production data (big dataset). PS: Anyone selling Data Lakehouse (Data Lake + Warehouse as one platform) is on meth.
- kartoonhero 5y agoPlease read up on Lakehouse. Data Lake + Merge support + DW performance is now possible. That is the game changer.
- strongbond 5y agoDo you work for Databricks?
- bpaneural 5y agoThey must do. But if you've been in this area for long enough, I'd put my money on Databricks, if anything, because of their open source integrity
- buttaphingas 5y agoDatabricks isn't open source, as they keep hold of all the IP that makes it much better than OS Spark. Whether you buy Snowflake or Databricks, you're buying proprietary software.
- ttmahdy 5y agoWith Snowflake data is locked away in a proprietary format not accessible by other compute platforms. You need to export/copy your data to a different system to train an ML model in python or R. With the Databricks, you can use python, R and Scala, (not just SQL) to interface with your data. You can use multiple compute engines (Spark, presto and other engines that support Delta) so you are not locked into one compute engine.
- dautkhanov 5y agoThanks Snowflake for removing the DeWitt clause, makes performance comparison more transparent. Would be best for Snowflake to complete official/audited TPC-DS benchmark so customers can compare apples to apples.
- jchw 5y agoBefore the Snowflake blog post, I did not know what Snowflake or Databricks were. I can only imagine that this rivalry is great for both of them, even if Databricks is somewhat on the advantage end, at least from a tactical standpoint; I admit though that they seem to be a bit unnecessarily defensive considering the position they're in with the exchange. In general though, I'm still not complaining. It's interesting to see a dispute like this unfold.
- qaq 5y agoSnowflake is 120B Market Cap Darling of Cloud Data warehouses I doubt obscurity is a problem they are trying to solve
- jchw 5y agoOf course they’re known among their pre-existing customer base of people and entities who already solve problems using tools like this. But it’s a subset of the multi-trillion dollar cloud industry, which itself is not the entire software engineering industry.
- scapecast 5y agoThe irony here is that what Databricks is doing to Snowflake is exactly what Snowflake did to AWS and Redshift. Same playbook - show that you’re better in a key metric that’s easy to understand (performance) to get the attention, but then pitch the paradigm change. In Snowflake’s case, that was separation of storage and compute. In Databrick’s case, it’s the Lakehouse Architecture. I think the reason why Snowflake is so nervous because they know they can’t win this game.
- glogla 5y agoIn what way is lakehouse architecture beneficial over something like Snowflake or BigQuery? I understand the appeal over having lake and warehouse as separate components, but with those native cloud warehouses, you can already do everything a lake does.
- doppelganger1 5y agoBig Query&Data Proc, Redshift&EMR, Synapse&HDR are tied to the cloud vendors. You can’t move easily from AWS stack to GCP without refactoring. Switching costs are higher. Snowflake and Databricks are multicloud. The different is that Snowflake is more like a SaaS solution and only does SQL. Databricks is more than just SQL. It has all the data science, machine learning information, built into it. Snowflake has Snowpark but it’s every limited and so you are more likely to have to buy more products to build out your capabilities and integrate them with Snowflake. With Databricks it is more out of the box in terms of capabilities. Databricks also runs in your cloud account which has trade offs. It can be harder to get going and more complex but you end up with a lot more flexibility and you own your data and have complete control over it. While Snowflake gives you control of your data with their tools, everything has to go through Snowflake and incur their tax to get to it. You pay for simplicity, which many customers are ok with because they see value in it. On the contrary, a lot of customers see value in having more control and options. This market is big enough for everyone - it’s really just about market share.
- turk- 5y agoWith a datawarehouse, you can only interface with your data in SQL. With big query and snowflake, your data is locked away in a proprietary format not accessible by other compute platforms. You need to export/copy your data to a different system to train an ML model in python or R. With the lakehouse, you can use python, R and Scala, (not just SQL) to interface with your data. You can use multiple compute engines (spark, Databricks, presto) so you are not locked into one compute engine. I recall being a junior programmer, and wishing I could talk to my MySQL database in python code to do some processing that was difficult to express in SQL, that day is finally here.
- boringg 5y agoAnd how soon is the S-1 for Databricks dropping?
- drej 5y agoWhat I find hilarious is that companies argue who can query 100 TB faster and try to sell this to people. I've been on the receiving end of offers by both of the companies in question and used both platforms (and sadly migrated some data jobs to them). While they can crunch large datasets, they are laughably slow for the datasets most people have. So while I did propose we use these solutions for our big-ish data projects, management kept pushing for us to migrate our tiny datasets (tens of gigabytes or smaller) and the perf expectedly tanked compared to our other solutions (Postgres, Redshift, pandas etc.), never mind the immense costs to migrate everything and train everyone up. Yes, these are very good products. But PLEASE, for the love of god, don't migrate to them unless you know you need them (and by 'need' I don't mean pimping your resume).
- autokad 5y agoits my experience if its just 10s of GBs then use 'normal' solutions. if TB then spark is great for that. note I have only used DataBricks & Spark, no snowflake.
- jeltz 5y agoPostgreSQL and MySQL can handle a few TB just fine. It is when you reach over 10TB that you need something else.
- tshanmu 5y agoResume driven development FTW!
- deleted 5y ago[deleted]
- StephenJGL 5y agoVery true. You have to understand the actual capabilities and your actual requirements. We work with petabyte size datasets and BigQuery is hard to beat. Our other reporting systems are still all in MySQL though.
- sanketsarang 5y ago
- benjaminwootton 5y agoIve been following this and it’s kind of embarrassing to watch. I love working with Databricks and Snowflake. They both knock it out of the park for their respective use case. They’re amazing products. It makes no sense to fall out about this though. For a 100TB dataset with a funky calculation, Spark will trounce Snowflake. For a 1 row dataset, Snowflake will return before the spark job has been serialised.
- imslowbutnice 5y agoWhat are you talking about. Spark isn't even used, and TPC DS is not a funky calculation at all. It's supposed to be a collection of typical datawarehouse type queries. Although I'm not really sure what funky means, but why would Spark trounce Snowflake on "funky" calculation at all. Do you mean an ML algorithm, and are you implying that TPC-DS has anything close to an ML Algorithm? And why would Snowflake perform better on returning one row, they are columnar stored.
- deleted 5y ago[deleted]
- nojvek 5y agoWhy would Spark trounce Snowflake. What makes it inherently so much faster at 100TB jobs? Also what kind of queries are we talking about?
- saj1th 5y ago> Why would Spark trounce Snowflake. What makes it inherently so much faster at 100TB jobs? These are the slides from a talk one of the co-founders (@rxin) gave at Stanford. https://web.stanford.edu/class/cs245/slides/LakehouseGuestTalk.pdf https://web.stanford.edu/class/cs245/slides/LakehouseGuestTa... It goes into the details of how this performance is achieved(and not just at 100TB). Part of this could be attributed to innovations in the storage layer(delta lake), and part of it is just the new query engine design itself.
- deleted 5y ago[deleted]
- bloodyplonker22 5y agoDatabricks is trying to punch up at the market leader. Every decent marketer knows that you should never do the opposite and punch down.
- AdamProut 5y agoI would say that TPC-DS and TPC-H are really table stakes benchmarks for data warehouses at this point in time (maybe they weren't 10 years ago). How to build a database that does well on them is well documented in the literature now[1][2][3][4] (maybe a few other papers). Its not easy to build such a database, but its "just" hard work and many companies have the $$ necessary to do that work. There isn't any magic or technical moat in the results for databricks (or snowflake, or redshift, etc.). I think Databricks is overly enthusiastic about their results as they have been trying to be competitive with cloud DWs on these benchmarks for a number of years now. They have finally caught up (by building deltalake and their photon query engine which implement a number of standard DW features). [1] http://www.vldb.org/pvldb/vol13/p1206-dreseler.pdf [2] https://stratos.seas.harvard.edu/files/stratos/files/columnstoresfntdbs.pdf [3] https://web.stanford.edu/class/cs245/readings/c- store.pdf [4] http://sites.computer.org/debull/A12mar/vectorwise.pdf
- thrtlvlmidnight 5y agoI agree with everything above. The main advantage the newer data warehouses have over the legacy on-prem incumbents is that they had the chance to build from scratch having learned from all of the challenges that the original players encountered. The public pissing contest is entertaining while also being silly and slightly cringe, but I think it's a nice story for Databricks nonetheless. They now have a performant SQL-based analytics engine that can credibly compete with the best DWs in the market today, and it's just one part of their overall platform. The sense I get is that Snowflake wants the conversation to be "no matter what you do, you need a data warehouse, and we're the best in the business at that." Databricks' Lakehouse approach is a fundamental challenge to that, and if they're getting this kind of performance from their analytics engine against the market-leading data warehouses today, that's a big momentum shift in their favour.
- imslowbutnice 5y ago(X-Posted) I dont get still how much optimization was done for the Databricks version Snowflake TPC-DS power run. This is what I am seeing so far (and i am foggy on) - DB1.Databricks generated the TPC-DS datasets from TPC-DS kit before time started. Databricks starts time then generated all queries. Then Databricks loaded from CSV to Delta format (also some delta tables were partitioned delta tables by date) and also computed statistics. Then all of the queries are executed 1-99 for TPCDS 100TB SF1. Databricks generated the TPC-DS datasets from TPC-DS kit before time started. Databricks starts time then generated all queries. Then load from S3 to Snowflake tables by - (i'm not sure about these next parts) - creating external stages and then "copy into" statements I guess? Or maybe just using copy into from an s3 bucket, that part doesnt matter much. But its not clear did they also allow target tables to be partitioned/clustering keys at all? Then all of the queries are executed 1-99 for TPCDS 100TB Its just hard to say exactly what "They were not allowed to apply any optimizations that would require deep understanding of the dataset or queries (as done in the Snowflake pre-baked dataset, with additional clustering columns)" means exactly. Like what does that exactly mean. At a glance though, this looks very impressive for Databricks, but just want to be sure before I submit to an opinion. SF1. Databricks generated the TPC-DS datasets from TPC-DS kit before time started. Databricks starts time then generated all queries. Then load from S3 to Snowflake tables by - (i'm not sure about these next parts) - creating external stages and then "copy into" statements I guess? Or maybe just using copy into from an s3 bucket, that part doesnt matter much. But its not clear did they also allow target tables to be partitioned/clustering keys at all? Then all of the queries are executed 1-99 for TPCDS 100TB Its just hard to say exactly what "They were not allowed to apply any optimizations that would require deep understanding of the dataset or queries (as done in the Snowflake pre-baked dataset, with additional clustering columns)" means exactly. Like what does that exactly mean. At a glance though, this looks very impressive for Databricks, but just want to be sure before I submit to an opinion.
- rdxm 5y agoohhhhh, a good old-fashioned DB benchmark shit-fight!!! wow. it's been a while since we've had one of these!!! paging Oracle, MS SQL Server, SAP, Sybase, etc, etc.... I'm gonna make some popcorn and crack open a beer!!!
- inetknght 5y agoSnowflake accuses other companies of lacking integrity? I really wish I could block all of Snowflake's domain from my inbox. Sadly, Google encourages spammers to just create a new email address. So I get a few emails each month from Snowflake who ask me to try their products. I've never done business with them and there's no unsubscribe link. Fuck Snowflake for thinking it has any room to talk about integrity.
- doppelganger1 5y agoWhat I find comical is they accuse Databricks of lacking integrity but they don’t actually call out anything except their benchmark was faster than what Databricks did in Snowflake. Databricks then reruns the benchmark and says the only reason that Snowflake’s was faster was because of the built in dataset they used. Databricks was able to match Snowflakes numbers using it but when they loaded the actual data set, it was much slower, which is how a proper TPC benchmark is supposed to happen. They then said that Databricks blog doesn’t match the TPC results, but when I looked at them, they do match. I guess Snowflake just expects people to take arguments at face value. Then I saw someone on LinkedIn complaining that Databricks must have used some beta version. I didn’t see a beta version being used, but that kind of goes out the window when Databricks follows up and then posts that they matched Snowflake when they used their built in TPC data set. This is funny and interesting to watch but also a distraction I feel. Amazon says it best when they say, “Leaders start with the customer and work backwards. They work vigorously to earn and keep customer trust. Although leaders pay attention to competitors, they obsess over customers.”
- xiaodai 5y agoSpark compares itself to Hadoop only on the front page. I wonder how Spark compares to Firebolt.
- funstuff007 5y agoI guess if anyone suggests "sampling" the data in meeting these days, they get their head blown off.
- uvdn7 5y agoNow I see that getting rid of the DeWitt clause is indeed great. Kudos to both companies.
- boublepop 5y agoSnowflake must be kicking themselves hard now for letting a story that was “Databricks is a viable alternative” turn into “Snowflake has absolutely no integrity and will fling mud even while they are gaming the statistics” Really can’t see what they can do now short of “bending” to Databricks and entering the competition. And naturally it’s no longer just enough that they show comparable performance. They have to hit their games stats somehow otherwise any news even of they beat Databricks will be reported as “see, we told you they where cheating”