5 ms·
I'm deeply disappointed in Databricks as an RDBMS. As a DS/DE, there's a lot to love (not all, but a lot). The easy provision of Spark clusters. The jobs API.
by jpau 6y ago
I'm deeply disappointed in Databricks as an RDBMS.
As a DS/DE, there's a lot to love (not all, but a lot). The easy provision of Spark clusters. The jobs API. DeltaLake (mostly). Easy notebooks (please don't create a prod system from these..). And Spark itself continues to improve, albeit in an increasingly crowded field.
But I've worked closely with BigCo SQL analysts on Azure Databricks, and their experience was terrible. For example:
- You cannot browse the data structure without an active cluster
- Starting a cluster can take ~5 minutes and, since you missed that moment, you may not submit your first query until 10-15 minutes.
- The SQL error messages are often (perhaps usually?) nonsense, so you have to operate without them.
- An unfortunate amount of downtime, followed by bizarre excuses.
- It's so darn slow, relative to equivalent queries on BigQuery or Snowflake.
- Even submitting a query can take a weird amount of time.
If Databricks-as-an-RDBMS were competing against Teradata, sure, let's have a chat.
But we're in 2021, and there's just no comparing the experience of the SQL analyst on Databricks-as-an-RDBMS vs. Snowflake/BigQuery.
I'm excited for the potential of Snowflake's SnowPark (though know little about it). Calling UDFs from SQL means you can create great features for SQL analysts, provided that they can build the momentum to need it.
- nattaylor 6y agoI got excellent performance in Databricks with well partitioned Parquet and Spark 2.4. What is making the queries slow? Data scanning?
- jpau 6y agoThey use DeltaLake + Spark 3.0, and are mostly careful to partition well. Their datasets are small. Most tables are ~50GB, the odd table up to ~2TB. The clusters typically are nothing shabby for this size, defaults to ~[4-12]x32GB. The queries that I have seen are typically not written well. Think view-on-view-on-view (there's a BigCo policy against them materialising data..), and where the filter is applied in the last step. The stuff of horrors, but something I've seen in more-than-one-BigCo. But we have compared some of those same queries on BigQuery vs. Databricks, and, I don't know if BigQuery's execution optimiser is better? Or if the BigQuery storage is better organising the data? Or if BigQuery is simply throwing more resource their way?
- wavesquid 6y ago> Think view-on-view-on-view (there's a BigCo policy against them materialising data..), and where the filter is applied in the last step. The stuff of horrors, but something I've seen in more-than-one-BigCo. That sounds wonderful (really). I was contracting for a BigCo where they materialised things all the time, and they would regularly end up running queries over multiple materialisations from different points in time, which invariably means that you always get wrong answers. I very much wished to put a stop to use of any materilised views, but didn't have the buy-in to make the policy.
- simo7 6y ago> That sounds wonderful (really) Was going to say. Most of the times all it takes is to have a proper data model. For analytics I favour de-nomarlized schemas and, if necessary, nested fields. Queries are much easier to write (fewer joins), much faster, no need to incrementally materialize (sigh), fewer backfills and no messy field definitions. What you often see instead is highly-normalized data models with an un-trackable amount of materialized views (usually on top each others) and some complicated tools/solutions to try to deal with all that mess. The cost of a bad design.
- marcinzm 6y agoIn my experience Athena on AWS beats Spark by an order of magnitude in terms of performance and price. Without the annoyance of having to start/stop/run a cluster. Snowflake is even faster than that since its storage is more optimized.
- willvarfar 6y agoAnd Athena is an old fork of Presto, provided as a service; modern Presto e.g. Starburst is much faster than Athena, and cheaper if you use it much as Athena has usage fees not hosting fees.
- mrbungie 6y agoI've seen and have compared Databricks clusters to a 10-15yo Teradata cluster and no way in hell I would use Databricks. Teradata is a lot faster for interactive workloads than Databricks. PS: I agree there's no comparing on Databricks vs Snowflake/BigQuery.
- snidane 6y agoTeradata is shared-nothing architecture. Of course it will outperform Databricks and Snowflake shared-disk model as the data is colocated on the compute nodes so it doesn't have to travel anywhere. Good luck on your budget trying to scale up shared-nothing database and making it scale up and down based on workload without downtime. You can achieve significant speedup resembling shared-nothing databases by pushing the data close to the query using caching. Snowflake does it out of the box as it maintains table metadata. Databricks can do it too, but you have to be careful and it sucks.
- mrbungie 6y agoFair enough, I'm just talking from the perspective of a SQL analyst who could use both Teradata and Databricks at work. Teradata was faster with no user-facing tuning vs tuned Databricks. And if you can pay for Teradata you may as well use it.
- anonymousDan 6y agoWait what? I thought spark is designed to work on top of a shared nothing cluster? Have you any resources describing the architecture of databricks specifically?
- snidane 6y agoIn Databricks you only get cloud storage (s3) as a persistent storage for data. So the data has to move from s3 to compute nodes. Self hosted HDFS will work more like shared nothing.
- anonymousDan 6y ago
- bkandel 6y agoYes, and I would add to this the (nearly) complete lack of IDE support makes working with Spark SQL quite painful.
- agambrahma 6y agoSigma Computing (https://www.sigmacomputing.com https://www.sigmacomputing.com) might be a good fit here too. (disclaimer: plug)
- Braxton_Hicks 6y ago> You cannot browse the data structure without an active cluster That's one of the reason's I'm interested in delta-rs [1], which has delta lake bindings for Python. Would love to read a delta lake table into a native python object without the need for spark. [1] https://github.com/delta-io/delta-rs https://github.com/delta-io/delta-rs
- jgalt212 6y agobut they have $400MM of 2020 revenue. How can they do that with such a disappointing product?
- jpau 6y agoDatabricks is good as a managed Spark platform. They have thought about how they can improve the DS experience. Inconsistent storage? DeltaLake. Slow Spark queries? Databricks Delta. Model management? MLFlow (I haven't adopted this, but can't pin down why -- on face value it seems great). Development environment? Databricks Connect. Cluster management? Core. But the same is not true for SQL analysts. Today's offering does not empathise with them. I'm unsure integrating Redash is a genuine reply to their needs. The upside here is that (1) Databricks (or at least, Databricks' marketing) appears to be prioritising this need, and (2) A lot of people are betting a lot money that they can do this well. Tomorrow looks sunny.
- Jugurtha 6y ago>Model management? MLFlow (I haven't adopted this, but can't pin down why -- on face value it seems great). Probably because you want your code to be about the problem you're trying to solve, not about tracking experiments. Similar to Anti-lock braking system or Electronic stability control systems in a car: you want them to be "on" by default, not to activate them every five minutes while driving.
- bostonsre 6y agoI would guess high cost, market saturation, and not enough research and development. It seemed like an ok product when we evaluated them, but the cost was prohibitive. After the hortonworks and cloudera merger, the future of the hadoop/spark open source ecosystem seemed extremely dull and we came really close to using either databricks or one of their competitors. We found that spark on kubernetes appears to have a bright future. So, we didn't have to double our spark cluster costs by going to databricks.
- victor106 6y ago> Calling UDFs from SQL means you can create great features for SQL analysts You can do that in Spark, no?
- jpau 6y agoYou sure can :) I see it as why the article supports Databricks as an RDBMS; it offers something others do not. You can't currently* do the same extensive UDFs in Snowflake or BQ and, sometimes, they are important. But with SnowPark coming, hopefully you won't have to make such a large sacrifice to SQL users' experience for it. * Currently you can do JavaScript UDFs and external functions in Snowflake, and BigQuery ML is worth mentioning here too. Those cover some, but not all, of what you might use a Spark UDF for in SQL.
- StreamBright 6y agoAnd on the top of that if i need transactions and SQL Data Lake wouldn’t be my first choice. There are other options.