3 ms·
> A lot of BigQuery users would be surprised to find they don't need BigQuery. No they wouldn't. a) BigQuery is the only managed, supported solution on GCP fo
by threeseed 1y ago
> A lot of BigQuery users would be surprised to find they don't need BigQuery.
No they wouldn't.
a) BigQuery is the only managed, supported solution on GCP for SQL based analytical workloads. And they are using it because they started with GCP and then chose BigQuery.
b) I have supported hundreds of Data Scientists over the years using Spark and it is nothing like BigQuery. You need to have much more awareness of how it all fits together because it is sitting on a JVM that when exposed to memory pressure will do a full GC and kill the executor. When this happens at best your workload gets significantly slower and at worst your job fails.
- winwang 1y agoHopefully, we can be another managed solution for those on GCP. And as for your second point, yep, Spark tuning is definitely annoying! BigQuery is a lot more than jusr the engine, and building a simple interface for a complicated, high-performance process is hard. That's a big reason why I made ParaQuery.
- threeseed 1y agoYou may want to look into DataMechanics who is another YC startup who tried something similar. They were acqui-hired by NetApp. If I remember they focused on SME space because in enterprise you will likely struggle against pre-allocated cloud spend budgets which lock companies into just using GCP services. I've worked at a dozen enterprise companies now and every one had this.
- winwang 1y agoEnterprises can deploy on their own GCP, and we're planning on releasing on GCP Marketplace. For a similar cost, what if their pipeline were 5x faster, and they don't have to dealing with managing the deployment themselves? Thanks for telling me about DataMechanics
- threeseed 1y agoa) Enterprises have almost entirely moved away from self-hosting software. GCP Marketplace is fine but I would probably also look at a Kubernetes option as many companies have GKE clusters. b) It won't be 5x faster though and I wrongly recommend you don't take a marketing attitude when selling this type of software. Because it will be mostly technical engineers and architects deciding on this and we aren't stupid. I have run GPU accelerated Spark clusters for years for enterprise companies and you will be able to accelerate the query part of the pipeline but that's like 20% of what a typical job does.
- winwang 1y agoa) Since being fully-managed is one of my value props, that's probably better for us. b) Of course I'm only accelerating the Spark/query part. Not sure what you mean. And in that case, I took a query which was 44 minutes on BigQuery and ran it with a "comparable" cluster on ParaQuery in 5.5 minutes. Perf is slightly variable, so maybe it's 40 minutes vs 6 minutes. In that case, ParaQuery would still be 6.5x faster, and >2x cheaper. That being said, it was just a benchmark ETL query with some random data (50b rows), and these things do vary between workloads. So yeah, without knowing more about the use case you're talking about, hard to say. Even Nvidia has a hard time optimizing certain TPS-DS queries btw, so it's not like I can just 5x anything!
- mritchie712 1y ago> No they wouldn't. haha, you're giving people way too much credit. Tons of people make bad software purchasing decisions. It's hard, people make mistakes.