4 ms·
Isn't a DW just part of the overall equation? These days I assume you'd want to do a lot more like train ML models as part of your pipeline, things that Spark/S
by curiousDog 8y ago
Isn't a DW just part of the overall equation? These days I assume you'd want to do a lot more like train ML models as part of your pipeline, things that Spark/Scala allow you to do that'd be harder than just SQL. I think most customers use both.
- indogooner 8y agoYes but BigQuery provides read/write libraries. So you can use Spark or Apache Beam to read the data into your cluster and then process it using your framework of choice - XGBoost/SparkML etc. The predictions can then be written back to BigQuery using those same libraries.
- curiousDog 8y agoSure but that would mean copying the data in/out. Instead you could get data-compute co-location with a proper HDFS cluster. But I think BigQuery launched SQL ML APIs recently to train simple distributed ML models on the fly.
- azurezyq 8y agoActually running and scaling a hdfs cluster is no easy task and quite expensive than cloud based object storage options. Just list a few: namenode HA, unbalanced r/w (a lot of old data but less used, wasting cpus for just making the machines up), cross dc replication. You may be able to get some of them from cloudera, but would that be cheap?
- smueller1234 8y agoA "proper" HDFS cluster also wastes incredible amounts of disk with it's naive redundancy. It's a bear to maintain (see sibling comment) and the disaster recovery story (at a DC level) is non-existent. If you invest a chunk of that money saved by not having to do 3 or 4 way data replication into good network, a lot of that move-compute-to-data complexity becomes unnecessary. The second big efficiency promise of the basic map reduce is doing lots of sequential I/O on those pesky disks which can do 100x sequential throughput compared to their pithy random read performance. Alas, in today's mixed workloads, you're certainly not getting those >100MB/s from a disk that might make a benchmark scream. So the network savings (you're still going to do that shuffle before your reduce anyway, aren't you?) from compute-next-to-storage becomes less useful. Best I can tell, the compute-next-to-storage trick does still matter as you get to large data sets (PB in a job), but then as you keep growing your infrastructure (and I'd wager Uber would be just about at this point), the benefits of disaggregation start weighing increasingly heavily. In particular, fleet management becomes considerably easier with less resource stranding.