3 ms·
Sure but that would mean copying the data in/out. Instead you could get data-compute co-location with a proper HDFS cluster. But I think BigQuery launched SQL M
by curiousDog 8y ago
Sure but that would mean copying the data in/out. Instead you could get data-compute co-location with a proper HDFS cluster. But I think BigQuery launched SQL ML APIs recently to train simple distributed ML models on the fly.
- azurezyq 8y agoActually running and scaling a hdfs cluster is no easy task and quite expensive than cloud based object storage options. Just list a few: namenode HA, unbalanced r/w (a lot of old data but less used, wasting cpus for just making the machines up), cross dc replication. You may be able to get some of them from cloudera, but would that be cheap?
- smueller1234 8y agoA "proper" HDFS cluster also wastes incredible amounts of disk with it's naive redundancy. It's a bear to maintain (see sibling comment) and the disaster recovery story (at a DC level) is non-existent. If you invest a chunk of that money saved by not having to do 3 or 4 way data replication into good network, a lot of that move-compute-to-data complexity becomes unnecessary. The second big efficiency promise of the basic map reduce is doing lots of sequential I/O on those pesky disks which can do 100x sequential throughput compared to their pithy random read performance. Alas, in today's mixed workloads, you're certainly not getting those >100MB/s from a disk that might make a benchmark scream. So the network savings (you're still going to do that shuffle before your reduce anyway, aren't you?) from compute-next-to-storage becomes less useful. Best I can tell, the compute-next-to-storage trick does still matter as you get to large data sets (PB in a job), but then as you keep growing your infrastructure (and I'd wager Uber would be just about at this point), the benefits of disaggregation start weighing increasingly heavily. In particular, fleet management becomes considerably easier with less resource stranding.