3 ms·
No, but you also won't be using all that data at the same time. On Hadoop (and Spark on HDFS) clusters you'll find that most disks are either not that big or a
by Kirth 10y ago
No, but you also won't be using all that data at the same time. On Hadoop (and Spark on HDFS) clusters you'll find that most disks are either not that big or are heavily under-used.
Our HBase cluster had 6 * 2Tb disks: about 8.5Tb of usable storage (the other 3.5Tb accounts for data replication/duplication) per host in the cluster. However, you need about 200 bytes in memory per kb on disk and should assign only 32Gb of heap to HBase. That's 2.5Tb wasted, per host. Couldn't just plug those disks out and use them somewhere else: you need all the disks in parallel to overcome the IO/bandwith bottleneck.