3 ms·
S3 has a number of challenges of its own, for example, file access is quite expensive (performance-wise). The Netflix engineering team has a ton of experience r
by squarecog 11y ago
S3 has a number of challenges of its own, for example, file access is quite expensive (performance-wise). The Netflix engineering team has a ton of experience running data processing "in the cloud", and have their share of war stories.
The scaling challenges the Twitter blog post addressed (disclaimer: I work at Twitter, was one of the first Hadoop people at the company, am now doing slightly different things) happen at fairly extreme scale ranges. The NN design leaves much to be desired, but it does work just fine in the vast majority of cases. The scaling challenges we are talking about involve thousands of nodes and hundreds of petabytes of data. This is not what one normally designs for, even in a "big data" system. Take this into consideration when exploring alternatives. Do you have a solid way of managing a few hundred petabytes and a few hundred million objects in S3? Does GlusterFS work at that scale?