3 ms·
Ah, so the thing I was attacking was the "if it fits in RAM, it isn't big data" meme. AWS is happy to sell you on the idea that you can use S3 as the persistent
by zten 6y ago
Ah, so the thing I was attacking was the "if it fits in RAM, it isn't big data" meme. AWS is happy to sell you on the idea that you can use S3 as the persistent storage and then bring up the compute whenever. If you bring up the compute on-demand, then the network link matters, whether it's EBS spoon-feeding you data or you using multiple VMs and the right object storage strategy in S3 to suck the data out as fast as possible.
You pay a premium to have this stuff loaded up on faster hardware. If you look at how Redshift works on their compute-optimized nodes, everything's sitting in local NVMe SSDs (a slightly different strategy is used on ra3 nodes). They handle the fact that this is ephemeral with automatic backups. Cluster restores aren't exactly fast. I actually agree with you; if you've got some monster SSDs attached, which are comparatively cheap, why focus on the RAM... I believe there are reasons to do that sometimes, but not everything demands quite that level of performance.
For this point:
>> You _could_ write something on your own that just forks out a bunch of threads in a single process to rip through the data, but why?
> Because it's simple and easy. I wrote one in a few days. Not much code.
I think it depends on what formats you're using and how it's laid out on disk. A lot of people reshape their data into a table-like structured or semi-structured format, and that makes it a candidate for putting it into a database like Postgres or Redshift, and other times it makes more sense to bring the database to the data (the Hadoop ecosystem of stuff.)
For example my company still has stuff that writes out row-like objects into S3, and it's not too hard to write a single process job that spawns threads, reads the input files, and does some computation on those. They're on S3, so throughput kinda sucks, and copying it to local SSD only makes sense if you want to make multiple passes. But the Parquet format alternative of this data is tremendously faster to work with and genuinely easier to use, and the only barrier to entry is that you run Spark on your local machine and commit to using Spark with Elastic MapReduce to do processing for this data. Sure, you might only end up filtering through tens or hundreds of gigabytes; terabytes is usually rare. But part of that is because you only read 10-20% of every input file to do the work. It also integrates extremely well with other stuff we're using - it's even easier to just use Snowflake against the same data set, for example.
Sorry, I think I'm rambling at this point and not presenting a really coherent argument.
- throwaway_pdp09 6y agoI'd pretty much define big data as quantity, so yes, if it fits in ram it ain't. But that's just my definition (though reasonable, surely). You have immediately confounded 'big data' with a ton of cloud tech. My point is that you arguably don't need big data frameworks, and therefore arguably sticking it in the cloud is pointless too. If you can buy a reasonable server and stick it in the corner of the office, do so (noise & security, yeah). > Sure, you might only end up filtering through tens or hundreds of gigabytes ...S3 ...Parquet ...Spark ...Elastic MapReduce ...Snowflake Dude, do you need all this stuff!? Really? For a poxy few hundred GB?