7 ms·
When you suddenly realize that your "big" data is not really that big!. Who needs a Hadoop/Spark cluster when you can run one of these bad boys
by scaleout1 10y ago
When you suddenly realize that your "big" data is not really that big!. Who needs a Hadoop/Spark cluster when you can run one of these bad boys
- tracker1 10y agoThat was kind of my thought as well... I worked on a small-mid sized classifieds site (about 10-12 unique visitors a month on average) and even then the core dataset was about 8-10GB, with some log-like data hitting around 4-5GB/month. This is freakishly huge. I don't know enough about different platforms to even digest how well you can even utilize that much memory. Though it would be a first to genuinely have way more hardware than you'll likely ever need for something. IIRC, the images for the site were closer to 7-8TB, but I don't know how typical that is for other types of sites, and caching every image on the site in memory is pretty impractical... just the same... damn.
- samstave 10y agoHeh, but I wonder what the default per account limits are on launching these... prolly (1) per account.
- cwyers 10y agoWhy would they put any kind of a limit on it?
- samstave 10y agoAll AWS accounts have a "limits" which have the default limits as to how many instances that you could launch in that region. The reason is so if you fuck up a scaling script for example you can't launch 1000 machines and take all the capacity and then bitch that you won't pay for it. It's a stop gap. However, aside from the hard limit of 100 S3 buckets, all other limits are configurable at the request of your AWS rep
- geofft 10y agoIt looks like that hard limit became a soft limit in August: https://aws.amazon.com/about-aws/whats-new/2015/08/amazon-s3-introduces-new-usability-enhancements/ https://aws.amazon.com/about-aws/whats-new/2015/08/amazon-s3...
- slaman 10y agobecause they can only put these into racks so fast
- nucleardog 10y agoAnd to prevent a run-away script from suddenly spooling up thirty of them. Besides issues with their hardware capacity, they're generally pretty good about refunding mistakes like that, so they're eating the cost...