4 ms·
Interesting that usually GCP is one upping AWS on some metric but in this case it isn't touching the current largest compute/memory instance on EC2, the X1 fami
by djb_hackernews 9y ago
Interesting that usually GCP is one upping AWS on some metric but in this case it isn't touching the current largest compute/memory instance on EC2, the X1 family with up to 128 vCPUs and 4TB of memory. Though the blog does allude to them testing such types in a closed beta, it still is a game of catch up still.
- RKearney 9y agoIt’s also unsurprising that this thread isn’t flooded with “disclosure: I work at google” posts. When GCP announces a feature that’s 1% better than AWS their employees flood the HN post, but when this story hits no one shows up.
- peatmoss 9y agoYikes, I'm a little out of the loop. I didn't realize you could get 4TB of RAM in a single machine on EC2. I've been seeing that medium data keeps getting bigger (i.e. the features of traditional RDBMS are eating away at the need for specialized / distributed stores for data analysis). But so too does it appear that small data is getting a lot bigger too—just load that dataset into memory for analysis. 4TB of memory allows for pretty big "small data." "I remember back when we used to do gradient descent to estimate linear models; back in the long ago when we didn't have 900 exabytes of memory attached to our NVidia Matrix Crusher 9000 linear algebra accelerator unit."
- sqldba 9y agoIs it possible that data isn't getting bigger - but that the people who work with it just want to process larger data sets than before? I mean before they'd train a model of 1,000 inputs and then test it against another 50 and call it a day. Now they want to train it against 1,000,000 inputs. Am I completely off base? It's not my area, though I work with databases, my observation is that developers always want to use the most data possible even when it doesn't really provide any benefit.
- adwf 9y agoThat is my experience recently. Developers storing 500GB on a database (pre-launch), with < 1GB of meaningful data. A bunch of json logs that they knew data science would want eventually, but couldn't be bothered to either pare down or put in a more sensible place. The thing is, it didn't really matter; Postgres still had a ton of performance left over even after the product went live. If you can still fit it in RAM, why waste $$$ of dev time over the $$ cost of a bigger instance.
- peatmoss 9y agoSorry, I was being a bit playful with language. What I mean is that, if you roughly define small, medium, and large data in terms of the strategies required to process, then the absolute size of the data that can be processed using simpler methods grows. And whether or not more data is needed or collectible varies by discipline. Astrophysics collects way more data than they used to because 1. they need it. 2. instrumentation allows it. Some kinds of data collection hasn't scaled up however. Surveying humans is expensive and labor intensive. And for many things that you might want to study about humans, you can't simply afix a sensor to them. So, what might have been only accomplished through big data, or medium data methods a few years ago can now be loaded into memory (i.e. small data strategies).