3 ms·
You can get a box with 4TB of ram on EC2 for $4/hr spot, so copy your data into /dev/shm and go hog wild. For lots of databases, most of their time is spent lo
by lsb 7y ago
You can get a box with 4TB of ram on EC2 for $4/hr spot, so copy your data into /dev/shm and go hog wild.
For lots of databases, most of their time is spent locking and copying data around, so depending on your workload you might getsignificant speedups in Pandas/Numpy if it's just you doing manipulations, and there are multicore just-in-time compilers for lots of Pandas/Numpy operations (like Numba/Dask/etc).
If you have lots of weird merging criteria and want the flexibility of SQL I'd say use a modern Postgresql with multicore selects on that 4TB box.
- dijit 7y agoHow long does it take to copy your data in? And what’s the bandwidth cost involved? People like to talk about the elasticity or compute, but startup is not free (or even cheap in most cases).
- cldellow 7y agoIf your data is in S3, my experience is that you can push ~20-40MB/core/sec on most instances. OP is probably talking about an x1e.32xlarge. According to Daniel Vassalo's S3 benchmark [1], it can do about 2.7GB/sec. So your 4TB DB might take ~30min to fetch. Bandwidth is free, you'd pay $2 for the 30 min of compute, and some fractions of pennies for the few hundred S3 requests. [1]: https://github.com/dvassallo/s3-benchmark https://github.com/dvassallo/s3-benchmark
- satanspastaroll 7y agoIt's to note that any data exported out of AWS will be billed at $0.09/GB, or $90/TB
- ramraj07 7y agoCurrently I have it loaded on redshift with as much optimization as possible, and the queries are far more analytical than end-user like (often having to self join on the same dataset). This works okay, but doesn't scale with more than a handful users at a time. I'll probably run some tests with the postgres suggestion but curious if this is still a better alternative or not