4 ms·
I do know Athena is built on top of Presto and it can pair with lots of different data sources. But somehow when talking about Athena I always end up thinking a
by fredliu 9y ago
I do know Athena is built on top of Presto and it can pair with lots of different data sources. But somehow when talking about Athena I always end up thinking about using big CSVs on S3 as data source, so what I really meant was Athena+S3 I guess.
Edit: regarding transaction though, I wonder if Athena could even be used as a "meta sharding" layer, when all of the underlying data sources support transaction, but data is too large to fit in a single instance. The advantage would be to not implement sharding logic in code. Not sure about the performance though. Just a thought.
Edit2: not sure if I read it correctly, but looks like right now Athena does not support transactions yet (http://docs.aws.amazon.com/athena/latest/ug/creating-tables.html http://docs.aws.amazon.com/athena/latest/ug/creating-tables....) although using Presto there could be a future that they do support.
- gopalv 9y ago> But somehow when talking about Athena I always end up thinking about using big CSVs on S3 as data source, so what I really meant was Athena+S3 I guess. CSV is a very unwieldy storage format for Athena + S3, since Athena charges you by data scan sizes. ORC is pretty much optimized for S3-like use-cases (of being read over HTTP), so you'll find that a columnar structure like that would be a much nicer way to store and query often. There's a pretty good run-down of the options and formats here - http://tech.marksblogg.com/billion-nyc-taxi-rides-aws-athena.html http://tech.marksblogg.com/billion-nyc-taxi-rides-aws-athena...