4 ms·
Not sure if it's related, and AWS Athena is not a Database in a strict sense (or in any sense at all since it's not a transactional DB), but it seems to offer b
by fredliu 9y ago
Not sure if it's related, and AWS Athena is not a Database in a strict sense (or in any sense at all since it's not a transactional DB), but it seems to offer both scale and lower cost if it fits your requirements (read only, DW like traffic)? Anybody has experience using it in prod as a DW replacement?
- dswalter 9y agoI would argue this question is not really related, but we use PrestoDB (the technology AWS Athena is built using), and it is particularly great for when data becomes too large to fit in a "regular" database. PrestoDB and Apache Impala are the two major open-source surprisingly-fast-distributed-SQL-query-planning-engines. If you pair either of them with a columnar data storage format like Parquet or ORC, you can make analytical workloads work on inconveniently-sized data. AWS Athena abstracts the cluster management for you and runs PrestoDB on spare EC2 instances under the hood.
- fredliu 9y agoAs mentioned in the other comment, I was actually only thinking about use cases like Athenta + csv_on_s3. and that's the really "low cost" approach I was referring to for analytical data.
- electrum 9y agoThe term database or RDBMS here seems to be mean runs SQL or processes [something close to] relational algebra. Based on that definition, Athena is definitely a database. Athena is built on Presto, which has full support for transactions. Each connector provides varying support (isolation level, multi-statement writes, etc.) for transactions, limited by the underlying data source.
- fredliu 9y agoI do know Athena is built on top of Presto and it can pair with lots of different data sources. But somehow when talking about Athena I always end up thinking about using big CSVs on S3 as data source, so what I really meant was Athena+S3 I guess. Edit: regarding transaction though, I wonder if Athena could even be used as a "meta sharding" layer, when all of the underlying data sources support transaction, but data is too large to fit in a single instance. The advantage would be to not implement sharding logic in code. Not sure about the performance though. Just a thought. Edit2: not sure if I read it correctly, but looks like right now Athena does not support transactions yet (http://docs.aws.amazon.com/athena/latest/ug/creating-tables.html http://docs.aws.amazon.com/athena/latest/ug/creating-tables....) although using Presto there could be a future that they do support.
- gopalv 9y ago> But somehow when talking about Athena I always end up thinking about using big CSVs on S3 as data source, so what I really meant was Athena+S3 I guess. CSV is a very unwieldy storage format for Athena + S3, since Athena charges you by data scan sizes. ORC is pretty much optimized for S3-like use-cases (of being read over HTTP), so you'll find that a columnar structure like that would be a much nicer way to store and query often. There's a pretty good run-down of the options and formats here - http://tech.marksblogg.com/billion-nyc-taxi-rides-aws-athena.html http://tech.marksblogg.com/billion-nyc-taxi-rides-aws-athena...