10 ms·
Amazon Athena: Query S3 Using SQL
- jakozaur 10y agoLooks very similar to Google Big Query. Even the pricing is same: $5 / TB of data scanned.
- estefan 10y agoWhen I tried it it was slower than bigquery. Plus you've got to mess about creating hive schemas.
- guywithabike 10y agoTFA states: "Amazon Athena uses Presto with ANSI SQL support and works with a variety of standard data formats, including CSV, JSON, ORC, and Parquet."
- bsg75 10y ago"Amazon Athena uses Apache Hive DDL to define tables."
- jackmaney 10y ago> Q: What data formats does Amazon Athena support? > Amazon Athena supports a wide variety of data formats like CSV, TSV, JSON, or Textfiles and also supports open source columnar formats such as Apache ORC and Apache Parquet. Athena also supports compressed data in Snappy, Zlib, and GZIP formats. By compressing, partitioning, and using columnar formats you can improve performance and reduce your costs. https://aws.amazon.com/athena/faqs/ https://aws.amazon.com/athena/faqs/
- spullara 10y agoI don't know why you are getting downvoted. For all those data formats you have to painstakingly make table schemas for them before you can query them. Not like Snowflake or BigQuery. One of the biggest strikes against Presto IMHO.
- bsg75 10y agoApache Drill might have been a better basis if they wanted to build a "query everything easily" based on an existing project.
- ktamura 10y agoIt's not Presto per se, but running any data processing workload against unoptimized data formats is the issue. Then again, both BigQuey and Snowflake require that you move data into their storage engine (Redshift too), and that's an additional step that's proportional to the size and complexity of your data. At the same time, it's stupid to store your logs as OLAP optimized formats and completely lose legibility. In sum, Athena trades off performance for convenience. No matter what database vendors say, you can't defy the principles of computer science.
- bsg75 10y agoYou don't replace them with an OLAP format, you can pair them with an OLAP engine to aggregate, filter, or analyze. Elastic Search and Splunk are one approach, SQL query engines are another. Apache Drill is a schema discovery on read approach that can handle some of this. Its not perfect, but it does simplify some of the process where its capabilities fit the task at hand.
- fhoffa 10y agoNote that BigQuery has been able to read files straight from GCS, Drive, and even Google Spreadsheets for a while: https://cloud.google.com/bigquery/federated-data-sources https://cloud.google.com/bigquery/federated-data-sources (I'm Felipe Hoffa and I work for Google https://twitter.com/felipehoffa https://twitter.com/felipehoffa)
- bsg75 10y ago"Amazon Athena uses Presto with ANSI SQL support and works with a variety of standard data formats, including CSV, JSON, ORC, and Parquet." I wonder if this is essentially a Presto SaaS product?
- maslam 10y agoYes
- neximo64 10y agoAny examples of queries and what this can do? S3 was file storage as far as i thought?
- raghavsethi 10y agoAthena (Presto) supports standard ANSI SQL - you can query data that's stored in S3.
- neximo64 10y agoHow does that work though, so say my bucket has 10,000 json files in it and I want all of them with the name attributes being like '%john'. Is that possible?
- spullara 10y agoIt looks really interesting but I'm surprised they launched it with the create table flow broken. The query you see here was generated by their wizard... https://www.dropbox.com/s/s4cw5x7yyrdl3ch/Screenshot%202016-11-30%2009.56.04.png?dl=0 https://www.dropbox.com/s/s4cw5x7yyrdl3ch/Screenshot%202016-...
- danso 10y agoThis looks very neat. I'm someone who deals with a lot of plaintext data from a variety of sources, and so I find using ack/grep and csvkit to be efficient enough for my purposes of exploration. I love using SQL and SQLite but rarely do it for "fun" -- that is, I'll use it when I've committed to building a project, but not for exploration. This seems like it could lighten the friction quite a bit. If anyone from AWS is here: how is this used internally at Amazon?
- ktamura 10y agoThe real question to ask is, will Amazon contribute back to open source? Presto itself is plenty proven and scalable: after all, it was created at Facebook.
- mrwnmonm 10y agoJohn Forstrom: Amazon Athena - welcome to 2010! https://twitter.com/jforstrom/status/804007642246938624 https://twitter.com/jforstrom/status/804007642246938624
- nulagrithom 10y agoWondering if I could use this like SQLite for Lambdas. I'd like to build some serverless apps, but the commitment to a monthly fee from DynamoDB puts me off. Could I use Athena to drive down my cost to zero as long as the app is unused?
- brilliantcode 10y agoDynamoDB is like $5 or $10 bucks a month? but I understand the need to keep it to a minimum. Athena is really interesting and if it can be as it is advertised "Serverless SQL" then they've got a killer product in the pipes: A future where developers no longer need to spend time on scaling, configuring, maintaining, strategizing deployments but upload code and instantly begin reaping the benefits of serverless tech. The only missing component that would be a killer feature is something that answers to Azure's Active Directory. It would be nice if we had serverless plug-and-play user authentication and access control that integrated with Lambda and Athena. I'd imagine some sort of "RoR on Serverless" type of framework that will scaffold out CRUD, User Management & REST Api is going to be in the works as well. The only potential downside I see at the moment for Serverless is the uncertainty surrounding cold boots, it will directly affect user experience. It's fine when you got enough traffic to keep things in the "warm" state but there needs to be no dead zone when the call to the API Gateway is taking many seconds waiting for Lambda function to fire.
- asteadman 10y agoRe: users auth. Isn't that what Cognito is supposed to be? I mean, I don't fully understand it, but I think so. As for the cold boot issue, I thought the standing solution was to have a "fast-exit" ping-like code-path within the lambda. Query it on a regular basis (you can even do it with a lambda scheduled-event). That way your lambda should be kept warm.
- brilliantcode 10y agoTIL Cognito! That completely flew under my radar, not sure why I didn't see it before (oh that's right I was heads down in Azure). With Athena the circle is complete for me. That fast exit ping thing is pretty cool, any more information regarding that? Your comment is probably the most valuable one I came across to date since signing up, I wish there was a way to award a gold star like on reddit :D There's very little objection at this point in moving to a Serverless architecture = Athena (SQL) + Lambda (CPU) + Cognito (User).
- buremba 10y agoI hope that Amazon contributes back to the Presto community.
- intrasight 10y agowhat does "point your data in S3" mean?
- justinsaccount 10y agoAre you talking about this? you left out a word. Simply point to your data in Amazon S3
- intrasight 10y agoStill makes no sense. Please explain if you understand.
- asteadman 10y agoTo me the obvious use case is querying your log files as stored on s3. Query for a specific combination of features, or do some (simple) processing on them. It's really only useful for a small list of file formats. Doesn't really do much for you if you primarily use s3 for binary data or static web hosting.
- nimrody 10y agoWould be useful if AVRO files were supported. This was the data can also be imported into Redshift if needed (Redshift does support Avro). Other formats are schema-less (JSON,CSV, etc.) or not supported by Redshift (ORC, Parquet). Perhaps less efficient for some queries (AVRO is not a columnar format) but still useful.
- balls187 10y agoTried it twice, and it crashed big time.
- balls187 10y agoAlso gives me a 500 on US-WEST-2
- cdevs 10y ago$5 a terabyte jeebus...don't f that query up
- bsg75 10y agoJust like with BigQuery, a carefully thought out partitioning scheme is critical, or your queries need to be carefully locked down to prevent excessive table scanning. I burned through my BigQuery trial credit fast, by not using partitions during a quick-and-dirty test.
- asafm 10y agoI wonder why they haven't chose Apache Drill over Presto. Anyone knows?
- nodesocket 10y agoAnybody have an example of storing NGINX access logs and using Athena to search them?
- dhananjayc 10y agoIs it possible to connect Athena to existing Hive Metastore?