8 ms·
Amazon Announces new Data Warehousing Product
- sologoub 14y agoMaybe a naive question, but how does this compare with Google Big Query?
- rorrr 14y agohttps://developers.google.com/bigquery/docs/pricing#table https://developers.google.com/bigquery/docs/pricing#table $1474 per TB per year for storage alone ($0.12 * 1024 * 12) plus $35.84 per TB queried Amazon is definitely cheaper.
- bravura 14y agoDoes anyone have insight into how painful it is for non-technical people to query their data warehouses? I'm building a tool that allows business people and non-technical analysts to query their data warehouses using natural language. (Currently, you must ask a technical person to write ad-hoc queries for you, or build you a dashboard. This bogs down your data people.) Does anyone have insight into the demand for such a product? [edit: I'd love to chat with anyone with insight into this topic. Reach me at Joseph at metaoptimize dot com]
- ironchef 14y agoThe closest things i've seen are exploratory data visualization products such as tableau (which is pretty awesome). The downside (or partial downside) is it can end up writing some nasty non-performant queries in certain aspects.
- kanwisher 14y agoMost of the time its usually easier to have people learn a touch of SQL and ask developers for harder queries. We used Tableau and after a couple of weeks, they had every query they every wanted saved.
- tom_b 14y agoNon-technical people don't query. They call me up, ask me to do a "quick report across the inventory db with the project cost data." I send it off to them. If they like it, we push a report (maybe with a couple of parameters) into production. My gut is that we aren't lacking for good technical options in analytics and data warehousing. To be honest, the lion's share of my work in data warehousing is helping the users know what questions to ask. But there is lots of room and probably several excellent lifestyle to 8-digit businesses for good BI.
- matwood 14y agoDoes anyone have insight into how painful it is for non-technical people to query their data warehouses? Depends. Back when I did DW stuff my general workflow was to speak with the analysts about what they were trying to accomplish. From there I would create the cubes and additional metrics. I would also set up all the processing schedules at this time. The analysts would then use an Excel plugin that provided a pivot table interface to any cubes for which they had access. It worked pretty well. For straight data access I would teach the them basic sql and/or build sql templates for them that they could extend. My goal was always teach a man to fish and get out of the way.
- johnrgrace 14y agoYes, it can be crazy painful to the point that Non-technical people just don't do the query unless it is business burning critical. In large part because the technical people often are tasked on projects from IT, and to get them to do a query required middle management department to department deal making which is slow and painful. At ExxonMobil, a place I worked, you're going to have VP's asking eachother and IT is going to hedge with, yea if we do this then project X will be late (it's going to be late anyway but they've kept quite about it and no one knows). My personal solution when I needed a query was to bring a six pack of beer down to IT friday afternoon, mostly because I wouldn't be given access to write queries because we had BI software.
- meritt 14y agoI would suggest reading some books on the topic of Dimensional Modeling [1] such as "The Data Warehouse Toolkit" [2]. The critical thing you need to expose to your users is the ability to ask for things which make sense in their world that are actually really difficult for even an engineer to code. Things like: "Show me average 9am-12pm sales on Mondays, Wednesday and Fridays for 1st quarter, 2012" [1] http://en.wikipedia.org/wiki/Dimensional_modeling [2] http://www.amazon.com/Data-Warehouse-Toolkit-Complete-Dimensional/dp/0471200247
- alex_anglin 14y agoSpeaking as someone who does his fair share of dimensional modeling, I would just point out that the example you cite could only involve two tables in a well designed dimensional model (sales fact and time/date dimension, I reckon). The challenge is in getting to that point. To speak to OPs point about difficulty in querying data warehouses, most business intelligence tools that I'm aware of provide semantic layer[1]-type capabilities, whereby the user interface of the tool is presented in the language of the business domain. Nevertheless, I still agree that this is still difficult work, unfortunately. That it is getting more complicated in some respects, such as through unstructured data, doesn't help either. [1] http://en.wikipedia.org/wiki/Semantic_layer http://en.wikipedia.org/wiki/Semantic_layer
- meritt 14y agoI guess I wasn't clear enough if I came across like my example was complex. It's one easily solved via DM and one that's extremely hard to execute in most non-dimensionally-modeled setups. That's exactly why I'm a huge advocate of DM instead of just throwing a ton of servers, hadoop & MR at everything.
- shuzchen 14y agoYou might want to look into rjmetrics or chart.io and see what they offer. I've been integrating with both, and it seems one of their goals is to (after the connection and datasources are set-up - that still requires technical knowledge) allow non-technical people access to analyze the data.
- jacques_chester 14y ago> Does anyone have insight into the demand for such a product? Enormous, and there are dozens of such tools available. Most of them work best if you build an actual data warehouse -- dimensionally structured, not normalised. This is because they can easily build query forms using the DW dimensions in a language that makes sense to end users.
- baltcode 14y agoSo is this the Amazon clone of Google's Spanner?
- scott_s 14y agoNo. Spanner is a globally distributed database which supports transactions. It is meant for applications which need to make frequent updates to a database, but the storage for the database may be distributed around the world. Redshift is a different usage model. You upload your data once, then ask questions of it - but you don't update it. Google does have something similar to Redshift: BigQuery (https://cloud.google.com/products/big-query https://cloud.google.com/products/big-query).
- zrail 14y agoNow if only Amazon would offer PostgreSQL on normal RDS.
- seiji 14y agoDoesn't RDS crap itself whenever there's a core AWS problem?
- semiquaver 14y agoYes, since it relies on EBS for persistence, which has been one of the flakiest parts of AWS so far.
- ceejayoz 14y agoAll it takes is an EBS outage to determine this. In the last one, RDS and ELB had issues (and were flagged as such on their status board) due to needing EBS, but I don't believe DynamoDB was.
- wahnfrieden 14y agoNot likely to happen as long as Oracle has its licensing grip on AWS.
- jat850 14y agoAnd PostGIS too, please - if anyone from AWS is listening/reading :)
- deleted 14y ago[deleted]
- kanwisher 14y agoShould be interesting if this will be a viable competitor to column oriented sql engines like Vertica or other OLAP solutions like SAP HANA. It would be nice if there was a simple SQL based olap solution that I can spin up for offline reporting that can scale terrabytes of data
- dude_abides 14y agoThis has the potential of really disrupting the enterprise data warehouse sector. All the MPP vendors today (HP Vertica, EMC Greenplum, Teradata) have exhorbitant pricing and ridiculous licensing. With their pricing - 1000 $ per TB per year, I would be really worried if I were Teradata (Not so much if I were IBM).
- huggyface 14y agoThis is essentially ParAccel as a service. So it is one of those data warehouse vendors. As to the pricing, as always it isn't so straightforward -- it's $999 per TB a year if you pay $3,000 upfront, and then another $1000 for a single, very weak database server.
- jasondc 14y agoA lot of large enterprises won't be comfortable hosting their data outside of their own data centers. The killer application is making a portable, on premises version of this functionality without the high price.
- monstrado 14y agoCloudera is doing just that with the recent announcement / open sourcing of Impala. Based on Amazon's description of their hosted product, the technology is very similar. Impala is still in beta, and columnar storage (trevni/avro) is right around the corner...with that, you can do petabyte scale queries for a very low cost. https://github.com/cloudera/impala https://github.com/cloudera/impala
- flanger 14y agoPlatfora is doing some interesting work with interactive, in-memory BI for Hadoop. They essentially do away with the traditional DW/ETL model and create ephemeral in-memory 'lenses' for querying and visualization.
- capkutay 14y agoImpala is married to Hadoop. What if your data infrastructure isn't built on hbase and too complex/large to integrate it easily? Would impala still serve that purpose?
- rpicard 14y agoWhat is the use case for something like this versus a regular RDS service?
- amock 14y agoThis can scale up much more than a single RDS database since it spreads the data across multiple machines, but it's not exactly a replacement for MySQL database. It's also possible that this doesn't make use of EBS, which could make it perform more predictably and protect it from failure when EBS fails.
- semiquaver 14y ago> It's also possible that this doesn't make use of EBS This quote from the product page seems to indicate that EBS is not used for primary data storage: "it runs on hardware that is optimized for data warehousing, with local attached storage and 10GigE network connections between nodes."
- bvdbijl 14y agoWhy isn't it a replacement? From what I read it seems it's like a very large and transparantly scalable SQL database
- ceejayoz 14y agoRDS has a maximum capacity of a terabyte, you'd need to shard to go beyond that.
- dgreensp 14y agoThe answer is in the term "data warehousing" -- http://en.wikipedia.org/wiki/Data_warehouse http://en.wikipedia.org/wiki/Data_warehouse -- which has implication that you're going to be doing data mining on vast amounts of data, often historical data like logs or transaction histories. Google has systems like this for analyzing its request logs. Think of how many HTTP requests hit Google's front-end servers per second or hour or day. Each one has a few dozen pieces of data associated with it -- URL, client IP, headers, etc. Suppose I want to make a bar chart of how many requests came from France containing a certain header, each day for the last year. The system can do this query quickly if the requests are already bucketed by time interval, organized by column, compressed, and stored so that exactly the information needed can be brought into RAM quickly. It is a little funny, when you step back, that "storing," "archiving," and "warehousing" are different things and Amazon has services for each. Try explaining the difference between S3, RDS, EBS, Glacier, and Redshift to a layperson.
- monstrado 14y agoI'm curious what technology they are using to power it. According to the website, the technology described seems very similar to what Cloudera recently open sourced (Impala), which sits along side Hadoop allowing ad-hoc MPP style querying on petabytes of data. https://github.com/cloudera/impala https://github.com/cloudera/impala
- jeremyjh 14y agoI'm guessing it is quite a bit different from that. It is a relational data warehouse. It supports a Postgres protocol and API, which sounds more like what Netezza has built. In fact, I would expect Netezza to be one of the most likely companies to partner with Amazon at this kind of price-point.
- javery 14y agoExcept Netezza is now owned by IBM.
- lazyjones 14y agoOther candidates: * Yahoo's Everest * Greenplum * Aster Data All mentioned here: http://www.cubrid.org/blog/dev-platform/database-technology-for-large-scale-data/ http://www.cubrid.org/blog/dev-platform/database-technology-... The Register wrote that Amazon's solution is a column-oriented database possibly based on Postgres, like Yahoo's: http://www.theregister.co.uk/2012/11/28/amazon_aws_redshift_data_warehousing/ http://www.theregister.co.uk/2012/11/28/amazon_aws_redshift_...
- 23david 14y agoHave to say that this is pretty amazing. The price is so low that it's a no-brainer to just give it a try. For the same 2TB capability, a Vertica license would run between $20-40K, with high annual subscription fees. The bigger question for me is why Amazon has been able to figure out the technical details necessary to run this kind of service for this price. It's just ridiculous. Talk about taking the oxygen out of the market...
- aptwebapps 14y agoThat seems to be their general strategy. Matthew Yglesias has posted about it several times. Here's one such: http://www.slate.com/blogs/moneybox/2012/10/26/amazon_profits_they_don_t_exist_but_the_company_keeps_on_keeping_on.html http://www.slate.com/blogs/moneybox/2012/10/26/amazon_profit...
- perlgeek 14y ago> The bigger question for me is why Amazon has been able to figure out the technical details necessary to run this kind of service for this price. I guess they grew the infrastructure for themselves, optimizing it bit by bit over the years. And then noticed that it could be sold too.
- justincormack 14y agoThey seem to have taken their own business requirements for amazon.com and reimplemented them on commodity hardware.
- deleted 14y ago[deleted]
- jbellis 14y ago"Amazon Redshift includes technology components licensed from ParAccel." http://finance.yahoo.com/news/amazon-services-announces-amazon-redshift-174300203.html http://finance.yahoo.com/news/amazon-services-announces-amaz...
- kzahel 14y agoIt seems that the price (~$1 / GB / year) in the best case (3 year reserved) is comparable to S3 at its lowest tiers (~$0.1 / GB / month)
- pierrend 14y agoIt's "Price per TB per Year" not GB.
- mgl 14y agoLooks impressive and very interesting, signed up to review and compare with Teradata/Netezza. Can we run more complex in-database processes implemented as stored procedures on this platform or is it going to be limited to pure SQL querying/analytics? And does anyone have an idea how to upload 1 TB of data to this service using Internet connection from your in-house company server? ;)
- maineldc 14y agoAWS has pretty good support for taking external drives and importing them to S3 which could then be used with this service: http://aws.amazon.com/importexport/ http://aws.amazon.com/importexport/ I am assuming that you have 1TB to start, not generating 1TB per day which obviously changes the equation.
- K2h 14y agoIt's called Redshift! wow.. I just finished reading the sci-fi book a few weeks ago - "Redshift Rendezvous" by John E Stith. I wonder if this is where the name comes from? In the book Redshift is the name of the space ship that runs cargo mission through folded space, the obvious problem that since you are traveling within just a few m/s of the speed of light just walking on the ship while underway causes color shift - thus redshift. I read that Stith has a physic degree and worked as an Engineer for NORAD Cheyenne mountain. That made me really interested in what novel he would come up with. http://www.neverend.com/short-bio-john-e-stith http://www.neverend.com/short-bio-john-e-stith
- jmoiron 14y agoRedshift is a real physical phenomena describing the way light wavelengths get "shifted" (stretched, to visualize) towards the red as they are seen coming from something moving away from the observer: http://en.wikipedia.org/wiki/Redshift http://en.wikipedia.org/wiki/Redshift http://en.wikipedia.org/wiki/Hubbles_law http://en.wikipedia.org/wiki/Hubbles_law
- hntldr_com 14y agosummary: Amazon Redshift is a fast and powerful, fully managed, petabyte-scale data warehouse service in the cloud. Amazon Redshift offers you fast query performance when analyzing virtually any size data set using the same SQL-based tools and business intelligence applications you use today. With a few clicks in the AWS Management Console, you can launch a Redshift cluster, starting with a few hundred gigabytes of data and scaling to a petabyte or more, for under $1,000 per terabyte per year.
- alexatkeplar 14y agoThis looks awesome - we'll definitely be plugging SnowPlow into this.
- polskibus 14y agoI cannot find information on whether Redshift supports queries in MDX. Lots of DWs today are run on Microsoft SQL Server Analysis Services and its MDX spec is now supported by several DW vendors. MDX support would mean it would be easy to switch the DW engine and leave your visualisation suite (or Excel, what the hell) and make it for an easy switch to the cloud - you'd just pick a different data source in your tool.
- deleted 14y ago[deleted]
- 23david 14y agoUpdate! The entire keynote is now available on youtube: http://www.youtube.com/watch?v=8FJ5DBLSFe4 http://www.youtube.com/watch?v=8FJ5DBLSFe4 The discussion about Amazon Redshift begins at 52:50 http://www.youtube.com/watch?feature=player_detailpage&v=8FJ5DBLSFe4#t=3175s http://www.youtube.com/watch?feature=player_detailpage&v...
- 23david 14y agoVery cool that this will support regular sql queries and queries can be sent using postgresql drivers. Postgresql drivers are super stable and supported everywhere. Driver support is usually overlooked with 'Enterprise' Data Warehousing solutions. I recall that it was really hard to get the Vertica drivers installed and stable under Linux. I took a few screenshots from the keynote and included one showing the mention of Postgresql and ODBC/JDBC support. Included here if you want to see for yourself: http://wp.me/p2sRpx-1e http://wp.me/p2sRpx-1e