6 ms·
Redshift is based on ParAccel, not on Postgres. ParAccel uses APIs similar to Postgres due to historical reasons, but not the technology. For a basic overview:
by TY 14y ago
Redshift is based on ParAccel, not on Postgres. ParAccel uses APIs similar to Postgres due to historical reasons, but not the technology.
For a basic overview: http://en.wikipedia.org/wiki/Paraccel http://en.wikipedia.org/wiki/Paraccel
As for the rest of the article, it feels like a basic Data Warehousing 101 re-discovered. It should have been titled "Analytics: Back To The Future" :-)
- bloomfilter 14y agoThanks for pointing it out, we have correct it in our post
- meritt 14y agoNo kidding. The amount of startups that have flocked to hadoop for "data analytics" over the past 5 years is extremely disheartening. Almost all of the cases are far more suitable for any off-the-shelf RDBMS much less a column-oriented one. Same thing with MongoDB. How much time and money would have been saved learning Database Theory/SQL/Data Warehousing/Dimensional Modeling instead of cramming everything into an unstructured data-store?
- mixedbit 14y agoWhich off-the-shell RDBMS can handle queries over 3 billion rows?
- jbverschoor 14y agopostgres?
- meritt 14y agoCounter-question: Which startup has a actual data table with over 3 billion rows?
- badgar 14y agoWhen you log every mousedown because the founder misunderstands A/B testing, 3 billion rows is easy to come by. Besides - you're busy changing the world, so you should expect to use the same technology as Facebook and Google.
- PanMan 14y agoWe have just crossed 2 billion items in our datastore. While not 3 billion yet, I expect that to happen later this year. Too bad Redshift can't handle JSON files: Converting everything will be annoying.
- fujibee 14y agoOur idea is to change from JSON on loading to Redshift, continuously. http://www.hapyrus.com/pages/flydata-for-redshift http://www.hapyrus.com/pages/flydata-for-redshift
- nkohari 14y agoWe do. http://adzerk.com/ http://adzerk.com/
- taligent 14y agoSAP HANA would be one but it is basically in memory so very, vey expensive.
- mbesto 14y agoAnd it's not a RDBMS. It's basically the same technology as RedShift, but not cloud based (yet).
- res0nat0r 14y agoActually HANA One is available in the AWS Marketplace: https://aws.amazon.com/marketplace/pp/B009KA3CRY/ref=mkt_ste_hp_car https://aws.amazon.com/marketplace/pp/B009KA3CRY/ref=mkt_ste...
- jacques_chester 14y agoIn 2007 I worked for a firm with a 4 billion row join table in PostgreSQL. Might've been 7 or 8, I don't recall which. It ran on a quad core server with 16Gb of RAM. Joins going through this table took about 2-3 seconds to complete.
- mixedbit 14y agoBut I suspect the join must have been over an indexed column, so it did not touched 4bln rows, otherwise 2-3 seconds would be hard to believe. The group by query in the article must access all 3bln rows, which makes a huge difference.
- jacques_chester 14y agoAll the columns were indexed. I remember it well, because I was trying to explain why having tens of gigabytes of indexes wouldn't help them much if they only had 16Gb of RAM. In terms of group-by performance, it depends a lot on the kind of data and how it's stored. For example, taking a sum on a columnar store is quite amenable to parallel solutions and a lot of databases will do that way.
- EwanToo 14y agoI can't think of an off the shelf RDBMS which can't handle queries on 3 billion rows. SQL Server can Oracle can Postgres can Even MySQL can (!) The limitations are almost always in the hardware, not the software. If you're looking at column based systems, you can look at Greenplum (does both row and column-based storage), InfiniDB (MySQL based), and all sorts of expensive but very fast appliance options like Netezza, Teradata, etc.
- nieksand 14y agoI think part of the issue why so many people have gone with Hive is that good, production-ready column stores are expensive. Redshift is posed to change that. If you're shopping in this space, Infobright is also worth checking out. And even for moderate data sizes (10+ GB per table), row store DBs tend to become painful. This is especially true when you need to support ad-hoc reporting queries, since the usual technique of matching your schema, indexes, and queries won't be effective any more. With true ad-hoc reporting, your only hope becomes lots of shallow indices rather than ones tuned to a particular query.
- meritt 14y agoDimensional modeling (I'm a fan of Kimball's approach) mitigates these problems quite well while still offering very flexible ad-hoc reporting. Works great on a row-based RDBMS, even better on columnar. http://en.wikipedia.org/wiki/Dimensional_modeling http://en.wikipedia.org/wiki/Dimensional_modeling http://www.amazon.com/Data-Warehouse-Toolkit-Complete-Dimensional/dp/0471200247 http://www.amazon.com/Data-Warehouse-Toolkit-Complete-Dimens... Redshift is indeed a solid product but all these comparisons against Hive are surprising, as that's not the right tool in the first place. Infobright, greenplum, aster, vertica, etc are the products which Redshift seeks to disrupt.
- jacques_chester 14y agoI realised a few years ago that pretty much every database course taught only teaches OLTP. OLAP never really gets a lookin. At my university, standard normalisation was taught in the "databases" course. OLAP was mentioned as part of the "advanced databases" course. The database course at that time blew about half its time on building PHP applications to talk to the database. I hate to second guess my professors, but I can't help but feel that a more productive use of the time would have been to teach normalised OLTP in the first half, and dimensionally modelled OLAP in the second half. Better yet, to divide them into two courses and spend some time talking about database history ("here's why network and hierarchical databases sucked") and maybe some introduction to how query planners work.
- 14y ago
- fludlight 14y agoHow much time and money would have been saved by running analytics on samples instead of the whole population?
- taligent 14y agoSeriously can you and your ilk just please stop. It's so exhausting to hear how much smarter you are and if we just educated ourselves we would realise the error of our ways. People who choose the technologies aren't stupid or masochistic. They understand their use case and the fact is that there are plenty of situations where SQL is suboptimal.
- TheAnimus 14y agoI don't think he is epeen waving (where e this time is education). Sometimes with technologies going through the Gartner Hype Cycle people choose the incorrect one, because of the buzz, the glamour around it. NoSQL is most definately in vouge, quite rightly, too many people often use heavy RBDMS when they are not required. But too many people perhaps are too quick to dismiss the regular database without actually understanding it. Any suggestion to avoid hype of technology, question your use cases fully is in my mind a good suggestion.
- bufo 14y agohttp://docs.aws.amazon.com/redshift/latest/dg/c_redshift-and-postgres-sql.html http://docs.aws.amazon.com/redshift/latest/dg/c_redshift-and...