7 ms·
Why We Chose Redshift
- deeviant 12y agoRedshift is like a prison, but with excellent accommodations. It's a great platform but it pretty much the perfect example of vendor lock-in.
- exelius 12y agoTo be fair, that kind of comes with the territory when talking about data warehousing. The data volumes are so large that migrating them is usually out of the question, and query languages vary between vendors pretty significantly.
- paladin314159 12y agoCompared to the lock-in of the AWS ecosystem in general, Redshift honestly isn't that bad. You can unload all of your data into S3 and then do whatever you want with it. I'd be surprised if most data warehousing solutions had such an easy way of exporting the data.
- vosper 12y agoIn addition, if you store your data in S3 and have Redshift load it from there then you don't even need to do an export - just leave your source data in S3 after Redshift's loaded it, and you're all ready to switch to another platform.
- not_kurt_godel 12y agoCan you explain what you mean by that? I fail to see how a PostgreSQL query interface could possibly qualify as a perfect example of vendor lock-in.
- exelius 12y agoIf you want to move to another DW platform, it's probably not going to be Postgres-based. As every vendor has a slightly different flavor of SQL with different behaviors, this will require redesigning your queries, schemas, and most if not all of your stored procedures. Depending on the company and age of the platform, this could be many thousands of hours of work. Really, vendor lock-in is pretty much a given with data warehousing platforms. Though these days, it's not uncommon for large companies to have multiple DW platforms all pulling data from each other. When one platform falls out of favor, the users just migrate themselves to another since most reporting systems not made by SAP or Oracle are compatible with pretty much everything.
- bsg75 12y agoIn contrast, Vertica, Greenplum, Netezza, Teradata Aster, and CitusDB are all based on PostgreSQL forks. In many cases, the client libraries behave like psql, and ease conversions at that level. As to SQL language differences, no DW platform uses "standard SQL", just as no two RDBMS use the exact same SQL dialect. I dare say database platform lock-in is a universal issue. Any migration will involve effort.
- eva1984 12y agoHow is Redshift a vendor-lock in though? Put your data in S3, in csv/tsv/json format, if you want to switch to other provider, just figure out how to import it, and your are all set. How to figure out the limitation of the different platforms and tuning and optimizing is the difficult part. Data migration is almost always painful and time-spending. When choosing your data provider, you have to be careful because it is very likely to be a long-term commitment. In that sense, in DW world there is always vendor lock-in. Only it is largely driven by the essence of the application itself, less so by the intention of the provider.
- otterley 12y ago"Resort" is probably a better analogy than "prison." Most people wouldn't choose to leave, since the accommodations are so nice, but for the expense.
- fsaintjacques 12y agoThat extra order of magnitude you pay in pricing you gain in response time.
- exelius 12y agoYeah, but a data warehouse isn't supposed to have great response times. Data warehouses are for large, low-value sets of historical data that you don't always know how you want to use. If you want to use data in real-time, you should be driving it from your transactional systems. Redshift and other data warehouse solutions are for doing reporting and dashboards, not triggering real-time reactions.
- luckydata 12y agoWell, used to be true, but now those systems are converging. -- Full disclosure, I work for a company working on exactly that problem called Treasure Data.
- exelius 12y agoMost companies are generally more concerned about reducing their data warehouse costs than they are about improving the performance of their data warehouses. Many companies implement a multi-tiered DW structure to get a mix of the two, but the core driver is managing the cost of storing petabytes of data while keeping performance acceptable.
- ernestipark 12y agoCurious if your funnels are just queries directly in Redshift or if there's more going on behind the scenes.
- silverrc21 12y agoAmplitude here - Most of our dashboards are powered separately from Redshift. We offer Redshift access as a way for our customers to answer more complex questions not offered by the dashboards.
- ripberge 12y agoWhy not power your dashboards with it? What do you use? I am considering using a columnar data store (maybe redshift) with a BI tool like bimeanalytics.com specifically to do dashboards.
- nemothekid 12y agoMy guess is latency - using Redshift for short lived, small queries might not be the best.
- deleted 12y ago[deleted]
- edwintorok 12y agoThe title should say 'Amazon Redshift'. At first I thought its going to be about redshift vs f.lux: http://jonls.dk/redshift/ http://jonls.dk/redshift/ Edit: Why the downvote? redshift (and flux) exist since before 2010, whereas Amazon Redshift got introduced just in 2012. I think it is reasonable to assume that someone who has never heard of Amazon Redshift would think of the open source project first (that exists in various distributions as packages), and not the Amazon service.
- untog 12y agoIf we took a poll I suspect the majority would be thinking of the Amazon service - I know I was. The date the projects were introduced isn't necessarily relevant.
- O____________O 12y agoThe UI colorizer is what I thought of immediately, too. If we took a poll I suspect the majority would be thinking of the Amazon service That's just personal projection, and is as irrelevant as an argument beginning with, "I think most people would agree that..." Personally, regardless of Amazon vs UI hack, I'm really tired of ambiguous naming in tech projects.
- chc 12y agoI'm much more tired of comments on ambiguous naming. There are at least two other people in my city who have my name, and many more who share either my first or last name. Somehow life goes on and this is not a topic of major controversy. But when two pieces of software have similar names, people just can't resist commenting endlessly and upvoting this content-free bikeshedding at the expense of actual discussion.
- jakobegger 12y agoYes! Especially when people use an existing word like 'redshift' as their product name. (It seems to be a popular choice! I remember there also was this astronomy software for the Mac called Redshift.) Of course, it gets even more ridiculous as the words get more common, eg see recent discussions about 'Paper' or 'Layout'
- ecaron 12y agoI wish he would talk about how they protect one customer from running a query that brings down the full stack. When we permitted Tableau to start talking to Redshift, we frequently encountered "Oh crap, Peter is running that query and and that's why everything is at a stand-still..."
- omgbear 12y agoYou can set up Workload Management[1] to restrict the amount of compute / query_slots each query/user can use. It splits the memory/compute into slices, and queries can use multiple slices, so you can get some fine-grained control, but it takes a bunch of work. [1] http://docs.aws.amazon.com/redshift/latest/dg/cm-c-modifying-wlm-configuration.html http://docs.aws.amazon.com/redshift/latest/dg/cm-c-modifying...
- sskates 12y ago*She :) I can understand why Ben Horowitz tries to default to using female pronouns.
- luckydata 12y agoHow do you guys handle the constant shifting of analytic schema that happens when handling a fast iterating application?
- silverrc21 12y agoWe update the table schema as we run into new fields, up to a limit. We also store the unstructured part of the data in a column that can be queried via json_extract_path_text.
- jlintz 12y agoIs each customer given their own redshift cluster for their data?
- sskates 12y agoNo, clusters are multi-tenant. We have a cap on the number of customers per cluster and we monitor usage to make sure no one customer is hammering the cluster.
- deleted 12y ago[deleted]
- deleted 12y ago[deleted]
- onewholovesfood 12y agoHas anyone used SnowflakeDB? Any feedback? Here's the website - http://www.snowflake.net/ http://www.snowflake.net/
- bnastic 12y agoO.T. but what the hell is a "Director Of Customer Success"?
- blumkvist 12y agoIt's pretty straightfoward I think. You have a complex product/service, with very diverse application scenarios. => Customer adoption is hindered by this complexity. => Customer is not getting value => Customer is angry and stops paying You hire a person who is familiar with the applications of your technology. He talks to customers to figure out what they want to do, how they plan to achieve it, what the hurdles. He helps them. Writes best practices, implementation plans, helps marketing to position and sales to close. It turns out so good that you hire many such people, who specialize in particular customer segments. Those people need management. You need Director of customer success. It's something between account manager/service/marketing.
- blumkvist 12y agoGreenplum (and associated tech) is partly open source now.