4 ms·
It really depends on your budget, data volume and nature. So each person can only offer advice. I did this a while back at my previous employer/startup (it got
by huy 9y ago
It really depends on your budget, data volume and nature. So each person can only offer advice.
I did this a while back at my previous employer/startup (it got acquired), we used AWS, PostgreSQL, Hadoop/Hive, using SQL + custom Ruby script for ETL + processing. You can read some of that writing here
https://engineering.viki.com/blog/2014/data-warehouse-and-analytics-infrastructure-at-viki/ https://engineering.viki.com/blog/2014/data-warehouse-and-an...
If I can distill them into bullet points, there probably are:
- SQL is great, stick with it.
- We started with PostgreSQL and scales a long way with it (using techniques like table partitioning, unlogged table, etc), then slowly split into Redshift, and later Hadoop/Hive.
- Handling events/behavioural data (JSON/semi-structure at large volume) are very different from handling transactional data (structured, lower volume).
- Take a more lean/incremental approach to it: get a basic DW setup, load data that you can immediately act on, act on it (build reports, run analysis, show to management), then repeat.
You can also check out my startup: https://www.holistics.io https://www.holistics.io , where I turned those experiences above into a data platform that automating DW + BI.
We allow you to do "lean ETL" on top of customers' DW infrastructure (be it Redshift, PostgreSQL, BigQuery or others). We work with small startup to unicorn tech companies.
- mattbillenstein 9y agoCurious why you went to Hadoop/Hive from Redshift? I usually hear of people going the other way. I would recommend staying away from Hadoop -- if you're a Java shop, it may make some sense if you're writing your own map/reduce jobs, but it seems particularly brittle and hard to debug although my info is a bit dated at this point. And generally if you must have Hadoop or Hive, the most elegant thing I've seen is EMR backed on S3 by .json.gz. You can spin up a cluster for a short time to run some jobs, and then spin it down when you're done -- hundreds of small instances seem to work well at this; but most of your interactive stuff should probably be in another system.
- vlahmot 9y agoWe do the EMR backed by s3 setup, only with snappy over gz as gz can't be split.
- mattbillenstein 9y agoAh, word, do you roll up the data by day? Or hour? I think in a situation where you roll it up by hour and you have a lot of files, it can be spread out pretty evenly on a large cluster.