8 ms·
I haven't actually built a cube myself, but I support a few right now. My boss, who built ours, always refers me to Ralph Kimball's "The Microsoft Data Warehou
by heynickc 11y ago
I haven't actually built a cube myself, but I support a few right now. My boss, who built ours, always refers me to Ralph Kimball's "The Microsoft Data Warehouse Toolkit" (you mentioned Oracle, so maybe there are synonymous toolsets out there)
I know you said you're aware of Ralph Kimball, but the first ~100 pages are broken into 1) Defining the Business Requirements and 2) Designing the Business Process Dimensional Model. It's really helped me wrap my head around the original design of the Facts and Dimensions as they relate to the business.
- thorin 11y agoThanks - amazon already tried to sell me this The Data Warehouse Toolkit: The Definitive Guide to Dimensional Modeling which I'd probably go for. I'm guessing the content would be similar but more database agnostic which suits my preferences. I feel comfortable doing the reporting and scripting but it would be good to validate my ideas for the designs when I get some hands on experience - hopefully soon! Maybe a new question but I wonder if the data warehouse has been superseded a little by the whole big data / multiple data sources thing?
- heynickc 11y agoYea I've wondered that myself. I picked up a book a while back called "Data Science for Business" http://shop.oreilly.com/product/0636920028918.do http://shop.oreilly.com/product/0636920028918.do Something that stood out to me was their discussion about "Big Data technologies" (Hadoop, HBase...) and how they "support" data mining techniques. So sometimes I wonder if big data technologies are more for the processing of the data, like the ETL needed for data warehouses - but since it's big, we need those special technologies to process it (and do additional data-sciencey things on it - like calculate a probability of churn, probability of loan default, etc). End results can still be those high-performance in-memory objects that we can slice and dice, just like our data warehouses, if that's how we need to see them. This is all coming from no practical experience with Hadoop / Big Data, just research so hopefully someone clarifies :)