3 ms·
Which is a false dichotomy, if you use the right storage mechanism for your column-oriented DB, the comparable relational DB storage will actually be larger sin
by fkdjs 14y ago
Which is a false dichotomy, if you use the right storage mechanism for your column-oriented DB, the comparable relational DB storage will actually be larger since it has to index every column to provide the equivalent of what you can do with column-oriented storage. Also, another myth he brings up is that you can't have complex joins with column-oriented DBs. It's completely possible.
- garysieling 14y agoFor practical purposes, you usually don't index every column. The index is a value-add for performance gains, which as far as I've seen requires custom implementation in map-reduce databases. When it is implemented, you would have the same storage problem. As an example you can look at Common Crawl, which has a public index of web page data. They provide a hadoop database of page source, and a smaller data set of page text. The page text database serves a similar function to an index. Using the text dataset instead of full HTML would be like an "Index Scan" in database optimizer terms. I don't think he said you can't do complex joins; he said people tend to denormalize the data before putting it into NoSQL databases.
- fkdjs 14y agoIf you compare the two, then you must compare apples with apples. That is, you must compare relational DBs where every column is indexed since with column-oriented DBs, you can search by arbitrary columns without having to worry about which column is indexed. You can say that in practice you don't need this, so you reduce functionality but you're no longer comparing apples to apples. Besides, it's nice to search by any column, just because relational DBs limit you doesn't mean it's not useful. map-reduce databases are something entirely different, although you can do map-reduce via joins if need be. You denormalize because, among other things, you don't have complex joins. With column-oriented DBs that can perform joins, denormalization is not necessary. NoSql is something entirely different.