4 ms·
He's missed the point by targeting column stores, however. Column stores are primarily aimed at data warehouses, which (generally) have the following
by jsrn 17y ago
He's missed the point by targeting column stores,
however. Column stores are primarily aimed at data
warehouses, which (generally) have the following
characteristics: [...]
He is not missing the point. Quote from the article:
"My conclusion is, on the whole, no. The column
store, when used to support existing petabyte
OLAP systems may be worth the grief, but for
transactional systems, at which the TRM is
aiming and from which column stores would extract,
not so much."
OLAP = Online Analytical Processing = Data Warehousing,
i.e. he primarily questions the use of column stores for
transaction processing (OLTP) and recognizes the
usefulness of column stores for large data warehouses.
- AlisdairO 17y agoThen he's criticising the applicability of column stores to a mode of use for which they're not designed. While there have been suggestions that correctly implemented column stores might provide acceptable OLTP performance, I don't see many people advocating them as a replacement for Oracle, DB2, or MSSQL. That seems to me to be missing the point, rather. "My interest is this: given the Next Revolution, do either a TRM or column store database have a purpose?" His post minimises the significance of data warehousing when it's a significant growth area, particularly for new players in the market. Column stores are applicable to DWs much smaller than petabyte scale - and indeed, I haven't seen any figures that suggest that Vertica has been scaled up that far yet. He's also ignored the impact of physical storage layer design on memory organisation and CPU cache performance. The use of SSDs is not a panacea that renders physical storage layout irrelevant. edit: quotes, and the realisation that 'minimalises' is not, in fact, a word.
- jsrn 17y agoThen he's criticising the applicability of column stores to a mode of use for which they're not designed. ah, ok. That makes sense, agreed. Here is question to you - not directly concerning the article - If we talk about all those new (or not so new) hash databases (a.k.a. "key/value stores" like Amazon Simple db, Google's equivalent which they make available with Google App Engine, Berkeley DB etc.) and if we talk about OLTP: I have been thinking lateley that those stores that often require the programmer to denormalize and trade relational features for performance could soon become largely irrelevant SSDs, and soon (OLTP databases are often not that big in my experience [compared to OLAP databases], many should fit into an SSD today). In contrast to column stores those key/value stores are explicitely marketed for transaction processing (if a RDBMS doesn't scale enough). As I understand it, most of the time performance and scaling problems with relational databases result from the database being disk bound - a problem that should largely vanish with SSDs. Do you agree?
- frig 17y agoYeah, that guy's writing is muddled and obscure, but I think your points are closer to (what I think is) his intended point: Column-oriented dbs are essentially a "heroic engineering" way to optimize a datastore for a particular workload -- continuous-read-heavy-batch-processing -- by engineering around the performance peculiarities of traditional hard drives. Clearly this is a sensible strategy for optimizing a datastore for certain workloads: at the moment that strategy delivers material performance improvements in the scenarios it's designed for (material enough that if the performance of your system on those workloads is economically important to you, it's worth the cost to build or buy a system that'd significantly speed things up). What I think he's saying is that the rise of ssd may make this engineering effort essentially useless outside of a handful of niches (essentially, the niches reduce to: data volumes too large to economically fit into ssd within the foreseeable future): - in a storage medium with heavy seek times, the engineering effort to implement a column store (instead of just using an off-the-shelf rdbms) can pay off in a somewhat broad range of usages - in ssd-ish storage media, the engineering effort isn't going to be worth the benefit, most of the time, compared to just using a stock rdbms The mention of the TRM fits into this picture of his intended claim. If you're not familiar the TRM is a mystery shrouded in an enigma: a bunch of big, credible names in the database world claimed have invented a radical new way of implementing the backend of a relational database that would've offered radically better performance characteristics (essentially it made joins so unbelievably 'cheap' that it was no longer necessary to denormalize for performance; supposedly the more-normalized you went the better TRM would perform). The issue with the TRM (transrelational model) is that: - the supposed core concept is patented, but doesn't explain en toto how it'd work (for obvious reason) - there's a ton of secrecy and ndas and so on surrounding anyone and everyone who got a good glimpse of the full picture -- the core inventors seem extremely protective of their ip, to the point pretty much nothing material has leaked about it's supposed workings - the company that was supposedly doing the first commercial implementation folded, ostensibly for non-technical reasons but again it's so secretive no one really knows what happened So it's a big mystery. There's basically a couple schools of thought on it: - it does actually work, but a comedy of errors / business climate / personality conflicts / whatever have prevented it from either being commercially implemented or from having a fuller picture of its workings disclosed publicly. Things have been quiet since then due to the protectiveness and penchant for secrecy on part of the principals. - it looked good on paper, but in doing the actual implementation some unavoidable complication turned up that prevented it from obtaining the needed performance (either at the time -- 2005ish -- or forever). The big names associated with it have kept quiet since then partly out of embarrassment (publicly endorsing a flop, kinda like hawking endorsing a free energy machine) and partly again out of concern for the principals' protectiveness - it was some kind of hypey thing that blew up in their faces; essentially a belief they could attract funding and customers by virtue of their reputations and claims of a revolutionary approach, combined with a belief that they could do an awesome-enough imlpementation of a non-revolutionary datastore approach -- basically do a best-practices, clean-room build of the current state of the art -- that no one would be the wiser I tend to think the second option is the likeliest story. All of that is a long windup for a very quick pitch: Assuming the TRM wasn't just bunk it would then be one of two things: - heroic engineering to bring revolutionary performance gains to systems built around disk-based storage - some kind of heroic datastore engineering that'd work better with a more ram-like storage system (eg ssds) That's a bit of a non-answer answer, but the connection to his main line of reasoning is something like: - if it's the former, then it's another example of heroic engineering made irrelevant by ssds, as even a 'traditional' rdbms can be made similarly performant on an ssd without all that effort; this is my take on the author's opinion - if it's the latter, then maybe there's something to be gleaned from it; I don't think this is what the author thinks So, yeah: from reading the rest of his blog (posts are wordy, but there's not that many of them) I think he's got the following idea: - implementing a traditional rdbms is hard, but mainly b/c of all the work you have to do to make it not perform like a dog under - ssds radically shake up the assumed performance contours of your persistent storage, enough so that a lot of the specifics designing an rdbms for performance might change - additionally, this guy has the impression this implementation might be "easy"; that is, a lot less work to do compared to writing a traditional rdbms from scratch There you have it, I think.