3 ms·
My biggest pet peeve in the world of data storage is that the most powerful ideas from relational data theory are still mostly locked up in books and papers, an
by ef4 12y ago
My biggest pet peeve in the world of data storage is that the most powerful ideas from relational data theory are still mostly locked up in books and papers, and most programmers who think they know what "relational" means aren't aware of the full picture.
The people who did all the pioneering work on the subject (mostly E.F. Codd and Chris Date) seem to have made several blunders in trying to popularize their ideas. For example, one of Date's later books ("Go Faster: The Transrelational(tm) Approach to DBMS Implementation") contains some very useful ideas. But it was written in 2001 and not published until 2011, because he was working with a now-failed startup that was trying to keep it all a trade secret.
Most of their writing is not available online. You have to buy their books. Which is an author's prerogative, but seriously limits the reach of the ideas.
The world thinks it already has relational databases that are good enough. Convincing it otherwise requires a web-savvy marketing approach that has so far been lacking.
- jules 12y agoI just checked out that book. The results are presented as some kind of revolution but: 1. The system presented is simply an inefficient way of doing a single column index on each column. 2. The book contains NO benchmark results. It only claims that it's efficient because of X, Y and Z without any numbers to back that up. Meanwhile the method involves a lot of pointer chasing which is extremely slow, especially on disk. Queries that require multi-column indices will also be extremely inefficient of course, because they will require scans. I wouldn't go so far as saying that the book is worthless, but its claims are certainly dubious at best.
- barrkel 12y ago"Queries that require multi-column indices will also be extremely inefficient of course, because they will require scans." If a query uses just the columns that are in a multi-column index, and doesn't sort in any order other than the index order, then the query may be much more efficient than a with a single-column index - since all the data may be retrieved from the index, rather than seeking to the row. I think you left something out.
- jules 12y agoWe can decide which indices we have, you know. The conditions that you state are not by far the only conditions when a multi column index helps you. Even if you do not sort by the index order, and you do not have all the data in the index, the index may still be way more efficient. For example if you have SELECT A WHERE X=3 AND Y=5 ORDER BY B with an index (X,Y).