7 ms·
Author of the GitHub repo here. I wanted a list of papers that would be essential in building database systems. It was little bit sad to see many members of th
by rxin 12y ago
Author of the GitHub repo here.
I wanted a list of papers that would be essential in building database systems. It was little bit sad to see many members of the community re-discovering and re-inventing the wheels from 30+ years of relational database and systems research. This list of papers should help build a solid foundation.
- tinco 12y agoIt's a nice list but I have the feeling you can read each of these articles and still not be able do write a halfway decent database because there's nothing on how to actually write data to the disk efficiently on any operating system. I for one would be very interested in that, although perhaps the deprecation of spinning disks has reduced the complexity of the issue.
- walterbell 12y ago> how to actually write data to the disk efficiently on any operating system Discussion between Postgres & Linux devs: http://lwn.net/Articles/590214/ http://lwn.net/Articles/590214/ & https://news.ycombinator.com/item?id=7521008 https://news.ycombinator.com/item?id=7521008
- neilc 12y agoperhaps the deprecation of spinning disks has reduced the complexity of the issue SSDs are much more complicated than spinning disks, so if anything the problem has gotten worse.
- tinco 12y agoAre they? see I really need articles about this :) I thought higher bandwidth and lower seek times meant less complexity, but I could be mistaken.
- walterbell 12y agoFirmware on the SSD is actively in the data path, e.g. wear-levelling to ensure that many writes to a single file does not cause a particular block to wear out much earlier than other blocks, reducing overall capacity of the drive.
- batbomb 12y ago"halfway decent database" is a very vague notion. What would that even be? Something similar to SQLite, MySQL, Cassandra, VoltDB? Are you talking about write ahead logging of flushing dirty pages from the buffer (aka sequential or random writes)? All of these database systems have extremely different host requirements and write characteristics. Even SQLite and MySQL can be pretty different depending on the storage layer (SQLite can use BerkeleyDB) The future isn't MySQL and spinning disks, the future is a tens of different databases that do things and handle reliability in a variety of ways. I know that complicates the issue even more, but my point would be that writing a halfway decent database is possible (or impossible, depending on your perspective) within a lot of this literature. ( Also, spinning disks even have abstractions since mostly you are using a RAID controller with a decent size of cache to alleviate random writes. )
- tinco 12y agoI completely agree with you with regards to the future of databases, it will definitely be running many databases. What I meant with halfway decent database is one that can actually guarantee the data it committed is actually on the hard drive in a readable state, and that can sustain random writes at a rate close to the theoretical limit of the persistent storage device. I'm not judging any databases in particular. I'm asking how can I make a database that's as awesome as MySQL or Cassandra. All these databases have already taught us a lot of what's in these articles, but one thing that isn't written about a lot is how they actually achieve high performance, not just in a theoretical way, but in a practical way. What system calls do they make, what schemes do they use to minimize seeks, how is the memory laid out. None of that high level "do we use Paxos of Raft", that's kids stuff.
- batbomb 12y agoYou should check out Database Systems: The Complete Book. Those papers have a lot of that information, but it's usually a bit buried. > What system calls do they make I think that's actually kind of out of scope for a paper, because already you are assuming an OS and a language. Many papers/articles assume neither. > what schemes do they use to minimize seeks Look for information on buffer management/buffer managers > how is the memory laid out Almost always in pages that correspond to disk blocks. Also it's dependent on lots of things (transactional model, index types, column vs row orientation, etc..) This is heavily tied into the buffer management stuff too.
- brasetvik 12y agoGreat initiative. This one is a classic as well: What Goes Around Comes Around (Michael Stonebraker, Joseph M. Hellerstein) – http://citeseerx.ist.psu.edu/viewdoc/summary?doi=10.1.1.113.5640 http://citeseerx.ist.psu.edu/viewdoc/summary?doi=10.1.1.113.... "This paper provides a summary of 35 years of data model proposals, grouped into 9 different eras. We discuss the proposals of each era, and show that there are only a few basic data modeling ideas, and most have been around a long time."
- billconan 12y agoThank you very much for this list!