3 ms·
This is a little misleading; an internal checkpoint happens upon transaction commit[1]. The comment makes it sound as if one or more transactions can commit be
by jstclair 14y ago
This is a little misleading; an internal checkpoint happens upon transaction commit[1]. The comment makes it sound as if one or more transactions can commit before a checkpoint writes them to disk.
[1]: http://msdn.microsoft.com/en-us/library/ms186259(v=sql.105).aspx http://msdn.microsoft.com/en-us/library/ms186259(v=sql.105)....
- spudlyo 14y agoThe transaction is actually written to disk twice. Once when it is written to the transaction log, and then again when a checkpoint happens and it's written to the underlying data blocks. Transaction log writes are sequential and fast, writes to the underlying data blocks are random and slow. The whole point of the checkpoint is to convert this sequential i/o to random i/o in an efficient manner. Still, that's not what we're talking about. We're talking about the difference between writing data to the transaction log, and writing data to the transaction log and then flushing it. Once you call fsync() and the data is flushed you know for a fact (provided the hardware isn't lying to you) the bytes are on the damn platter, and not in some OS buffer cache.
- Nitramp 14y ago... you'd think. However in practice most operating systems and/or file systems and/or disks cheat. fsync() is usually buffered in the hard disk itself. You can disable that and force the hard disk to truly write out on fsync, but that is so prohibitively slow that people rarely do that. If you want absolute durability, you'll have to have hard disks running on some battery buffered power supply, which is a common configuration. On the other hand, in a proper database system, at least the data files won't be corrupted by a missing fsync, so it'll come up. Figuring out whether that one commit did or did not go out in the very rare event of a fatal power failure and a just pending fsync() and the commit making it out of the network stack in time is probably a Heisenberg-esque inquiry into obscure realms of uncertainty.
- spudlyo 14y agoFirst of all, ext(2|3|4) and XFS filesystems honor fsync. That's what all the noise about Firefox and SQLite was all about. Secondly the HP RAID controllers I'm familiar with disable the drive write cache by default, and throw dire warnings if you try to turn it on, and make you ACK your choice like: Without the proper safety precautions, use of write cache on physical drives could cause data loss in the event of a power failure.... The bottom line is that Database pros insist on and get actual durability, and with a battery backed write cache, it's not painful.
- Nitramp 14y agoThe file system does, but does the hard drive? But I agree, getting battery backed write caches is not that hard, it's just not a default config.