3 ms·
... are Apple's manpages never read? https://developer.apple.com/library/mac/documentation/Darwin/Reference/ManPages/man2/fsync.2.html https://developer.apple.
by pudquick 13y ago
... are Apple's manpages never read?
https://developer.apple.com/library/mac/documentation/Darwin/Reference/ManPages/man2/fsync.2.html https://developer.apple.com/library/mac/documentation/Darwin...
"For applications that require tighter guarantees about the integrity of their data, Mac OS X provides the F_FULLFSYNC fcntl. The F_FULLFSYNC fcntl asks the drive to flush all buffered data to permanent storage. Applications, such as databases, that require a strict ordering of writes should use F_FULLFSYNC to ensure that their data is written in the order they expect. Please see fcntl(2) for more detail."
https://developer.apple.com/library/mac/documentation/Darwin/Reference/ManPages/man2/fcntl.2.html https://developer.apple.com/library/mac/documentation/Darwin...
"F_FULLFSYNC - Does the same thing as fsync(2) then asks the drive to flush all buffered data to the permanent storage device (arg is ignored). This is currently implemented on HFS, MS-DOS (FAT), and Universal Disk Format (UDF) file systems. The operation may take quite a while to complete. Certain FireWire drives have also been known to ignore the request to flush their buffered data."
OS X has aggressive file buffering in memory, and it's getting more aggressive all the time. For example, cfprefsd, introduced in 10.8 (https://developer.apple.com/library/mac/releasenotes/DataManagement/RN-CoreFoundationOlderNotes/ https://developer.apple.com/library/mac/releasenotes/DataMan...) made it so that when a system application read a preferences file, it stayed in memory and ignored the disk version, until cfprefsd eventually synced it back to disk. In 10.9, the behavior is much worse to the point that as soon as a pref is in cfprefsd, it's unlikely to leave it until the user logs out / the machine reboots.
In this instance, OS X has, for quite some time, had "defrag on the fly" for files under 20MB in size. On access of the file, it's read into memory and kept there in its entirety until memory pressure from other processes triggers a sync it back to disk. When it comes to writing a small file back to disk, OS X will "get around to it" when it's damned well ready unless you force its hand using the fcntl options above.
Unfortunately, the bit about "This is currently implemented on HFS, MS-DOS (FAT), and Universal Disk Format (UDF) file systems" covers pretty much the range of filesystem types that OS X can natively read+write on - but one that might get past this is ExFAT. I'd be surprised if that was the case, but it is natively supported read+write on OS X and would be something quick and easy to test (set up an ExFAT volume for the database) and possibly verify this is the root cause.
(Additionally, third-party read+write access to filesystems like NTFS via Paragon / Tuxera may be able to confirm this as well.)
More reading material (MySQL has been dealing with this since 2005): http://lists.apple.com/archives/darwin-dev/2005/Feb/msg00072.html http://lists.apple.com/archives/darwin-dev/2005/Feb/msg00072...
- jamesaguilar 13y agoLooks like a free $10k for you if you're right! Let's see!
- gigq 13y agoI believe to claim the reward you have to reproduce the issue, there is already a patch out for the F_FULLFSYNC change. https://code.google.com/p/leveldb/issues/attachmentText?id=197&aid=1970005000&name=0001-On-Mac-OS-X-fsync-does-not-guarantee-write-to-disk.-.patch https://code.google.com/p/leveldb/issues/attachmentText?id=1...
- cryptocoin 13y agoI'm not sure why you got that conclusion, as leveldb already received that fix some months ago, see https://code.google.com/p/leveldb/issues/detail?id=197 https://code.google.com/p/leveldb/issues/detail?id=197
- maaku 13y agoBitcoin is using an older version of leveldb (although, as mentioned, this fix is backported in a pull request).
- cypherpunks01 13y agoThere are two patches linked in the OP that switch to using F_FULLFSYNC on OSX. The OP says that people are still encountering db corruption even on branches with these fixes. https://github.com/sipa/bitcoin/commit/b28d8b423bddc860c5858a9df2982ce825835350 https://github.com/sipa/bitcoin/commit/b28d8b423bddc860c5858... https://github.com/gmaxwell/bitcoin/commit/e7bad10c12ce9b5d424ac273c1c977b88469d46c https://github.com/gmaxwell/bitcoin/commit/e7bad10c12ce9b5d4...
- pudquick 13y agoWe'll, I'm glad someone apparently IS reading :) But again - I'd point to the work of other longstanding database projects that are available on OS X as a source of "how we ensured data correctness".