3 ms·
> On the db side, in MySQL, it's possible to disable the doublewrite buffer (which is another form of overhead, which, in theory, ZFS should avoid), but nobody
by throwdbaaway 5y ago
> On the db side, in MySQL, it's possible to disable the doublewrite buffer (which is another form of overhead, which, in theory, ZFS should avoid), but nobody really knows if that is 100% reliable or not.
It is. Actually, that's the best thing about running MySQL on ZFS. The doublewrite buffer is not just another form of overhead, the doublewrite buffer is *the bottleneck* for any workload with moderate amount of write (until it got revamped in 8.0: https://dev.mysql.com/worklog/task/?id=5655 https://dev.mysql.com/worklog/task/?id=5655).
- pizza234 5y ago> It is. Actually, that's the best thing about running MySQL on ZFS. The doublewrite buffer is not just another form of overhead There has been an interesting discussion on the MariaDB mailing list, questioning this, called "Is disabling doublewrite safe on ZFS?"¹. The last post has been from the InnoDB lead, saying²: > I believe that it is technically possible for a copy-on-write filesystem like ZFS to support atomic writes, but for that to be possible in practice, the interfaces inside the kernel must be implemented in an appropriate way. It seems even he is not 100% sure on this topic. ¹=https://lists.launchpad.net/maria-discuss/msg05205.html https://lists.launchpad.net/maria-discuss/msg05205.html ²=https://lists.launchpad.net/maria-discuss/msg05219.html https://lists.launchpad.net/maria-discuss/msg05219.html
- throwdbaaway 5y agoIt doesn't sound like the guys understand how ZFS CoW works? So let's say the machine crashed after InnoDB had written 4KB out of a 16KB data file page to the filesystem. Due to CoW, the old data file page would still be there, unmodified. In contrast, on XFS or EXT4, without the doublewrite buffer, the data file page would be corrupted. I have a lot of respect for Marko Mäkelä, as he seems like the only one who tries to steer InnoDB away from an evolution dead end. For example, in https://jira.mariadb.org/browse/MDEV-24449 https://jira.mariadb.org/browse/MDEV-24449, he fixed a corruption bug that had existed since the very first InnoDB commit. At the same time, he also tries to simplify the internals, unlike Oracle who just churns out features along with regressions. But there are always more corruption bugs within InnoDB. I am not even sure the log checkpoint mechanism is sound, where it does a fsync of the data files first, before getting the checkpoint position from memory, which may have moved forward after the fsync. This behavior also exists since the very first InnoDB commit.
- pizza234 5y agoThanks for the information! > he also tries to simplify the internals, unlike Oracle who just churns out features along with regressions. Sadly, very true. I don't know how it was on 5.7 and before, however, at least since v8.0, even patch updates have an alarm probability of breaking existing installations, due to either bugs, or subtle changes in existing functionality.