3 ms·
We've evaluated it for our cloud database service, but ended up going back to ext4 due to memory management bugs. In particular this one: https://github.com/zf
by lfittl 10y ago
We've evaluated it for our cloud database service, but ended up going back to ext4 due to memory management bugs.
In particular this one: https://github.com/zfsonlinux/zfs/issues/3645 https://github.com/zfsonlinux/zfs/issues/3645
- kim0 10y agoThis kind of thing is what's better on platforms like freebsd
- Zancarius 10y agoI ran into a bug [1] that appeared to be a deadlock triggered by rsync of a relatively modest directory with tons of small files. The only thing that fixed it was `spl_taskq_thread_dynamic=0` to stop the dynamic spawning of ZFS-related kernel threads. Rather than a memory management issue, it'd stall the copy and peg my CPUs indefinitely. I suspect it's fixed in 0.7.0 based on some of the other related bugs I've run into since, but I've been reluctant to upgrade as of yet. Otherwise, ZFS on a home file server has been great. [1] https://github.com/zfsonlinux/zfs/issues/3808 https://github.com/zfsonlinux/zfs/issues/3808
- tobias3 10y agoThe Ubuntu ZFS module sets spl_taskq_thread_dynamic to zero by default. My problem was that the large number threads with spl_taskq_thread_dynamic=0 cause it to OOM, at one point it had an OOM at mount (with 32GB RAM). When I set spl_taskq_thread_dynamic=1 I had the same deadlock issue (That's where I stopped trying to use it and went back to FreeBSD).
- Annatar 10y agoThe fix is not to work around the out of memory killer in Linux, but to turn off memory overcommit altogether: the system should not lie to applications that it has more virtual memory than is actually available, because that is extremely detrimental to correctness of operation. illumos based systems, for instance, never lie about such things.
- tobias3 10y agoZFSOnLinux is a kernel module and not a user space application (and you cannot OOM kill kernel threads) and the overcommit probably isn't direct, e.g. there are 64 worker threads who all need some memory to do the work and tell the kernel that memory allocation must not fail. If the memory isn't available either Linux starts killing userspace processes or gets stuck. The same thing can happen in Solaris except there ZFS doesn't use a separate memory mechanism which does not properly react to memory pressure.