3 ms·
I am looking forward to testing this feature on our DBs running in Azure (on our own VMs - not hosted by Azure). We heavily use ZFS for compression (5:1 to 8:1
by rtp4me 1y ago
I am looking forward to testing this feature on our DBs running in Azure (on our own VMs - not hosted by Azure). We heavily use ZFS for compression (5:1 to 8:1) and tend to do lots of sequential scans. ZFS can hit about 150MB/sec on large scans, but with the new async-io feature and worker queues, I am hoping we can double or triple our seq scan performance. Time for testing!
- adgjlsfhk1 1y agowhy are you doing compression at the fs layer rather than the db layer? postgres supports compression and I'd assume it would do a better job than the filesystem
- rtp4me 1y agoBased on our testing, we get much better compression using ZFS than PGSQL. As a side benefit, we also get snapshots and the ability to easily find out how much compression our DBs are getting. According to a quick google search (to refresh my memory), PGSQL compression (eg: TOAST) targets specific large data values within tables, while ZFS compresses all data written to the ZFS pool.
- tempest_ 1y agoPostgres doesnt have great compression. Depending on your database you can reduce the size by 50% sometimes with high zstd zfs configurations. No lunch is free though. Aside from the obvious cpu cycles spent compressing configuring zfs / postgres is a pain in the ass and really depends on the trade offs and use cases.
- antonkochubey 1y agoHow is your experience overall running Postgres on a CoW filesystem? I thought that was something highly frowned upon - e.g. Postgres deployments on btrfs recommend setting chattr +C to the pg_data folder, essentially disabling CoW.
- supermatt 1y agoI’m not the parent so can’t comment on their experience, but ZFS groups the writes into transaction groups so it isn’t as write heavy as btrfs.
- rtp4me 1y agoAfter lots of trial and error, we have everything running pretty well. We mainly use ZFS for the cost savings on virtual drives (compression=zstd). Previously, we were using XFS and the DB sizes were >5TB or more, and now we can use much smaller disk sizes (1TB or so) and still get usable performance. That said, ZFS presents some challenges for a few reasons: - As you probably already know, PGSQL relies heavily on system RAM for caching (effective_cache_size). That said, ZFS and OS cache are NOT the same thing, thus you need to take this into consideration when configuring PGSQL. We normally set PGSQL effective_cache_size=512MB and use `zfs_arc_min` and `zfs_arc_max` options to adjust ZFS ARC cache size. We typically get a +95% hit rate on ZFS (ARC) caching. - ZFS is definitely slower than XFS or EXT4 and it took a while to understand how options like `zfs_compressed_arc_enabled`, `zfs_abd_scatter_enabled`, and `zfs_prefetch_disable` affect performance. In particular, the `zfs_compressed_arc_enabled` option determines if the ZFS cache data is compressed in RAM as well on disk. When enabled, this option can seriously affect latency since the data has to be uncompressed each time it is read/written. That said, a very nice side affect of `zfs_compressed_arc_enabled=on` is the amount of data in the cache. From my understanding, if you get 5:1 data compression on disk, you get the same for ARC cache. Thus, if you give ZFS 12GB of cache, you get about 60GB of data in ZFS memory cache. - Getting ZFS installed requires additional kernel modules + the kernel header files, and these files have to match the version of ZFS you want to run. This is especially important if you update your kernel very often (thus requiring new ZFS modules to be built and installed). Lots of blog posts are on the 'net describing some of these challenges. It's worth checking them out...
- abrookewood 1y agoHave you ever done a blog post describing these recommendations? It's super interesting to me, but not well discussed as far as I can tell.