6 ms·
Running PostgreSQL on Compression-enabled ZFS
- old-gregg 13y agoCan this simply be an artifact of terrible disk I/O on AWS or overall difference between ZFS/ext3? Do you think the results would have been similar if you were to use no-compression-ZFS instead of ext3 on a proper database hardware? Basically trying to figure out if the low performance of uncompressed dataset is specific to AWS/ext3. Thanks.
- mfenniak 13y agoold-gregg makes a great observation here. The addition of a ZFS benchmark without-compression is needed to isolate the compression as a factor in the speedup. That aside, I thought this was a wonderful article with non-intuitive findings. Very interesting, CirtusDB [edit: er, CitusDB]. :-)
- codewright 13y agoCitus
- cwsteinbach 13y agoold-gregg's point is valid. Since we didn't benchmark ZFS without compression we can't say for sure how much of the performance improvement is attributable to compression vs. just ZFS. As far as AWS goes, we have noticed ephemeral disks connected to the same instance can exhibit fairly large performance differences, and attempted to control for that in our tests by reusing the same disk for each test run.
- mattbillenstein 13y agoI always question benchmarks on ec2 because of "noisy neighbor" effects - maybe that's in the noise, but someone trying to replicate your results would perhaps see significantly different results depending whether their VM landed on a busy node or not... Tag this one #ymmv +1 on testing uncompressed ZFS I did see a blog post about MySQL with similar results (at least compression was a significant win) some time ago - disks are so slow compared to what throughput modern CPUs are capable of on these sorts of compression algorithms.
- kevinastone 13y agoYou're also testing on a c1.xlarge that gives you excess CPU compared to I/O, so it's potentially biasing your results.
- cwsteinbach 13y agoWhile I can't claim that we logged CPU load while running these tests, I can say that I watched the output of top and iotop and that the CPU load was relatively light. It's also worth pointing out that Amazon describes the I/O performance of c1.xlarge instances as "high". We also considered using an hs1.8xlarge "High Storage" instance for these tests, but eventually decided that we were more interested in testing against conventional disks as opposed to SSDs.
- lucian1900 13y agoDid you use instance storage? EBS? Provisioned IOPS? There are vast differences between those three.
- bsg75 13y agoI have seen similar advantages comparing XFS to ZFS+compression on a local server (Centos 6.3, ZFS on Linux 0.61). Using a 2 disk striped volume for PostgreSQL 9.2, I get an average of 2.5X compression (as reported by ZFS), and a 1.5 to 2X time reduction in database restores (single threaded or 8 jobs in parallel). Given this development box has relatively slow 7200 RPM disks, the tradeoff of more CPU time for less disk transfer makes sense. Edit: My use case is an OLAP server. I can't state how the tradeoffs affect OLTP performance.
- mbell 13y agoAnother issue is that ZFS is extremely aggressive with caching data in ram (L1ARC). That can eat up memory you'd rather give to the database heap and also tends to skew benchmarks.
- bsg75 13y agoPlacing limits on ARC (25% for me currently) limits that effect.
- wiredfool 13y agoYes, but postgres' design in that area actually helps. Postgres relies on the OS caching the data tables for the most part. There is some caching in shared buffers, but generally that's not a huge portion of the memory of your db system. (10-25%, and not more than a few gigs)
- rscale 13y agoI run Postgres on ZFS, and simply limit the amount of memory dedicated to the ARC.
- trotsky 13y agoUsing compression with zfs on solaris derived platforms serving as a san/nas backend for vsphere appears to speed up every workload backed on rotating storage. Well, not if vmdks use guest full disk encryption, but that's understandable.
- ars 13y agoI wouldn't recommend doing benchmarking on a virtual server. You have no idea how busy the real server is, (noisy neighbors, etc), so it's impossible to have comparable results from benchmark to benchmark.
- jaytaylor 13y agoFWIW, If you use one of the largest instance types (4x large or whatever), the VM will probably be on it's own host which would mean you're unlikely to have neighbors ;)
- skeletonjelly 13y agoWhen benchmarking, it's best to remove assumptions based on "probably" though right?
- marshray 13y agoIt's cloud.
- reeses 13y agoYou still have to deal with the storage fabric.
- atoponce 13y agoI'm curious of the rest of the architecture. Each benchmark needs to be tested separately, as ZFS is likely caching the reads in the ARC. We also need a benchmark of ZFS without compression enabled. However, we're not showing how bad ext3 is, but that the end result still shows the stellar performance, compression or not.
- GalacticDomin8r 13y agoext3 does suck for certain workloads, one of them being large scale db's. And by suck, I mean dangerous. Unless you really want to set barrier on the fs and watch your IO plummet to 45 record player speeds.
- cwsteinbach 13y ago> as ZFS is likely caching the reads in the ARC Each of the seven queries we used in our benchmark required a sequential scan of the 32GB dataset. It's unlikely that the ARC had any impact on the results since the EC2 instance had only 7GiB of memory.
- iso8859-1 13y agoWhat is the reason for using ext3 over ext4?
- fsiefken 13y agoThe Btfrs and Reiser4 filesystems also support transparent compression and might currently be a better alternative to increase Postgresql query speed. Btfrs supports gzip, LZO, LZ4 and Snappy and is in the mainline linux kernel, Reiser4 is still maintained and available as a patch on Linux 3.8.5 (latest is 3.8.8) and supports LZO and gzip (alternatively there are also the embedded NAND flash medium compatible filesystems F2FS and UBIFS which both improve on the JFFS2 filesystem and it's transparent compression). For I/O bound queries SSD drives (in your preferred raid configuration) also will speed up the system. Btfrs has built-in support for TRIM SSD already, Reiser4 TRIM/SSD support is being discussed among the remaining developers.
- laumars 13y agoReiser4 is in a weird place after the conviction of Hans. I'm not sure I'd want to trust a production system on it. And I've been less than impressed with BtrFS on the test systems I've ran it on (though I'm aware there's others who swear by it - I'm only talking about my experiences). ZFS is a fantastic file system, but I can't help wondering if part of the issue is the fact that the benchmarking was conducted on a virtual machine. ZFS is better suited for raw disks than virtual devices (again, just my anacdotal evidence. I've never ran benchmarks myself).
- yunong 13y agoDo you have any benchmarks to support your claim? Statements such as "Btrfs ... might currently be a better alternative" without benchmarks are worthless. Anyway -- I'd be interested to see benchmarks of Btrfs on GNU/Linux vs ZFS on illumos -- I suspect that ZFS "might currently be a better alternative". Simply ratcheting off a set of features and stating that Btrfs is "better" is dubious at best, and perhaps mis-leading. As the OP stated in his blog post, ZFS has a rich feature set -- which we find invaluable in our own postgres stack -- features such as incremental snapshots, a real copy on write filesystem, etc.
- GalacticDomin8r 13y agoThis isn't the first time benchmarks like this have been done and these results are consistent with the earlier ones. It shouldn't surprise most people that enabling transparent compression gives these benefits. Why you ask? Well what is the largest bottleneck in a system? Disk IO - by far. So all ZFS is doing is transferring workload to a subsystem you likely have plenty of(CPU) from one that you have the least of(Disk IO/latency)
- nemothekid 13y agoIf I'm reading this right, with ZFS compression enabled I am seeing 1/3rd disk usage and 3x increase of speeds in query times just from switching the filesystem. Stats like that make me very skeptical. Does this mean that I can get a 3x increase in speed while cutting my disk space down by a third just by switching to ZFS? If so, why isn't everyone doing this?
- lwat 13y agoThe way I make sense of this is that you need fewer (slow) disk reads to get the same amount of data into RAM, so that might explain the speedup? I agree that it sounds too good to be true though.
- rosser 13y agoYour read is correct. Once CPU time spent in decompression became less than disk wait time for the same data uncompressed, the reduced IO with compression started to win — sometimes massively. As powerful as processors are these days, results like these aren't impossible, or even terribly unlikely. Consider the analogous (if simplified) case of logfile parsing, from my production syslog environment, with full query logging enabled: # ls -lrt ... -rw------- 1 root root 828096521 Apr 22 04:07 postgresql-query.log-20130421.gz -rw------- 1 root root 8817070769 Apr 22 04:09 postgresql-query.log-20130422 # time zgrep -c duration postgresql-query.log-20130421.gz 19130676 real 0m43.818s user 0m44.060s sys 0m6.874s # time grep -c duration postgresql-query.log-20130422 18634420 real 4m7.008s user 0m9.826s sys 0m3.843s EDIT: I'm not sure why time(1) is reporting more "user" time than "real" time in the compressed case.
- caf 13y agozgrep runs grep and gzip as two separate subprocesses, so if you have multiple CPUs then the entire job can accumulate more CPU time than wallclock time (so it's just showing you that you exploited some parallelism, with grep and gzip running simultaneously for part of the time).
- tracker1 13y ago
- danbruc 13y agoThe result doesn't really surprise me - many operations are bound by the available bandwidth. There is even a compressor named Blosc [1] that speeds up operations by moving compressed data between memory and L1 cache and (de)compressing it there instead of moving the uncompressed data. [1] http://blosc.pytables.org/ http://blosc.pytables.org/
- petsos 13y agoCan someone give us an overview of the state of ZFS on Linux? Last time I had checked it was implemented over fuse. Has this changed?
- cdjk 13y agoThere are kernel modules here, which is what I assume they're using: http://zfsonlinux.org http://zfsonlinux.org The licensing problems only apply to distributing CDDL and GPL code that have been compiled into the same binary, not running a CDDL-licensed module in a GPL kernel - I think. My experience with ZFS (which is awesome, btw) comes from FreeBSD.
- BUGHUNTER 13y agoZoL looks pretty good - unfortunately if you want Samba on ZoL, of course with snapshots and ACLs, you will have problems, as ACL mapping is not implemented, if I understood things well. That is a real pitty, ZFS is great, Samba is great, Linux is great and having these three things working smoothly together without having to spend weeks of research on how to get it running would help many admins to finally get away from commercial clown & bloat systems. However, the groundwork is done and if we are lucky in 2014 the Linux + Samba4 + ZFS dreamteam will be available as a stable replacement.
- jacob019 13y agoWould love to see these performance metrics on a powerful system with pcie or raided SSD's. Would be interesting to find the tipping point where the extra CPU time outweighs the IO reduction. Even if the DB layer performs better total application response time could be negatively impacted for CPU intensive work loads as the compression steals cycles from the application layer.
- jamhan 13y agoIs it just me or is "Compression Ratio" a poor label for the graph in that article? Normally, when one uses "Compression Ratio", it is the opposite of those numbers, i.e. EXT3 storage would be 1:1, ZFS-LZJB would be 2:1 (not 0.5), and ZFS-gzip would be 3.33:1 (not 0.3). It's a small thing I know but it turns convention on its head in its current form. A better label would be perhaps "Storage Size Ratio".
- marshray 13y agoI don't see a problem with them expressing the ratio as a decimal since it becomes a simple multiplier of the original file size 38GB x 0.3. But it's downright misleading to show the vertical axis from something other than 0.0 to 1.0 when comparing ratios. They start it at 0.2. In reality, LZJB is saving 50% of the space whereas gzip saves 70%. But a naive glance at the graph implies gzip look roughly 3 times smaller/better than LZJB. Classic "How to Lie with Statistics" stuff.* I would have expected better from an "analytics" database. * Not saying they intend to lie here but it's representative of the classic text https://en.wikipedia.org/wiki/How_to_Lie_with_Statistics https://en.wikipedia.org/wiki/How_to_Lie_with_Statistics
- jamhan 13y agoIf you read in any other article something like the following: "Taking Product X as having a baseline compression ratio of 1, Product Y had a compression ratio of 0.5 and Product Z had a compression ratio of 0.3", I'm pretty sure 99.9999% of the HN population would interpret that as Products Y and Z having worse compression than X, not better. That's my point.
- marshray 13y agoThis academic-looking paper (first hit I tried from Wikipedia) gives the standard definition of "compression ratio" as compressed/uncompressed size (section 4.2), consistent with the linked article. I'm pretty sure you're impression of 99.9999% of the HN population is wrong.
- cafard 13y agoI tried running Oracle on ZFS for a while, with fairly terrible results. A bit of examination showed that ZFS was fine for table scans but had bad performance with indexes. It may be possible to tune one's way around this, but I simply dumped ZFS in favor of Automated Storage Management.