25 ms·
Neat. My reading of these changes imply that they finally made the L2ARC's info survive a reboot. For some background: ZFS, a modern WAFL clone [1], has a rep
by cokernel_hacker 13y ago
Neat. My reading of these changes imply that they finally made the L2ARC's info survive a reboot.
For some background:
ZFS, a modern WAFL clone [1], has a replacement algorithm called ARC [2] which can concisely be described as a hybridized MRU/MFU (Most Recently Used/Most Frequently Used) replacement algorithm to decide which pages make the most sense for keeping in memory. There is considerable literature surrounding replacement algorithm design, I have little to say about ARC other than it is patented [3] and can be outperformed by newer algorithms.
Note that this is quite different from the traditional approach of FS/buffer-cache design. One usually expects the OS kernel to manage the buffer-cache for you (OS X has it's Unified Buffer Cache (UBC), NT has it's Cache Manager (Cc), etc.). However ZFS includes it's own incredibly complex caching subsystem around. I do not know why they didn't want to improve or modify the Solaris kernel's segmap subsystem but there are consequences to this design. Notably, ZFS's memory usage is quite a bit higher because of ARC.
The idea of performing read-caching in memory with ARC seemed like such a good idea to the ZFS designers that they allow for a second level of ARC to take place: L2ARC. L2ARC essentially runs the ARC algorithm between SSDs and HDDs to, hopefully, speed up the performance of random reads in a ZFS storage pool.
Now to steer back towards what this code dump seems to be about. If you recall from before, ZFS's ARC is a replacement algorithm based on usage and it needs to know which things to put where. This so-called persistent L2ARC remembers where things were on a L2ARC device so that the storage pool can take advantage of the fact that data is on the SSD on, say, a reboot.
Huh? Why did this require extra code? Remember, ARC was about caching: it didn't need to remember anything. When coming back online, complicated things happen: transactions get replayed, metadata integrity needs to be rechecked, etc. Implementing a persistent cache that is crash safe is incredibly difficult but not uncommon: auto-tiering [4] solutions like Fusion Drive [5] have to provide this kind of safety.
[1] http://en.wikipedia.org/wiki/Write_Anywhere_File_Layout http://en.wikipedia.org/wiki/Write_Anywhere_File_Layout
[2] http://en.wikipedia.org/wiki/Adaptive_replacement_cache http://en.wikipedia.org/wiki/Adaptive_replacement_cache
[3] http://patft1.uspto.gov/netacgi/nph-Parser?patentnumber=6996676 http://patft1.uspto.gov/netacgi/nph-Parser?patentnumber=6996...
[4] http://en.wikipedia.org/wiki/Automated_Tiered_Storage http://en.wikipedia.org/wiki/Automated_Tiered_Storage
[5] http://en.wikipedia.org/wiki/Fusion_Drive http://en.wikipedia.org/wiki/Fusion_Drive
- dmpk2k 13y agocan be outperformed by newer algorithms I'm interested to hear more about this. There's CAR... what else? As an aside, ZFS's ARC algorithm differs a fair bit from IBM's -- a case of theory which then met reality. I don't recall the details though, alas.
- cokernel_hacker 13y agoOff of the top of my head: CLOCK-Pro [1] - an approximation of LIRS [2] It's quite popular, I remember that MySQL and one of the BSDs use it. [1] http://www.cse.ohio-state.edu/hpcs/WWW/HTML/publications/papers/TR-05-3.pdf http://www.cse.ohio-state.edu/hpcs/WWW/HTML/publications/pap... [2] http://www.cse.ohio-state.edu/hpcs/WWW/HTML/publications/papers/TR-02-6.pdf http://www.cse.ohio-state.edu/hpcs/WWW/HTML/publications/pap...
- gnoway 13y agoA modern WAFL clone? I've never read that before. The article you link to doesn't assert that either. Can you provide mode information?
- cokernel_hacker 13y agoThe original ZFS paper references WAFL wrt its similarity a number of times. It seems like the biggest distinction that the paper claimed was that ZFS had pooled storage and WAFL was network oriented. WAFL's biggest idea of the day was "write-anywhere" (the WA in WAFL). Write-anywhere is another way of phrasing copy-on-write which is a fancy way of saying _never overwrite_. The idea, while simple, can be built upon to yield features like cheap snapshots and reasonable data integrity. Perhaps "clone" is a bit too much but the similarity is definitely there. FWIW, the NetApp folks sued Oracle because they also thought it looked similar [2] [1] http://users.soe.ucsc.edu/~scott/courses/Fall04/221/zfs_overview.pdf http://users.soe.ucsc.edu/~scott/courses/Fall04/221/zfs_over... [2] http://www.netapp.com/us/company/news/press-releases/news-rel-20100909-oracle-settlement.aspx http://www.netapp.com/us/company/news/press-releases/news-re...
- 13y ago