9 ms·
The Sorry State of Copy-On-Write File Systems
- transfire 11y agoSomeone should write an article on the sorry state of file system in general. ZFS and BTRFS are improvements, though still not quite there yet. But distance between user and storage still seems vast and primitive. Perhaps Seagates Kinetic drives are the future we need? (http://www.seagate.com/tech-insights/kinetic-vision-how-seagate-new-developer-tools-meets-the-needs-of-cloud-storage-platforms-master-ti/ http://www.seagate.com/tech-insights/kinetic-vision-how-seag...)
- transfire 11y agoWonder if anyone has ever thought about building a SQL server directly in a hard drive?
- zokier 11y agoI think more interesting question would be if anyone has build practical userland on top of SQL (or any reasonably rich database for that matter). Here is some discussion about SQL on raw devices: http://dba.stackexchange.com/questions/80036/is-there-a-way-to-store-a-postgresql-database-directly-on-a-block-device-not-fi http://dba.stackexchange.com/questions/80036/is-there-a-way-...
- protomyth 11y agoThe Newton skipped SQL but had an object database with queries.
- scurvy 11y agoThere are a few drives out there which have a k/v interface as opposed to block interface. Seagate Kinetics come to mind: http://www.seagate.com/tech-insights/kinetic-vision-how-seagate-new-developer-tools-meets-the-needs-of-cloud-storage-platforms-master-ti/ http://www.seagate.com/tech-insights/kinetic-vision-how-seag...
- kibibu 11y agohttps://en.wikipedia.org/wiki/WinFS https://en.wikipedia.org/wiki/WinFS edit: apologies, you meant on the controller presumably
- fredkbloggs 11y agoBecause what everyone wants is more invisible closed-source software running on underpowered CPUs? Not to mention the limitations on performance, capacity, and reliability scaling that would be inherent in binding the database instance to a single disk device. Or were you suggesting that each of these tiny controllers running deeply proprietary software should also form a distributed database in concert with host software or HBA controller firmware? To say that this is not a good idea would be putting it mildly. The existing efforts to build more functionality into disk drives, starting with FDE and now with some object-store-like interfaces, sounds real appealing at first. Less software to write, functionality everywhere, disks seem to Just Work today so wouldn't it be nice if they could Just Work in some more ways too. However, as soon as you start thinking about building something larger or more interesting than a toy, it becomes apparent that the disk drive is the wrong place for this. The knowledge of the problem is not present at that level of the system, and the interfaces and processing power are inadequate to express it. The problems of error recovery, scaling, and debuggability cannot be solved there. It might be useful for a few of the very smallest consumer-grade applications where none of these concerns are significant (nor likely to be solved by a higher-level vendor anyway), but it's not generally viable.
- jhayward 11y agoOne of the big log search startups was doing postgres query engine integration on drive controllers at least 10 years ago. (sorry, don't recall which) And as far back as 1964 IBM was doing K-V in hardware on the drive in their count-key-data devices. [1] [1] https://en.wikipedia.org/wiki/Count_key_data https://en.wikipedia.org/wiki/Count_key_data
- gnoway 11y agoDiscussion from 6 months ago: https://news.ycombinator.com/item?id=9128404 https://news.ycombinator.com/item?id=9128404
- jsprogrammer 11y agoIsn't COW a fundamentally hard problem? How does one expect a complete, comprehensive solution?
- oconnore 11y agoPeople who use filesystems like this seem to completely misunderstand what RAID is for. They seem to care about their data somewhat, which of course means they have off-site backups. Given that they have reliable backups, redundancy is only useful to the extent that it improves uptime. Uptime is best maintained by eliminating single points of failure. Raid is a good first step, but at some point, creating a massive disk array on a single motherboard/cpu/disk-controller is anything but that. And they don't seem to be complaining of downtime. Given that they have already achieved data safety, and don't care about uptime, the only reasonable explanation I can surmise is that they're confused.
- gnoway 11y agoIn your opinion, what is the best way to address the use case of spreading I/O across multiple spindles for non-uptime-related performance reasons?
- oconnore 11y agoUsing any of the btrfs supported RAID levels will improve read performance, and anything but level 1 will improve write performance. I like Raid10.
- toomuchtodo 11y agoWho even cares about spindles anymore? Need fast? SSD on the PCI bus. Need lots of storage? Spinning disk where you care not about performance.
- simoncion 11y ago> Who even cares about spindles anymore? Folks who have a shitload of data to store, but want more perf than a single drive can give them?
- rapala 11y agoDoes the off-site backup also contain the write that happened 1.5 seconds ago? What about the corruption that happened just before taking the backup?
- yellowapple 11y agoTranslation: "ZFS and btrfs are 'incomplete' because their handling of striped RAID is incomplete." Disregarding the fact that things like mdadm still exist, and further disregarding the fact that the vast majority of filesystems out there don't bother with trying to implement RAID at all (probably because - again - there are things like mdadm that do that already), RAID5/6 are generally a bad idea compared to a RAID1 or RAID10. Both RAID5 and RAID6 make multiple very incorrect assumptions: * Failure of multiple drives in a short period of time is rare (in reality, if one drive fails (excluding bathtub-curve-related infant mortality), the likelihood of subsequent drive failures increases significantly) * The failed drive can be replaced and repopulated quickly (in reality, this is becoming less and less true as drives get bigger and bigger, thus taking more and more time to rebuild the failed array member; SSDs buy some time here, but that's not sustainable) * Bit rot / "cosmic rays" are rare (in reality, silent data errors happen all the time, as has been demonstrated [0] as an example of why RAID5/6 is woefully insufficient for even the most basic protection against data corruption) Basically, if you want something striping, and care at all about data integrity, go with RAID10. Only use RAID5/6 if you don't care about data loss (whether due to a comprehensive backup policy or a comprehensive redundancy policy on a machine-level), though in that case you might as well just cut to the chase and use RAID0. I really wish this article would've dug into some of the real shortcomings of these filesystems; btrfs in particular would be incredibly useful compared to a more traditional LVM approach if it supported file encryption and swap subvolumes (for either of these things, LVM (+ LUKS for encryption) is necessary with or without btrfs). Instead they get criticized over not supporting things that a sane sysadmin wouldn't touch with a ten-foot pole in this day and age. [0]: http://www.miracleas.com/BAARF/Why_RAID5_is_bad_news.pdf http://www.miracleas.com/BAARF/Why_RAID5_is_bad_news.pdf
- stephengillie 11y agoNot to mention the "write hole" where a single write requires 4 reads on RAID5. RAID0+1 provides better write speeds because less compute and fewer reads are needed. Striped RAIDs are a thing of the past. The closest we could come today is, effectively, "spanning datacenters".
- 11y ago
- Vegemeister 11y agos/tripple/triple/g