12 ms·
Working with Files Is Hard (2019)
- deleted 2y ago[deleted]
- wruza 2y agoNo mention on ntfs and windows keywords in the article, for those interested.
- yahayahya 2y agoIs that because the windows APIs are better? Or because businesses build their embedded systems/servers with Windows?
- wruza 2y agoI doubt that, was just curious how it might compare in the article.
- p_ing 2y agoCertainly depends on which APIs you ultimately use as a developer, right? If it is .NET, they're super simple, and you can get IOCP for "free" and non-blocking async I/O is quite easy to implement. I can't say the Win32 File API is "pretty", but it's also an abstraction, like the .NET File Class is. And if you touch the NT API, you're naughty. On Linux and macOS you use the same API, just the backends are different if you want async (epoll [blocking async] on Linux, kqueue on macOS).
- pjc50 2y agoThe windows APIs are certainly slower. Apart from IOCP I don't think they're that much different? Oh, and mandatory locking on executable images which are loaded, which has .. advantages and disadvantages (it's why Windows keeps demanding restarts)
- pjdesno 2y agoAlthough the conference this was presented at is platform-agnostic, the author is an expert on Linux, and the motivation for the talk is Linux-specific. (Dropbox dropping support for non-ext4 file systems) The post supports its points with extensive references to prior research - research which hasn't been done in the Microsoft environment. For various reasons (NDAs, etc.) it's likely that no such research will ever be published, either. Basically it's impossible to write a post this detailed about safety issues in Microsoft file systems unless you work there. If you did, it would still take you a year or two of full-time work to do the background stuff, and when you finished, marketing and/or legal wouldn't let you actually tell anyone about it.
- wmf 2y agoUniversities can get Windows source code under NDA and do research on it but nobody really cares about such work.
- pjdesno 2y ago"Getting windows source code under NDA" doesn't necessarily mean "can do research on it". If you can't publish it, it's not research. If the source code is under NDA, then Microsoft gets the final say about whether you can publish or not, and if the result is embarrassing to Microsoft, I'm guessing it's "or not".
- continuational 2y ago> Pillai et al., OSDI’14 looked at a bunch of software that writes to files, including things we'd hope write to files safely, like databases and version control systems: Leveldb, LMDB, GDBM, HSQLDB, Sqlite, PostgreSQL, Git, Mercurial, HDFS, Zookeeper. They then wrote a static analysis tool that can find incorrect usage of the file API, things like incorrectly assuming that operations that aren't atomic are actually atomic, incorrectly assuming that operations that can be re-ordered will execute in program order, etc. > When they did this, they found that every single piece of software they tested except for SQLite in one particular mode had at least one bug. This isn't a knock on the developers of this software or the software -- the programmers who work on things like Leveldb, LBDM, etc., know more about filesystems than the vast majority programmers and the software has more rigorous tests than most software. But they still can't use files safely every time! A natural follow-up to this is the question: why the file API so hard to use that even experts make mistakes?
- Retr0id 2y ago> why the file API so hard to use that even experts make mistakes? I think the short answer is that the APIs are bad. The POSIX fs APIs and associated semantics are so deeply entrenched in the software ecosystem (both at the OS level, and at the application level) that it's hard to move away from them.
- __loam 2y agoPOSIX is also so old and essential that it's hard to imagine an alternative.
- jcranmer 2y agoNot really, there's been lots of APIs that have improved on the POSIX model. The kind of model I prefer is something based on atomicity. Most applications can get by with file-level atomicity--make whole file read/writes atomic with a copy-on-write model, and you can eliminate whole classes of filesystem bugs pretty quickly. (Note that something like writeFileAtomic is already a common primitive in many high-level filesystem APIs, and it's something that's already easily buildable with regular POSIX APIs). For cases like logging, you can extend the model slightly with atomic appends, where the only kind of write allowed is to atomically append a chunk of data to the file (so readers can only possibly either see no new data or the entire chunk of data at once). I'm less knowledgeable about the way DBs interact with the filesystem, but there the solution is probably ditching the concept of the file stream entirely and just treating files as a sparse map of offsets to blocks, which can be atomically updated. (My understanding is that DBs basically do this already, except that "atomically updated" is difficult with the current APIs).
- Retr0id 2y ago> they found that every single piece of software they tested except for SQLite in one particular mode had at least one bug. This is why whenever I need to persist any kind of state to disk, SQLite is the first tool I reach for. Filesystem APIs are scary, but SQLite is well-behaved. Of course, it doesn't always make sense to do that, like the dropbox use case.
- ziddoap 2y ago>SQLite is the first tool I reach for. Hopefully in whichever particular mode is referenced!
- Retr0id 2y agoWAL mode, yes!
- nodamage 2y agoBefore becoming too overconfident in SQLite note that Rebello et al. (https://ramalagappan.github.io/pdfs/papers/cuttlefs.pdf https://ramalagappan.github.io/pdfs/papers/cuttlefs.pdf) tested SQLite (along with Redis, LMDB, LevelDB, and PostgreSQL) using a proxy file system to simulate fsync errors and found that none of them handled all failure conditions safely. In practice I believe I've seen SQLite databases corrupted due to what I suspect are two main causes: 1. The device powering off during the middle of a write, and 2. The device running out of space during the middle of a write.
- ablob 2y agoI believe it is impossible to prevent dataloss if the device powers off during a write. The point about corruption still stands and appears to be used correctly from what I skimmed in the paper. Nice reference.
- SoftTalker 2y agoOnly way I know of is if you have e.g. a RAID controller with a battery-backed write cache. Even that may not be 100% reliable but it's the closest I know of. Of course that's not a software solution at all.
- praptak 2y agoExt4 actually special-handles the rename trick so that it works even if it should not: "If auto_da_alloc is enabled, ext4 will detect the replace-via-rename and replace-via-truncate patterns and [basically save your ass]"[0] [0]https://docs.kernel.org/admin-guide/ext4.html https://docs.kernel.org/admin-guide/ext4.html
- gavinhoward 2y agoI wonder if, in the Pillai paper, I wonder if they tested the SQLite Rollback option with the default synchronous [1] (`NORMAL`, I believe) or with `EXTRA`. I'm thinking that it was probably the default. I kinda think, and I could be wrong, that SQLite rollback would not have any vulnerabilities with `synchronous=EXTRA` (and `fullfsync=F_FULLFSYNC` on macOS [2]). [1]: https://www.sqlite.org/pragma.html#pragma_synchronous https://www.sqlite.org/pragma.html#pragma_synchronous [2]: https://www.sqlite.org/pragma.html#pragma_fullfsync https://www.sqlite.org/pragma.html#pragma_fullfsync
- userbinator 2y agoI don't get it. The only times I've had problems with filesystem corruption in the past few decades was with a hardware problem, and said hardware was quickly replaced. FAT family has been perfectly fine while I've encountered corruption on every other FS including NTFS, exFAT, and the ext* family. Meanwhile you can read plenty of stories of others having the exact opposite experience. If you keep losing data to power losses or crashes, perhaps fix the cause of that? It doesn't make sense to try to work around it.
- nodamage 2y ago> If you keep losing data to power losses or crashes, perhaps fix the cause of that? I keep telling my users to make sure to plug their phones in before the battery dies, but for some reason they keep forgetting...
- userbinator 2y agoThen that's entirely their fault. They deserve all the corruption they get.
- userbinator 2y agoSeems like I hit a nerve. Apparently teaching users responsibility is a bad thing? No wonder things are "hard". Because otherwise many in this godforsaken industry wouldn't need to be employed.
- dooglius 2y agoPhones shut down when close, but before they hit zero battery
- PaulHoule 2y agoThere was that time (2009 or so?) I wrote 2 million files to a single directory on NTFS and that filesystem was never the same again. It didn't seem to be a hardware problem. I used to be really careful to not put a crazy number of files in a directory on Linux and Windows storing them in subdirs like b7/b74a/b74a56 where the digits are derived from a hash of the file name but lately I've had some NTFS volumes with a 1M file directory that seem to be OK. Hardware problems also manifest in mysterious ways. On both Windows and MacOS I had computers that seemed to be OK until I did an OS update which caused enough IO that a failing HDD was pushed over the edge and the update failed; in one case I was able to roll back the update but not apply the update, in another case the machine was trashed. Careful investigation (like taking the disk out and inspecting it on another computer) revealed a hard drive error although there was no clear indication of this in the UI and the average person would blame to software update
- ryao 2y ago> On Linux ZFS, it appears that there's a code path designed to do the right thing, but CPU usage spikes and the system may hang or become unusable. ZFS fsync will not fail, although it could end up waiting forever when a pool faults due to hardware failures: https://papers.freebsd.org/2024/asiabsdcon/norris_openzfs-fsync-failure/ https://papers.freebsd.org/2024/asiabsdcon/norris_openzfs-fs...
- ein0p 2y agoZFS on Linux unfortunately has a long standing bug which makes it unusable under load: https://github.com/openzfs/zfs/issues/9130 https://github.com/openzfs/zfs/issues/9130. 5.5 years old, nobody knows the root cause. Symptoms: under load (such as what one or two large concurrent rsyncs may generate over a fast network - that's how I encountered it) the pool begins to crap out and shows integrity errors and in some cases loses data (for some users - it never lost data for me). So if you do any high rate copies you _must_ hash-compare source and destination. This needs to be done after all the writes are completed to the zpool, because concurrent high rate reads seem to exacerbate the issue. Once the data is at rest, things seem to be fine. Low levels of load are also fine.
- deleted 2y ago[deleted]
- ryao 2y agoThere are actually several distinct issues being reported there. I replied responding to everyone who posted backtraces and a few who did not: https://github.com/openzfs/zfs/issues/9130#issuecomment-2614107797 https://github.com/openzfs/zfs/issues/9130#issuecomment-2614... That said, there are many others who stress ZFS on a regular basis and ZFS handles the stress fine. I do not doubt that there are bugs in the code, but I feel like there are other things at play in that report. Messages saying that the txg_sync thread has hung for 120 seconds typically indicate that disk IO is running slowly due to reasons external to ZFS (and sometimes, reasons internal to ZFS, such as data deduplication). I will try to help everyone in that issue. Thanks for bringing that to my attention. I have been less active over the past few years, so I was not aware of that mega issue.
- einpoklum 2y agoThe article wrap up with this salient point: > In conclusion, computers don't work (but I guess you already know this...
- paulddraper 2y agoThey work. Just not all the time.
- 1vuio0pswjnm7 2y agoNo Javascript or SNI: https://archive.wikiwix.com/cache/index2.php?rev_t=&url=https%3A%2F%2Fdanluu.com%2Fdeconstruct-files%2F https://archive.wikiwix.com/cache/index2.php?rev_t=&url=http...
- edgarvaldes 2y agoAs per HN headlines, files are hard, git is hard, regex is hard, time zones are hard, money as data type is hard, hiring is hard, people is hard. I wonder what is easy.
- ssivark 2y agoTo reuse another HN headline, all this is probably because no one really cares X-)
- paulddraper 2y agoComplaining :)
- D-Coder 2y agoSelection error. The stuff that always works doesn't get posted here.
- jheriko 2y agothis whole thing is a story about using outdated stuff in a shitty ecosystem. its not a real problem for most modern developers. pwrite? wtf? not one mention of fopen. granted some of the fine detail discussion is interesting, but it doesn't make practical sense since about 1990.
- rep_lodsb 2y agoThe article is about the hardware and kernel level APIs used for interacting with storage. Everything else is by necessity built on top of that interface. "fopen"? That is outdated stuff from a shitty ecosystem, and how do you think it's implemented?
- AutistiCoder 2y agoit's a good thing I'm a Web developer. closest I come to working with files is localStorage, but that's thread safe.