3 ms·
A better example might be a SAN briefly becoming unavailable due to a transient issue with your ISCSI network?
by LIV2 8y ago
A better example might be a SAN briefly becoming unavailable due to a transient issue with your ISCSI network?
- dooglius 8y agoYeah, this more or less clears it up, I was assuming an error implied disk failure.
- pgaddict 8y agoRight, this is about "ephemeral" failures which are becoming more common thanks to accessing storage over network, virtualization, thin provisioning etc.
- mprovost 8y agoNFS has been around forever and has always had a bad reputation due to problems like this. It mostly handles transient failures by waiting (indefinitely) for the server to return, but it's unclear what a better option is.
- masklinn 8y agoAnd that likely "hid" this issue for quite a long time, according to Tomas Vondra (https://youtu.be/1VWIGBQLtxo https://youtu.be/1VWIGBQLtxo): data loss would just be blamed on NFS being NFS and kinda crappy, and not necessarily properly investigated in full (why waste time on NFS shitting the bed yeah?), but it's likely the incorrect checkpointing / fsync assumptions were the culprit in at least some of the issues. Though an other factor is that people now run a lot more DBs, on a lot more environments, with a lot less reliability, and concurrently the database improved, so things which were rare and lost in the noise when run on "big iron" with expensive drive controllers become visible signal.
- macdice 8y agoFWIW here is a standalone test that shows Linux NFS exhibiting behaviour that would corrupt a PostgreSQL database: https://www.postgresql.org/message-id/CAEepm=1FGo=ACPKRmAxvb53mBwyVC=TDwTE0DMzkWjdbAYw7sw@mail.gmail.com https://www.postgresql.org/message-id/CAEepm=1FGo=ACPKRmAxvb... You can also tweak that test so that ENOSPC is discovered at close() time. Now you have a system that has thrown away data that PostgreSQL has already evicted from its own buffers, and there is no way to get it back (other than replaying the WAL, which is what PANIC achieves, as unpleasant a solution as it is, especially if it just happens again, and again, ...). The recent change in 11.2 adds a PANIC on error there. But I'm not sure it's sufficient in Linux NFS, because even on the tip of the master branch of Linux (by my inexpert drive-by reading, at least), the errseq_t stuff doesn't seem to have made it into the NFS client code, so it's still using the old single AS_EIO flag. That probably exposes at least one race that is discussed in this thread: https://www.postgresql.org/message-id/flat/CA%2BhUKGKa-HtBHJaBUJuZHsKwvVkxW2nE0W8BqRVOhNr5yNgiDA%40mail.gmail.com https://www.postgresql.org/message-id/flat/CA%2BhUKGKa-HtBHJ... I think we need to do something to make space allocation eager for NFS clients (a couple of concrete approaches are discussed) so that ENOSPC is excluded as a possibility after we have evicted data from PostgreSQL's buffer, and then I think we need Linux NFS to adopt errseq_t behaviour, and PostgreSQL to adopt the "fd passing" design discussed on the pgsql-hackers mailing list (to make sure the checkpointer's file descriptor is old enough to see all relevant IO errors). Or we need direct IO. TL;DR We are not out of the woods on NFS.
- wbl 8y agoWhy would you run a database on NFS?
- macdice 8y agoWell, I wouldn't. But people do. It makes more sense to use a SAN IMHO. I'm told it's not uncommon to use NFS for Oracle. One interesting thing is that they have their own NFS client implementation instead of trusting the kernel (they also do direct IO by default, though I'm not actually sure whether their NFS or DIO support came first).