4 ms·
So, one busy process performs a file operation that triggers a system restore checkpoint, and the OS locks the entire drive during this file operation? Sounds s
by strictfp 7y ago
So, one busy process performs a file operation that triggers a system restore checkpoint, and the OS locks the entire drive during this file operation? Sounds strange to me.
Is the problem that the checkpointing critical section has the same duration as the triggering file operation?
I get that there must be some sort of critical section for setting a checkpoint, but I don't understand why it takes so long, and why it would be affected by how busy the userspace process that triggered it is.
I would expect it to have a short barrier-style critical section; drain all outstanding writes, record some checksum or counter from a kernel data structure, and then release all writers again.
In my mind this should be kernel code only, entirely unaffected by userspace, and if designed nicely, quite fast.
So I guess I don't get what is going on here.
- kevingadd 7y agoMy guess would be that the system restore checkpointing functionality ends up holding a lock by necessity while it manipulates internal state. It would be really hard to make that sort of algorithm lock-free and preserve data integrity guarantees - I certainly wouldn't want to be responsible for writing it. Obviously the lock shouldn't be held so long and so often though...
- brucedawson 7y agoMy understanding is that the system restore checkpoints happen every five seconds. They hold a lock, which seems reasonable. The problem is that for some reason on this machine the checkpoint process was taking a really long time. I also don't understand why it was taking so long. It normally doesn't. Something went terribly wrong. > and if designed nicely, quite fast. Yep, should be. But it wasn't. If everything worked as it should then I'd never get to write any blog posts!