3 ms·
Applications should account for it and measure it themselves, when an application chooses to wait (to persist writes, for example), then it's hard to argue the
by zaarn 4y ago
Applications should account for it and measure it themselves, when an application chooses to wait (to persist writes, for example), then it's hard to argue the kernel should take that as load. A similar argument could otherwise be made if an app is waiting on a futex or other resource. I don't think an fsync counts either. The solution there, IMO, is to use IO fences. Tell the OS that async writes and reads may not reorder beyond point X in time for this thread (or all threads, optionally for all FDs or just a specific one).
Then your app, like Oracle, simply issues a fence and can be done with it. If the app crashes before the fence is persisted, then it must be able to resume work from a previous one (a simple example would be that between WAL checkpoints, a fence is issued). The app won't have to worry that only some specific writes completed vs some others not beyond what the fence permitted. Additionally a good mechanism might be a call to wait for a fence to be persisted to disk.
Simple example for the WAL use case:
1. Write new transaction to WAL
2. Issue an IO fence for the WAL file
3. Write new data to database file
4. Wait for Fence from 2
5. Return success
This is roughly equivalent to the synchronous example:
1. Write new transaction to WAL file
2. Issue fsync
3. Write new data to database file
4. Return Success
In case writes to the database file are lost, you can recover from the WAL (as intended). Notably Step 4 of the Async Example is a case where the thread is waiting but it can do other useful work while that is happening. The same thread can offload the work and simply issue more IO in the meanwhile and return success to the client once it sees the correct wait-for-fence returning in the async queue. And it won't have to wait for AIO event completion/checkpoints like currently, that make the system load non-indicative of app load (though frankly, system load is never indicative of app load, apache2 doesn't increase load if it runs out of workers).