6 ms·
What I’m looking for is a governor that makes sure processes like X/Wayland that handle user input are always available to handle user input at the cost of othe
by Osiris 4y ago
What I’m looking for is a governor that makes sure processes like X/Wayland that handle user input are always available to handle user input at the cost of other processes.
If I’m running a video encoding using all available CPUs, I don’t want that to cause lag with all my other processes.
This happens to me a lot when running node processes like unit tests that running on all cores. It slows everything else down. I even set the nice value on node to always be the lowest priority and it only helps a little.
I’ve actually had Linux become completely unresponsive when running many large node instances. That shouldn’t happen.
- pengaru 4y agoYour compositor could already be run at realtime priority... the risk is if the process has bugs which may result in infinite loops, you may lose the ability to interact with the system.
- PlutoIsAPlanet 4y agoWouldn't you also need some kind of priority system in the graphics system? If you have a game open using all the available GPU resources, doesn't matter if you can't block input if you can't get any updates to the screen.
- pengaru 4y agoWell, for anything actually putting pixels on the screen under the compositor's purview, the compositor controls the framerate of the game and there needs to be context switching involving the scheduler to get things displayed. I'm not clear on what mechanisms exist today for preventing things like a DoS of GPU resources in something of a single entry to the GPU driver without returning to userspace. We've all seen things like GPU stalls reported in dmesg, so clearly the drivers are tasked with preventing that sort of hang where a process enters the driver and stays there far too long.
- jlokier 4y agoThe Linux kernel since 2.6.25 provides CPU time limits on real-time priority so that runaway real-time tasks don't prevent other interactive processes from having some CPU time. So you can still interact with the system. "man 7 sched" provides a lot of detail about the real-time scheduling policies and options. (Of course if the compositor is essential to the kinds of interaction you need, e.g via the GUI, you'll still be stuck if the compositor stops working. But that's not a real-time priority problem.)
- pengaru 4y agoNice! The last time I programmed something with SCHED_FIFO was long ago, and involved a lot of sysrq.
- invalidator 4y agoIt will always have some impact because you're thrashing the caches. A couple more things to try: Reduce swappiness. Linux tries to swap out idle pages to free memory for block cache, even when there's only light memory pressure. If you have a dozen processes working hard and your UI is untouched for several minutes, it can get swapped out and lag. sudo sysctl vm.swappiness=1 Set the CPU affinity for your build processes to exclude CPU 0, ensuring one is always free for UI: taskset 0xFFFFFFFE nice your_build_command_here Neither is perfect but they can help.
- jeffbee 4y agoOddly, bazel will defeat your attempts to use taskset to corner it. Bazel users need to actually remove CPUs from its control group.
- post-factum 4y agoSetting the nice level is not enough. Instead, `SCHED_IDLE` policy should be applied to the workload that is being run in the background. [1] may help if you run such a workload from CLI. [1] https://codeberg.org/post-factum/litter https://codeberg.org/post-factum/litter
- arakageeta 4y agoSCHED_IDLE is problematic. There is no priority inheritance with SCHED_OTHER threads. SCHED_OTHER threads block for prolonged periods of time when they contend for a mutex in the kernel with a SCHED_IDLE thread. The SCHED_IDLE thread holds the mutex, but it is unable to complete its critical section because it keeps getting preempted.
- wongarsu 4y agoWindows handles this by giving the foreground window time slots that are three times longer than normal (unless you are on a server version or change the setting). It also boosts the priority of threads that get woken up after waiting for IO, events, etc, as opposed to being CPU bound. Which I guess Linux's fair scheduler also kind of ends up doing, even if via an entirely different mechanism (simply by observing that they used less CPU). It's interesting how different operating systems take completely different approaches to their schedulers. Linux seems to try to make quite sophisticated schedulers, trying out very different concepts, but keeps the scheduler very seperated and "blind" to the rest of the system. Meanwhile Windows has an incredibly simple scheduler (run the highest-priority runnable thread, interrupt it if its timeslot is over or a more important thread become ready, repeat) and puts all the effort into letting the rest of the kernel nudge priorities and timeslot lengths based on what the process or thread is doing.
- ixmerof 4y agoWhat's your source of knowledge for Windows scheduler? You made me curious that we can actually know this
- wongarsu 4y agoMy primary source would be the Windows Internals book, which conveniently talks about this in the free sample chapter of the 5th edition [1] (thought I imagine it's worth getting the seventh edition to see what changed since then). I'd say that's pretty authorative, having at least one senior kernel developer in the author list. There are also various blog posts around the net, but they seem to be using this book/chapter as their source too. Of course if you really want to know, the source code of Windows XP got leaked and is easy enough to find on github, so it should be possible to verify. I don't really know my way around that ball of source code though. 1: https://www.microsoftpressstore.com/articles/article.aspx?p=2233328&seqNum=7 https://www.microsoftpressstore.com/articles/article.aspx?p=...
- saint_yossarian 4y agoHave you tried a low-latency kernel, like Liquorix?
- binkHN 4y agoI haven't looked to see how Android handles this internally, but, when developing Android apps, the "Main" thread is for UI work. Anything else is supposed to be offloaded to other threads and some Android libraries will simply crash if you make them try to do non-UI work on the Main thread.
- charleslmunger 4y agoAndroid sets the priority of the top app's main thread and render thread to -10. Devices often also configure a cpuset reservation for the top app.
- roomey 4y agoHave you tried xfce window manager? By far the last laggy desktop I've ever used. It's so light, even when a system is highly loaded, it still works fine (in my experience anyway)
- Osiris 4y agoI use i3. I'm talking about the entire X Window System not getting enough priority on the CPU when there is an oversubscription to CPU resources. I don't think it matters what the window manager is if mouse and keyboard events aren't even being handled.
- Dalewyn 4y ago>If I’m running a video encoding using all available CPUs, I don’t want that to cause lag with all my other processes. Run the encoder at a lower priority?
- Vecr 4y agoUnder /usr/bin/chrt --idle 0 as well if you can.
- michaelmrose 4y agoYou can do this with ionice/nice/renice at the start of the operation—or with a wrapper script—but various daemons have existed to hang out and apply a policy for at least 22 years starting if I recall correctly with "and "the autonice daemon. ananicy ulatencyd or if you are lazy call a script with a cron job.
- atq2119 4y ago> I’ve actually had Linux become completely unresponsive when running many large node instances. This is almost certainly not a scheduling problem but a swapping problem. Linux is famously bad at handling these kinds of situations. About five years ago I decided to never use swap anymore, and I couldn't be happier with that decision.
- iforgotpassword 4y ago> About five years ago I decided to never use swap anymore, and I couldn't be happier with that decision. How does that make sense? Either your workload fits into RAM, so no swapping occurs and everything is responsive. Or your workload is too large to fit into RAM so a) you have a swap partition and you system starts swapping and everything becomes really unresponsive, but the task will finish. b) you don't have swap so the kernel kills the process using most memory, which is most likely the thing you're just trying to run. How can you be happy with option b? I mean obviously get more RAM either way, but I'd rather have that safety net for the occasional memory intensive task. Obviously not a solution if you're short on RAM by a huge margin but if it's just a GB or two it's acceptable, because it's likely that the kernel can just swap out stuff from background programs that don't actively access their stuff right now.
- Asooka 4y agoBecause the failure mode is gentler. With swap, once you start a process that needs more RAM than you have, you enter a swapping death spiral, the system becomes completely unresponsive and you have to hard reset. Without swap, you only get one process killed.
- iforgotpassword 4y agoNo, as I described that is not the case. Not every process running needs all the memory it allocated all the time. A GB or two can easily be swapped out from your browser, Networkmanager, file manager, email client, etcpp. swapping only becomes a problem if pressure is so high that you constantly swap in and out the same pages. You can see this well from the kennel's proactive swapping. If I run a bog standard desktop Linux distro and leave swappiness at the default 60, after a few hours of normal usage where I don't even get close to filling up my RAM with actual memory allocations from processes, the kernel will have swapped out almost a GB of data from several background services. Because a lot of programs allocate some memory for something and then never look at it in a long time. So it makes more sense to swap that to disk and instead use that ram as page cache. Also, with current NVMe speeds, swapping has become a lot more bearable. Really the only place where I can see why you wouldn't use swap is servers with a well defined, dedicated workload. Here I indeed want a fast and noticable fail mode in case a new version of whatever I'm running might have a memory leak, or just vastly different behavior regarding memory usage.
- arglebargle123 4y agoFor some time now I've been spawning intensive multicore jobs with nproc-1 or even nproc-2 workers on my Linux laptops. It's too easy to end up with an unresponsive machine when you spawn an all core test/build/compute job - and as nice as enforced "my code is compiling" breaks are people tend to get the wrong idea when they see you on your phone or with a book in hand while you wait.