9 ms·
Similarly, there are many annoying Linux bugs: `pthread_create` can sometimes return back a garbage thread value or crash your program entirely without any way
by bArray 7y ago
Similarly, there are many annoying Linux bugs:
`pthread_create` can sometimes return back a garbage thread value or crash your program entirely without any way to catch it or detect it [1]. High speed threadding is hard enough as it is, without the kernel acting non-deterministically.
Un-killable processes after copy failure (D or S state) [2]. If the kernel is completely unable to recover from this failure, is it really best to make the process hang forever, where your only available option is to restart the machine? I ran into this with a copy onto a network drive with a spotty connection, that actual file itself really didn't matter - but there was no way to tell the kernel this.
Out Of Memory (OOM) "randomly" kills off processes without warning [3]. There doesn't appear to be a way to mark something as low-priority or high-priority and if you have a few things running, it's just "random" what you end up losing. From a software writing stand-point this is frustrating to say the least and makes recovery very difficult - who restarts who and how do you tell why the other process is down?
[1] https://linux.die.net/man/3/pthread_create https://linux.die.net/man/3/pthread_create
[2] https://superuser.com/questions/539920/cant-kill-a-sleeping-process/541493#541493 https://superuser.com/questions/539920/cant-kill-a-sleeping-...
[3] https://serverfault.com/questions/84766/how-to-know-the-cause-of-a-oom-error-on-linux https://serverfault.com/questions/84766/how-to-know-the-caus...
- throwaway2048 7y agoThere is now a way to mark processes as OOM-killer exempt https://backdrift.org/oom-killer-how-to-create-oom-exclusions-in-linux https://backdrift.org/oom-killer-how-to-create-oom-exclusion... Part of the issue with processes stuck in D state (waiting for the kernel to do something) is that it is deeply tied into kernel assumptions about things like NFS, NFS is stateless, and theoretically severs can appear and disappear at will, and operations will keep working when it comes back. You can make NFS a hell of a lot less annoying in this regard by mounting it with soft or intr flags, however if the network disappears or hiccups, you WILL lose data (the network is NEVER reliable, in fact the entire model of NFS is arguably wrong to begin with)
- marmaduke 7y ago> the entire model of NFS is arguably wrong to begin with On local networks (everything attached to a single switch) with good hardware, it is reliable, and soft/intr is the worse choice among others. To wit NFS is one of the commonly supported VM storage options (libvirt, VMware, etc).
- yetanotherme 7y agoBest of the bad options doesn't make it a good option. It's far from reliable and one of the most common reasons I encounter for deadlocked Linux systems. Well, to be fair, it also hangs a fair share of BSDs and sometimes even a Solaris. If at all possible, I'd suggest avoiding it.
- marmaduke 7y agoWhat do you propose for a shared file system?
- mkesper 7y agoThere is no use in specifying (no)intr as of kernel 2.6.25, it can always be interrupted by sigkill. See e.g. https://access.redhat.com/solutions/157873 https://access.redhat.com/solutions/157873 There are many things "known" about NFS that are legacy.
- joosters 7y agoThat's not a problem with NFS, that's a fundamental issue with computers. Things fail. Nothing protects you from losing data, your local log-structured filing systems won't save you either. They'll help you protect an existing state from corruption, but they don't protect you from loss. That new request you just received when a hardware failure occurred? Say goodbye to it, you've no way of ensuring it will make it to storage when the disks have caught fire. Later on, when you've put out the blaze, all that algorithms can do is tell you when things started to get lost.
- 7y ago
- joshumax 7y agoWhat are you referring to regarding pthread_create()? Last time I checked I thought that it would return an undefined thread* only when giving a nonzero return code, which while certainly isn't handled in a lot of multi-threaded applications could be checked before anything else is done with the newly created thread.
- bArray 7y agoThere appears to be cases under high create/destroy scenarios where it returns zero, but failed to allocate memory for a thread during create. This is with tonnes of available memory (hence not a mapping error) and no exceptions thrown. That said, it's still entirely possible I've made a mistake. Please see here: https://github.com/electric-sheep-uc/black-sheep/blob/master/source/src/main.cc#L214 https://github.com/electric-sheep-uc/black-sheep/blob/master... The idea is that you do preliminary processing of a camera frame before sending a neural network over the top.
- datenwolf 7y ago> This is with tonnes of available memory (hence not a mapping error) and no exceptions thrown. Could be a memory fragmentation issue EDIT: Also since you're using C++ threads: You should really use move semantics, because right now you have two points of failure acting on the same thing: `new` operator may fail on creating the threads instance, and the underlying pthread_create may fail as well.
- datenwolf 7y agoComment on my mentioning on move semantics, because I feel that this really ought to be pointed out: In all the C++ standard library implementations of std::thread the only member variable of that class is the native handle to the system thread itself; there are no additional member variables! This means that the size of a std::thread object is equal to the size of a native handle, usually the size of a pointer, but sometimes smaller. If you create std::thread by `new` you're essentially creating a pointer to a "possibly a pointer", which comes with all the inefficiencies associated with it: Double indirection, small size allocations tend to fragment memory. And at the end of the day to actually use it, you have to at least lob around that outside pointer around on the stack anyway. So there is zero benefit at all of using dynamic allocation for std::thread. Don't do it! Just create the std::thread instances on the stack, they're just handled/pointers with "smarts" around them, and you can copy them around just efficiently as you can copy a pointer or an integer. Better yet, if you're not trying to "outsmart" the compiler you'll often get copy elision where applicable.
- idoubtit 7y agoPoint 3 is wrong. OOM killing is not random. Each process is given a score according to its memory usage, and the highest score is chosen by the kernel. The way to mark priority in killing is to adjust this score through /proc. All of this is documented in `man 5 proc` from `/proc/[pid]/oom_adj` to `/proc/[pid]/oom_score_adj`. http://man7.org/linux/man-pages/man5/proc.5.html http://man7.org/linux/man-pages/man5/proc.5.html
- bArray 7y ago> OOM killing is not random. Yeah I know, hence the use of "random". > The way to mark priority in killing is to adjust this score through /proc. Haven't heard about this, thanks for the heads up!
- fluffything 7y ago> > The way to mark priority in killing is to adjust this score through /proc. > > Haven't heard about this, thanks for the heads up! While going down that road is technically correct, it is a road full of pain. A slightly less painful strategy is to disable overcommit. That way, if memory pressure is high, and a process calls `malloc`, that call will fail if there is not enough memory, and that process will fail. If you only have a couple of processes in your system that are using most of the memory and you can control them, it is simpler to just making them resilient to these kind of errors, than to try to mess with the process score to control the OOM killer.
- enneff 7y agoBut any process can be trying to allocate at the time your system runs out of memory, and most applications are not authored to handle malloc failing. Process failure seems easier to work around, from what I’ve seen. Would love to hear more from someone who has contrary experience.
- AnIdiotOnTheNet 7y agoAt least you can then properly blame the software for doing the wrong thing and potentially patch it. The kernel should not implement global behavior that encourages improper memory allocation failure handling.
- azinman2 7y agoI have to say, reading all these replies about the OOM killer makes Linux look quite bad. These proc scores are not an elegant solution. I far prefer Darwin’s launchd which lets you set actual memory limits (soft and hard) that gives you warnings before you cross a threshold. Now this is more consumer OS oriented, but something equivalent for servers that let you express preferences in a more natural way seems desirable.
- the8472 7y agomemory limits are also available in linux via cgroups or ulimit.
- jimpudar 7y agoYou can use ulimit to set soft and hard limits for all sorts of system resources (including memory) on Linux.
- azinman2 7y agoTrue — I forgot about this. But can you do that on a process via some config _before_ the process is created?
- jimpudar 7y agoYeah, usually you call ulimit before calling the process. The new process inherits the limits. If you want to modify the limits of an _existing_ process, you can use prlimit.
- IcePic 7y ago..but if the problem is programs not checking malloc() return codes since it will not return failures, then ulimits will in themselves not help the program. It will help the OS to stay alive which is good, but we still need to deal with the programs who expect to run all the way into a swamp and sink without malloc giving them "bad news".
- TheDong 7y agosystemd lets you configure soft and hard limits on a per-service level, almost identically to launchd. See MemoryHigh= and MemoryMax= [1]. This does depend on cgroupsv2, but it works on most modern distros. [1]: https://www.freedesktop.org/software/systemd/man/systemd.resource-control.html#MemoryLow=bytes https://www.freedesktop.org/software/systemd/man/systemd.res...
- epiphanitus 7y agoWhat do you recommend doing when Linux Freezes? It doesn't come up a lot, but when it does it can be kind of unnerving since the three-finger-salute doesn't work. I would also love to know if anybody has a solution for getting video to play properly in Firefox. I know it's not a bug per se, but it would be nice to not have to switch between browsers all the time. I've been using Ubuntu for about a year now and otherwise its been a very positive experience.
- bArray 7y ago> What do you recommend doing when Linux Freezes? It doesn't come up a lot, but when it does it can be kind of unnerving since the three-finger-salute doesn't work. Unfortunately, I don't really have a solution for this. I have an SSD as the main disk and _even now_, when I hit this too hard Linux grinds to a halt. No mouse, no keyboard, just heat, fans and disk light. One thing that sometimes works for me is the old CTRL+ALT+FX mashing, but not always. Once you can get a shell you can type into you're okay, but of course this doesn't always work. > I would also love to know if anybody has a solution for getting video to play properly in Firefox. I know it's not a bug per se, but it would be nice to not have to switch between browsers all the time. What do you mean? It's been reliable for quite a long while? There were two issues I used to have, one was not having my graphics card setup (it was running from the CPU) and the other was not having flash (when that was something). > I've been using Ubuntu for about a year now and otherwise its been a very positive experience. Yeah, I think it makes one of the better daily drivers.
- epiphanitus 7y agoThank you for your answer. It's good to know that dealing with freezes is not a problem only faced by newbs like myself. Re: playing video, for some reason I can't play South Park on either Firefox or Chrome. Videos on twitter also won't load with Firefox, though they work fine with Chrome.
- bArray 7y ago> Thank you for your answer. It's good to know that dealing with freezes is not a problem only faced by newbs like myself. Yeah, it's another bug with Desktop based Linux. The problem is that Linux "basically" treats the GUI like any other program, when things get heavy everything gets roughly evenly screwed. OSes designed to be centralized around a GUI on the other hand usually guarantee that GUI related processes get a minimum amount of time on the CPU to ensure they don't freeze. Linux should 100% be doing this. Doing lots of hardware I/O shouldn't mean you lose the mouse or keyboard. When you lose control of input, you think the machine isn't doing something, when in actual fact it's doing tonnes, it's just not showing you. Even something like Android suffers from this under heavy load, it's really crap. The real joke is, it's probably a difficult kernel fix. You would need some kind of watch dog timer for the kernel to make sure it's not getting too bogged down with any one particular task and then interrupt ones that are (some tasks don't like to be interrupted) [1]. You then need make sure that all of the heavy kernel calls don't make guarantees about the call being completed (i.e. blocking), which to be completely honest should be the default position to take anyway. I'm not completely up-to-date on this, but my bet is that the issues come from everywhere. Programs can read/write arbitrarily large amounts of data (RAM, disk, network, bus, etc), when I believe you should be able to ask the kernel what size it would like you to read/write based on I/O activity and the capabilities of the device. If your program is bogging down the kernel, it should lover your recommended block read/write size. Better yet, this would be compatible with existing software, as they could choose to ignore this, possibly with it "punishing" programs that eat up lots of kernel time by making them wait longer for their next opportunity. There's a bunch of algorithms for time splicing tasks, but the most optimal appears to be max-min with very little organizing overhead [2]. > Re: playing video, for some reason I can't play South Park on either Firefox or Chrome. Videos on twitter also won't load with Firefox, though they work fine with Chrome. Hmm, that shouldn't happen, sounds like you potentially have some system-wide badness. A few things to try: * Disable any customization you made (extensions, add-ons, etc) - see if it is one of these interfering * Make sure you have all updates and you're running an up-to-date version of Ubuntu (this issues have possibly been patched already) * Make sure you have the correct drivers installed for your GPU But... One thing I did note was that JavaScript coin miners have gotten so bad that I can't run certain sites without ad-blockers anymore (uBlock Origin is generally recommended). I remember my CPU sitting at one core maxed out just because of the JS engine. I generally run uBlock Origin + NoScript on every page and manually enable temporary scripts on pages I trust. One of the biggest offenders for crazily heavy JS was actually Facebook. [1] https://en.wikipedia.org/wiki/Watchdog_timer https://en.wikipedia.org/wiki/Watchdog_timer [2] https://uhra.herts.ac.uk/bitstream/handle/2299/19523/Accepted_Manuscript.pdf?sequence=2&isAllowed=y https://uhra.herts.ac.uk/bitstream/handle/2299/19523/Accepte...
- altmind 7y agoThere is nothing more frustrating than unkillable processes stuck io iowait(D) state. There's no reason for this behavior to exist. And its so easy to hang forever - network blink, your NFS client gets stuck and your programs too.