9 ms·
3,200% CPU Utilization
- kachapopopow 2y agoYep, ran into this way too many times. Performing concurrent operations on non thread-safe objects in java or generally in any language produces the most interesting bugs in the world.
- Espressosaurus 2y agoWhich is why you manage atomic access to non-thread-safe objects yourself, or use a thread-safe version of them when using them across threads. Multithreading errors are the worst to debug. In this case it's dead simple to identify at design time and warning flags should have gone up as soon as he started thinking about using any of the normal containers in a multithreaded environment.
- kachapopopow 2y agoTell that to inexperienced developers or making a massive single-thread project have multi-threaded capabilities.
- baggy_trough 2y agoMulti-threading - ain't nobody got time for that.
- mrkeen 2y agoYeah, our software politely waits for one customer to finish up with their GETs and POSTs before moving onto the next customer. We have almost one '9' of uptime!
- baggy_trough 2y agoThere are better ways than threading.
- mrkeen 2y agoYeah, like pretending you aren't
- baggy_trough 2y agoI don't know what you mean.
- saagarjha 2y agoasync/await, for example.
- mrkeen 2y agoIf you write: int balance = 0; int get {...} void set(int balance) {...} void withdraw(int amt) { if amnt <= balance { balance -= amnt; } } it will be flagged immediately in code review as a race condition and/or something that doesn't guarantee its implied balance>=0 invariant. Because threading. But lift the logic up into a REST controller like pretty much all web app backends: GET /balance/get POST /balance/set POST /balance/withdraw And it will sail straight through review. (Because we pretend each caller isn't its own Thread.)
- baggy_trough 2y agoYou're right, we can't get away from concurrency control with database, cache, even the file system. I'm still happy to never have to think about in-process multithreading bugs.
- stuff4ben 2y agoI've been that developer making a single-threaded app multi-threaded. Best way to learn though!
- BobaFloutist 2y agoEvery time I think I'm sorta getting somewhere in my understanding of how to write code I see a comment like this that reminds me that the rabbithole is functionally infinite in both breadth and depth. There's simply no straightforward default approach that won't have you running into and thinking through the most esoteric sounding problems. I guess that's half the fun!
- mrkeen 2y agoIt's not that bad. We just don't have the equivalent of GC for multi-threading yet, so the advice necessarily needs to be "just remember to take and release locks" (same as remembering to malloc and free). Hopefully someone will invent something like STM [1] in the distant year of 2007 or so [2]. It has actual thread-safe data structures. Not just the current choice between wrong-answer-if-you-dont-lock and insane-crashing-if-you-dont-lock. [1] https://www.adit.io/posts/2013-05-15-Locks,-Actors,-And-STM-In-Pictures.html https://www.adit.io/posts/2013-05-15-Locks,-Actors,-And-STM-... [2] https://youtu.be/4caDLTfSa2Q?feature=shared https://youtu.be/4caDLTfSa2Q?feature=shared
- LegionMammal978 2y agoRust takes pride in its 'fearless concurrency' (strict compile-time checks to ensure that locks or similar constructs are used for cross-thread data, alongside the usual channels and whatnot), while Go takes pride in its use of channels and goroutines for most tasks. Not everything is like the C/C++/C#/Java situation where synchronization constructs are divorced from the data they're responsible for.
- neonsunset 2y agoSynchronization primitives in Go are just as divorced as elsewhere, sometimes even more so - it does have channels, but Goroutines cannot yield a value, forcing you to employ a separate storage location together with WaitGroup/Mutex/RWMutex (which, unlike Rust's RWLock, is separate too, although C# lets you model it to an extent). This results in community developing libraries like https://github.com/sourcegraph/conc https://github.com/sourcegraph/conc which attempt to replicate Rust's Futures / C#'s Tasks.
- sunshowers 2y agoThe usual issue is code evolution over time, not the initial version which tends to be okay. You really want to have tooling strictly enforce invariants, and do so in a way that fails closed rather than open. In other words, use Rust.
- foobarian 2y agoI ran into my share of concurrency bugs, but one thing I could never intentionally trigger was any kind of inconsistency stemming from removing a "volatile" modifier from a mutable field in Java. Maybe the JVM I tried this with was just too awesome.
- hashmash 2y agoWere you only testing on x86 or any other "total store order" architecture? If so, removing the volatile modifier has less of an impact.
- bob1029 2y agoI've universally found that even when I am convinced that I am OK with the consequences of sharing something that isn't synchronized, the actual outcome is something I wasn't expecting.
- loeg 2y agoThe only things that should be shared without synchronization are readonly objects where the initialization is somehow externally serialized with accessors, and atomic scalars -- C++ std::atomic, Java has something similar, etc.
- ivanjermakov 2y agoSome (maybe most?) operations on Java Collections perform integrity checks to warn about such issues, for example map throwing ConcurrentModificationException
- smarks 2y agoConcurrentModificationException is typically thrown from an iterator when it detects that it’s been invalidated by a modification to the underlying collection. It’s harder to check for the case described in this article, which is about multiple threads calling put() concurrently on a non thread safe object.
- kachapopopow 2y agoConcurrentModificationException does not check threads, it triggers when it is already too late. It also triggers on the same thread if you remove while iterating an iterator
- saagarjha 2y agoThis is kind of a hot take but I actually prefer debugging races in C/C++ for this reason. Yes, the language prescribes insane semantics (basically none) when it happens, but in practice you’ll get memory corruption or other noisy issues pretty often, and the fact that races are mostly illegal means you can write something like thread sanitizer without needing source code changes to indicate semantics. Meanwhile in Java you’ll never have UB but often you’ll have two fields be subtly out of sync and it’s a lot harder to track this kind of thing down.
- lucianbr 2y ago> I always thought of race conditions as corrupting the data or deadlocking. I never though it could cause performance issues. But it makes sense, you could corrupt the data in a way that creates an infinite loop. Food for thought. I often think to myself that any error or strange behavior or even warnings in a project should be fixed as a matter of principle, as they could cause seemingly unrelated problems. Rarely is this accepted by whoever chooses what we should work on.
- cryptonector 2y agoNot everyone agrees. The SQLite team, for example, famously refuses to fix warnings reported by users. I myself do try to fix most warnings -- some are just false positives.
- swatcoder 2y ago> Rarely is this accepted by whoever chooses what we should work on. You need to find more disciplined teams. There are still people out there who care about correctness and understand how to achieve it without it being an expensive distraction. It a team culture factor that mostly just involves staying on top of these concerns as soon as they're encountered so there's not some insurmountable and inscrutable backlog that makes it feel daunting and hopeless or that makes prioritization difficult.
- saulpw 2y agoMost teams are less disciplined than they should be. Also, job/team mobility is very low right now. So the question becomes, how do you increase discipline on the team you're on?
- rapind 2y agoFor very small teams, exploring new platforms and / or languages that compliment correctness is an option. Using a statically typed language with explicit managed side effects has made a huge difference for me. Super disruptive the larger the team though of course.
- hyperhello 2y agoShould I read this as the Java TreeMap itself is thread unsafe, and the JVM is in a loop, or that the business logic itself was thread unsafe, and just needed to lock around its transactions?
- otterley 2y agohttps://docs.oracle.com/en/java/javase/21/docs/api/java.base/java/util/TreeMap.html https://docs.oracle.com/en/java/javase/21/docs/api/java.base... "Note that this implementation is not synchronized. If multiple threads access a map concurrently, and at least one of the threads modifies the map structurally, it must be synchronized externally. (A structural modification is any operation that adds or deletes one or more mappings; merely changing the value associated with an existing key is not a structural modification.) This is typically accomplished by synchronizing on some object that naturally encapsulates the map. If no such object exists, the map should be "wrapped" using the Collections.synchronizedSortedMap method. This is best done at creation time, to prevent accidental unsynchronized access to the map..."
- layer8 2y agoTo add to this: Java’s original collection classes (Vector, Hashtable, …) were thread-safe, but it turned out that the performance penalty for that was too high, all the while still not catching errors when performing combinations of operations that need to be a single atomic transaction. This was one of the motivations for the thread-unsafe classes of the newer collection framework (ArrayList, HashMap, …).
- wging 2y ago> still not catching errors when performing combinations of operations that need to be a single atomic transaction This is so important. The idea that code is thread-safe just because it uses a thread-safe data structure, and its cousin "This is thread-safe, because I have made all these methods synchronized" are... not frequent, but I've seen them expressed more often than I'd like, which is zero times.
- bsder 2y agoDoes a ConcurrentSkipListMap not give the correct O(log N) guarantees on top of being concurrency friendly? java.util.concurrent is one of the best libraries ever. If you do something related to concurrency and don't reach for it to start, you're gonna have a bad time.
- kccqzy 2y agoIt's often a better design if you manage the concurrency in a high-level architecture and not have to deal with concurrency at the data structure level.
- bsder 2y agoDesigning your own concurrency structures instead of using ones designed by smart people who thought about the problem more collective hours than your entire lifetime is unwarranted hubris. The fact that ConcurrentTreeMap doesn't exist in java.util.concurrent should be ringing loud warning bells.
- mrkeen 2y agoOnce they expose data structures that allow generic uses like "if size<5, add item" I'll take another look. Until then, their definition of thread-safety isn't quite the same as mine.
- kccqzy 2y agoYou can take a look at Haskell's Software Transactional Memory then. Or you can take a look at something like the Linux kernel's Read-Copy-Update (RCU) abstraction, add some persistent data structures and a retry loop on top. It's indeed a very programmer friendly way of doing concurrency.
- layer8 2y agoThe GP comment is not about designing your own concurrency data structures. It’s about the fact that if your higher-level logic doesn’t take concurrency into account, using the concurrent collections as such will not save you from concurrency bugs. A simple example is when you have two collections whose contents need to be consistent with each other. Or if you have a check-and-modify operation that isn’t covered by the existing collection methods. Then access to them still has to be synchronized externally. The concurrent collections are great, but they don’t save you from thinking about concurrent accesses and data consistency at a higher level, and managing concurrent operations externally as necessary.
- mvc 2y agoHaha, I was recently running a backfill was quite pleased when I managed to get it humming along at 6400% CPU on a 64vcpu machine. Fortunately ssh was still receptive.
- layer8 2y agoAnother way to get infinite loops is using a Comparator or Comparable implementation that doesn’t implement a consistent total order: https://stackoverflow.com/questions/62994606/concurrentskipset-compareto-instant-infinite-loop https://stackoverflow.com/questions/62994606/concurrentskips... (This is unrelated to concurrency.) Whether it occurs or not can depend on the specific data being processed, and the order in which it is being processed. So this can happen in production after seemingly working fine for a long time.
- TYMorningCoffee 2y agoHave you seen this before in person? It would make a great blog post. I haven't personally encountered a buggy comparator without a total order.
- layer8 2y agoI have seen a lot of incorrect Comparators and Comparable implementations in existing code, but haven’t personally come across the infinite-loop situation yet. To give one example, a common error is to compare two int values via subtraction, which can give incorrect results due to overflow and modulo semantics, instead of using Integer::compare (or some equivalent of its implementation).
- smarks 2y agoInteresting. I haven’t seen an infinite loop either, but I can imagine one if a comparator tries to be too “clever” for example if it bases its comparison logic on some external state. Another common source of comparator bugs is when people compare floats or doubles and they don’t account for NaN, which is unequal to everything, including itself! In Java, the usual symptom of comparator bugs is that sort throws the infamous “Comparison method violates its general contract!” exception.
- hollowcelery 2y agoI knew someone who missed out on a gold medal at the International Olympiad of Informatics because his sort comparator didn’t have a total order.
- w10-1 2y agoVery well said, and very nice to see references to others on point. As a sidebar: I'm almost distracted by the clarity. The well-formed structure of this article is a perfect template for an AI evaluation of a problem. It'd be interesting to generate a bunch of these articles, by scanning API docs for usage constraints, then searching the blog-sphere for articles on point, then summarizing issues and solutions, and even generating demo code. Then presto! You have online karma! (and interviews...) But then if only there were a way to credit the authors, or for them to trace some of the good they put out into the world. So, a new metric: PageRank is about (sharing) links-in karma, but AuthorRank is partly about the links out, and in particular the degree to which they seem to be complete in coverage and correct and fair in their characterization of those links. Then a complementary page-quality metric identifies whether the page identifies the proper relationships between the issues, as reflected elsewhere, as briefly as possible in topological order. Then given a set of ordered relations for each page, you can assess similarity with other pages, detect copying (er, learning), etc.
- moffkalast 2y agoAnyone else mildly peeved by how CPU load is just load per core summed up to an arbitrary percentage all too often? Why not just divide 100% by number of cores and make that the max, so you don't need to know the number of cores to know the actual utilization? Or better yet, have those microcontrollers that Intel tries to pass off as E and L cores take up a much smaller percentage to fit their general uselessness.
- mystifyingpoi 2y agoIDK but the current convention makes it easy to see single-threaded bottlenecks. So if my program is using 100% CPU and cannot go faster, I know where to look.
- ahoka 2y agoThis is “Irix” vs “Solaris” mode of counting, the latter being summed up to to 100% for all cores. I think the modern approach would be to see how much of its TDP budget the core is using.
- shrx 2y agoIf you think that's confusing, let me introduce you to CPU load average values on android... https://stackoverflow.com/questions/10829546/how-to-read-the-stock-cpu-usage-data https://stackoverflow.com/questions/10829546/how-to-read-the...
- jonathanlydall 2y agoWhat about spotting a cycle by using an incrementing counter and then throwing an exception if it goes above the tree depth or collection size (presuming one of these is tracked)? Unlike the author’s hash set proposal it would require almost no memory or CPU overhead and may be more likely to be accepted. That being said, in the decade plus I’ve used C# I’ve never found that I failed to consider concurrent access on data structures in concurrent situations.
- saagarjha 2y agoIt’s not a horrible idea but adding useful preconditions in the face of data races is quite hard in general. This is just one of the ways the tree can break.
- TYMorningCoffee 2y agoThat's much better. Constant memory. The number of nodes is guaranteed to be less than or equal to the height of the tree.
- exDM69 2y agoYes, this is a good idea. I've done this before with binary search and tree structures. I avoid unbounded loops wherever possible and the unavoidable cases are rather rare. It's not a fix but it is a good mitigation strategy. Infinite loops are one of the nastiest kind of bugs. Although they are easy to spot in a debugger they can have really unfriendly consequences like OP's "can barely ssh in" situation. Infinite loops in library code are particularly unpleasant.
- lolc 2y agoWhat I like here is the discovery of the extra loop, and then still digging down to discover the root cause of the competing threads. I think I would have removed the loop and called it done.
- lcfcjs6 2y agoJava is usually shockingly inefficient.
- deskr 2y agoExceptions in threads are an absolute killer. Here's a story of a horror bughunt where the main characters are C++, select() and thread brandishing an exception: https://news.ycombinator.com/item?id=42532979 https://news.ycombinator.com/item?id=42532979
- TYMorningCoffee 2y agoI remember reading that article but being unable to understand it due to my lack of knowledge in the area. I will have to give it another go.
- loeg 2y agoJust to probe the code review angle a little bit: shared mutable state should be a red/yellow flag in general. Whether or not it is protected from unsafe modification. That should jump out to you as a reviewer that something weird is happening.
- svilen_dobrev 2y agoi found this article very/deeply informative , memory-wise , concurency vs optimizations and troubles thereof: Programming Language Memory Models https://research.swtch.com/plmm https://research.swtch.com/plmm https://news.ycombinator.com/item?id=27750610 https://news.ycombinator.com/item?id=27750610
- Someone 2y agoFTA: The code can be reduced to simply: public void someFunction(SomeType relatedObject, List<SomeOtherType> unrelatedObjects) { ... treeMap.put(relatedObject.a(), relatedObject.b()); ... // unrelatedObjects is used later on in the function so the // parameter cannot be removed } That’s not true. The original code only does the treeMap.put if unrelatedObjects is nonempty. That may or may not be a bug. You also would have to check that a and b return the same value every time, and that treeMap behaves like a map. If, for example, it logs updates, you’d have to check that changing that to log only once is acceptable.
- TYMorningCoffee 2y agoGood point. It should be replaced with an if not empty check.
- hinkley 2y agoThe author has discovered a flavor of the Poison Pill. More common in event sourcing systems, it’s a message that kills anything it touches, and then is “eaten” again by the next creature that encounters it which also dies a horrible death. Only in this case it’s live-locked. Once the data structure gets into the illegal state, every subsequent thread gets trapped in the same logic bomb, instead of erroring on an NPE which is the more likely illegal state.
- thinkingemote 2y ago"I could barely ssh onto it" Is there a way to ensure that whatever happens (CPU, network overloaded etc) one can always ssh in? Like reserve a tiny bit of stuff to the ssh daemon?
- LtdJorge 2y agoNice? Or maybe give the Systemd slice a special cgroup with a reservation.
- Aloisius 2y agoI'd consider doing the inverse and nice the JVM with a lower priority instead in certain situations.
- beisner 2y agoOn Linux I’ve done this by pinning processes to a certain range of CPU cores, and the scheduler will just keep one core free or something. Which allows whatever I need in terms of management to execute on that one core, including SSH orUI.
- homebrewer 2y agoCreate a systemd override (by using systemctl edit sshd) and add MemoryMin=32M (or whatever makes sense for your system). This makes sure sshd is never pushed out into swap. https://www.freedesktop.org/software/systemd/man/latest/systemd.resource-control.html#MemoryMin=bytes,%20MemoryLow=bytes https://www.freedesktop.org/software/systemd/man/latest/syst... You can also use sibling knobs to increase the CPU and IO weights of the unit, for example but setting this to something higher than 100: https://www.freedesktop.org/software/systemd/man/latest/systemd.resource-control.html#CPUWeight=weight https://www.freedesktop.org/software/systemd/man/latest/syst...
- masklinn 2y agoIs there a way to do that for alternate tty? From time to time I’ll run something dumb on my machine (e.g. GC aggressive the wrong repo) and if I don’t catch the ramp up the machine will lock up until the oom killer find the right process. Sufficiently locked up, accessing alternate ttys to kill the offending process won’t work either. I guess I could just reserve ssh then ssh into it from an other computer but…
- neonsunset 2y agoIn practice, it's rarely an issue in C# because it offers excellent concurrent collections out of box, together with channels, tasks and other more specialized primitives. Worst case someone just writes a lock(thing) { ... } and calls it a day. Perhaps not great but not the end of the world either. I did have to hand-roll something that partially replicates Rust's RWLock<T> recently, but the resulting semantics turned out to be decent, even if not providing the exact same level of assurance.
- xyst 2y agoAt least it was using all of the cores. The CPU running this application was cooking.
- procaryote 2y agoTL;DR; don't use thread unsafe data structures from multiple threads at once
- scottlamb 2y ago> Could an unguarded TreeMap cause 3,200% utilization? I've seen the same thing with an undersynchronized java.util.HashMap. This would have been in like 2009, but afaik it can still happen today. iirc HashMap uses chaining to resolve collisions; my guess was it introduced a cycle in the chain somehow, but I just got to work nuking the bad code from orbit [1] rather than digging in to verify. I often interview folks on concurrency knowledge. If they think a data race is only slightly bad, I'm unimpressed, and this is an example of why. [1] This undersynchronization was far from the only problem with that codebase.
- TYMorningCoffee 2y agoOh no I didn't know this can happen with HashMaps too! Something to do with the linked list they use for collisions?
- vineet_rc 2y ago[dead]
- scottlamb 2y agoThere's a dead comment replying to yours with this link: https://mailinator.blogspot.com/2009/06/beautiful-race-condition.html https://mailinator.blogspot.com/2009/06/beautiful-race-condi... ...which I don't think I've ever seen before, but nicely explains the problem I saw, and coincidentally is from the same year!
- MadVikingGod 2y agoI was excited to see that not only does this article cover other languages, but that this error happens in go. I was a bit surprised because map access is generally protected by a race detector, but the RedBlack tree used doesn't store anything in a map anywhere. I wonder if the full race detector, go run -race, would catch it or not. I also want to explore if the RB tree used a slice instead of two different struct members if that would trigger the runtime race detector. So many things to try when I get back to a computer with go on it.
- doctor_phil 2y agoWhy does the fix need to remember all the nodes we have visited? Can't we just keep track of what span we are in? That way we just need to keep track of 2 nodes. In the graphic from the example we would keep track like this: low: - high - low: 11 high: - low: 23 high: - low: 23 high: 26 Error: now we see item 13, but that is not inside our span!
- johnklos 2y agoDoes it not strike anyone else as odd that if someone said they had a single CPU, and that CPU were running a normal priority task at 100%, and that caused the machine to barely allow ssh, we'd say there's a much bigger problem than that someone is letting on? No 32 core (thread, likely) machine should ever normally be in a state where someone can "barely ssh onto it". Is Java really that janky? Or is "barely ssh onto it" a bit hyperbolic?
- sc68cal 2y agoI've absolutely experienced situations like this where the system is maxed out and ssh becomes barely usable. The problem is that all processes run with pretty much the same priority level, so there's no differentiation in the scheduler between interactive ssh sessions and other programs that are consuming all the CPU, and most of the time you only realize this fact far after the point where you can SSH in and renice the misbehaving processes.
- spockz 2y agoMy macOS install got into a state before the Christmas holiday where something would spawn #corecount amount of process. They seemed to be related to processing thumbnails for contents of directories. This would go on for a few minutes in which the machine was completely unresponsive including the mouse. So it seems possible still to bring a high core count machine to its knees. But something then is indeed very wrong.
- macspoofing 2y agoOoof. The core collections in Java are well understood to not be thread-safe by design, and this should have been noticed. OP should go through the rest of the code and see if there are other places where collections are potentially operated by multiple threads. >The easiest way to fix this was to wrap the TreeMap with Collections.synchronizedMap or switch to ConcurrentHashMap and sort on demand. That will make the individual map operations thread-safe, but given that nobody thought of concurrency, are you sure series of operations are thread-safe? That is, are you sure the object that owns the tree-map is thread-safe. I wouldn't bet on it. >Controversial Fix: Track visited nodes Don't do that! The collection will still not be thread-safe and you'll just fail in some other more subtle way .. either now, or in the future (if/when the implementation changes in the next Java release). >Sometimes, a detail oriented developer will notice the combination of threads and TreeMap, or even suggest to not use a TreeMap if ordered elements are not needed. Unfortunately, that didn’t happen in this case. That's not a good take-away! OP, the problem is you violating the contract of the collection, which is clear that it isn't thread-safe. The problem ISN'T the side-effect of what happens when you violate the contract. If you change TreeMap to HashMap, it's still wrong (!!!), even though you may not get the side-effect of a high CPU utilization. --------- When working with code that is operated on by multiple threads, the only surefire strategy I found was to to make every possible object immutable and limiting any object that could not be made immutable to small, self-contained and highly controlled sections. We rewrote one of the core modules following these principles, and it went from being a constant source of issues, to one of the most resilient sections of our codebase. Having these guidelines in place, also made code reviews much easier.
- OskarS 2y ago> That will make the individual map operations thread-safe, but given that nobody thought of concurrency, are you sure series of operations are thread-safe? That is, are you sure the object that owns the tree-map is thread-safe. I wouldn't bet on it. This is a very important point! And, in my experience, a common mistake by programmers who aren't great at concurrency. Lets say you have a concurrent dynamic array. The array class is designed to be thread-safe for the explicit purpose that you want to share it between threads. You want to access element 10, but you want to be sure that you're not out of bounds. So you do this: if (array.size() > 10) { array.element(10).doSomething(); } It doesn't matter how thread-safe this array class is: these two operations combined are NOT thread-safe, because thread-safety of the class only means that `.size()` and `.element()` are on their own not going to cause any races. But it's entirely possible another thread removes elements in between you checking the size and accessing the element, at which point you'll (at best) get an out-of-bounds crash. The way to fix it is to either use atomic methods on the class which does both (something like `.element_or_null()` or whatever), or to not bother with a concurrent dynamic array at all and instead just use regular one you guard with a mutex (so the mutex is held during both operations and whatever other operations other threads perform on the array).
- opentokix 2y agoclick Java close
- tasty_freeze 2y agoyou missed the part where he reproduced it in a number of other languages, and some where he was unable to reproduce it.
- hn_acc1 2y agoThe mention of "could barely ssh in" reminds me of a situation in grad school where our group had a Sun UltraSparc 170 (IIRC) with 1GB HD and 128 or 256 MB of RAM, shared by maybe 8 people in a small research group relating to parallel and distributed processing. Keep in mind, Sun machines were rarely rebooted, ever. So I guess the new user / student was trying to do things in parallel to speed things up when they chopped up their large text file into N (32 or 64) sections based on line number (not separate files), and then ran N copies of perl in parallel, each processing its own set of lines from that one file. Not only did you have a large amount (for back then) of RAM used by N copies of the perl interpreter (separate processes, not threads, mind you!) processing its data, as well, any attempt to swap was interleaved with frantic seeking to a different section of the same file to read a few more lines for one of N processes stalled on IO. Also, probably the Jth process had to read J/N of the entire file to get to its section. So the first section of the file was read N times, the next N-1, then N-2, etc. We (me and the fellow-student-trusted-as-sysadmin who had the root password) couldn't even get a login prompt on the console. Luckily, I had a logged-in session (ssh from an actual x-terminal - a large-screen "dumb" terminal), and "su" prompted for a password after 20-30 minutes of running it. After another 5-10 minutes, we had a root session and were able to run top and figure out what was going on. Killing the offending processes (after trying to contact the user) restored the system back to normal. Edit: forgot to say: had the right idea, but totally didn't understand the system's limitations. It was SEVERELY I/O limited with that hard drive and relatively low RAM, so just processing the data linearly would have easily been the best approach unless the amount of data to be kept would have gotten too large.
- aqueueaqueue 2y agoSerious though. What 100% means on cpu dashboards is as inconsistent as what time zone dates are in.
- krick 2y agoWow, that's pretty high-effort blog post (taking time to compare across 12 mainstream languages).
- memoryfault 2y agoThis is a fairly common bug in multithreaded use of .NET Dictionary, too.
- fulafel 2y ago> A while back my machine was so messed up that I could barely ssh onto it. 3,200% CPU utilization - all 32 cores on the host were fully utilized! You wouldn't notice the load from 32 cpu pegging threads or processes on a 32-core host when ssh'ing in. Sounds like the OP is maybe leaving out what the load on the machine was, maybe it was more like thousands of spinning threads?
- kristianp 2y agoHas anyone noticed an inability to horizontally scroll the code samples on chrome Android phones? I've noticed it on a few different blogs. The window has to be dragged at lower point on the screen to scroll the code section.
- TYMorningCoffee 2y agoYes. Sorry about that. Will look into it when I have time.
- TYMorningCoffee 2y agoI think I tracked it down to the table in the article adding a scroll bar for the entire page. This scrollbar fights with the code blocks scrollbar. I will try to refactor the table to fit in a phone width somehow.
- TYMorningCoffee 2y agoFixed! It was because the table overflowed. This created a scroll bar for the entire page. The entire page scrollbar had a weird interaction with the code block scroll bars. Now that there is no more whole page scroll bar, it works.
- kristianp 2y agoYep, works fine now. Thanks!
- jchw 2y agoSomewhat relatedly, I've always appreciated Go adding some runtime checks to detect race conditions, like concurrent map writes. It's not as good as having safe concurrency from the ground up, but on the other hand it's a lot better than not having detection, as everyone makes mistakes from time to time, and imperfect as it may be, the detection usually does catch it quickly. Especially nice since it is on in your production builds; a big obstacle with a lot of debugging tools is they're hard to get them where you need them...
- khana 2y ago[dead]
- xandrius 2y agoDoes CPU utilisation calculations work this way? - 1 CPU at 100% = 100% - 10 CPUs at 100% = 100% - 1 CPU at 100% + 9 at 0% = 10% Is that not right? Or is CPU utilisation not usually normalised?
- TYMorningCoffee 2y agoIt depends which OS. Windows yes, Linux no.
- bmm6o 2y agoI've seen this in production C#. Same symptoms, process dump showed a corrupt dictionary. We thought we were using ConcurrentDictionary everywhere but this one came in from a library. We were using .net framework, IIRC .net core has code to detect concurrent modification. I don't know how they implement it but it could be as simple as a version counter. It's odd he gets so hung up on npe being a critical ingredient when that doesn't appear to be present in the original manifestation. There's no reason to think you can't have a bug like this in C just because it doesn't support exceptions. To me it's all about class invariants. In general, they don't hold while a mutator is executing and will only be restored at the end. Executing another mutator before the invariants are reestablished is how you corrupt your data structure. If you're not in a valid state when you begin you probably won't be in a valid state when you finish.
- TYMorningCoffee 2y agoIt came down to poor logic. I got hung up on it because I got unlucky and couldn't reproduce it with uncaught NPE so I incorrectly concluded that uncaught NPE was a necessary condition.
- cytocync 2y ago[dead]
- yburkov 2y agojunior level error, dont understand whats to talk here about... Use ConcurrentSkipListMap and never write our own concurrent data structure unless you Doug Lea
- NovaX 2y agoHow about the senior engineer level error of forking Doug Lea's concurrent data structures only to make them operate in worst-case time complexity? Found this doozy recently. [1] https://gist.github.com/ben-manes/6312727adfa2235cb7c5e25cae523ad0 https://gist.github.com/ben-manes/6312727adfa2235cb7c5e25cae...