10 ms·
Mimalloc Cigarette: Losing one week of my life catching a memory leak (Rust)
- kibwen 2y agoLevel 1 systems programmer: "wow, it feels so nice having control over my memory and getting out from under the thumb of a garbage collector" Level 2 systems programmer: "oh no, my memory allocator is a garbage collector"
- seanthemon 2y agoAt the very bottom of everything is a garbage collector..
- amelius 2y agoLevel 3 system programmer: "get me out of this straight jacket and give me my garbage collector back so I can get stuff done"
- forrestthewoods 2y agoNo. Just no. For as painful as the debugging story was I have spent vastly more amounts of time working around garbage collectors to ship performant code.
- 0x457 2y agoWhat, you don't like doing GC only N requests (ruby web servers), disabling GC completely during working hours (java stock trading), fake allocating large buffers (go's allocate and don't use trick)?
- pton_xd 2y agoAin't nothin' wrong with configuring V8 to have unbounded heap growth, disabling the memory reducer, and then killing the process after a while.
- mike_hearn 2y agoThe Java shops you're thinking of didn't disable GC during working hours, they just sized the generations to avoid a collection given their normal allocation rates. But there were / are also plenty of trading shops that paid Azul for their pauseless C4 GC. Nowadays there's also ZGC and Shenandoah, so if you want to both allocate a lot and also not have pauses, that tech is no longer expensive.
- 0x457 2y ago> The Java shops you're thinking of didn't disable GC during working hours, they just sized the generations to avoid a collection given their normal allocation rates. Well, I just trivialized it. However, in one case in mid 00s, I saw it disabled completely to avoid any pauses during trading hours.
- wpollock 2y agoUsed to do this in C decades ago. Worked on Unix but I doubt it works on Linux today, unless you disable memory overcommit completely.
- neonsunset 2y agoI'd wager it was an issue with the language of choice (or its GC) being rather poorly made performance-wise or a design that does not respect how GC works in the first place :)
- ComputerGuru 2y agoThat's not how system programmers think..
- amelius 2y agoA new generation of system programmers is tired of solving the same old boring memory riddles over and over again and no borrow checker is going to help them because it only brings new riddles.
- __s 2y agogc replaces riddles with punchlines
- troutwine 2y agoI agree. If we were to try and pin a thought process to an additional level of systems programmer it’d involve writing an allocator that’s custom to your domain. The problem with garbage collection for the systems’ case is you’re opting into a set of undefined and uncontrolled runtime behavior which is okay until it catastrophically isn’t. An allocator is the same but with less surface area and you can swap it at need.
- amelius 2y agoMeanwhile an OS uses the filesystem for just about everything and it is also a garbage collected system ... Why should memory be different?
- troutwine 2y agoI'm not tracking how your question follows. If by garbage collection you mean a system in which resources are cleaned up at or after the moment they are marked as no longer being necessary then, sure, I guess I can see a thread here, although I think it a thin connection. The conversation up-thread is about runtime garbage collectors which are a mechanism with more semantic properties than this expansive definition implies and possessing an internal complexity that is opaque to the user. An allocator does have the more expensive definition I think you might be operating with, as does a filesystem, but it's the opacity and intrinsic binding to a specific runtime GC that makes it a challenging tool for systems programming. Go for instance bills itself as a systems language and that's true for domains where bounded, predictable memory consumption / CPU trade-offs are not necessary _because_ the runtime GC is bundled and non-negotiable. Its behavior also shifts with releases. A systems program relying on an allocator alone can choose to ignore the allocator until it's a problem and swap the implementation out for one -- perhaps custom made -- that tailors to the domain.
- matklad 2y agoThe answer is clear: just don’t have a malloc implementation in your process' address space!
- poikroequ 2y agoA bump allocator is all anyone really needs
- thebruce87m 2y agoWelcome to embedded! It’s no heaps of fun!
- hinkley 2y ago> no heaps Angry upvote
- eschneider 2y agoI'm always surprised how much I don't miss dynamic allocation. :)
- ckocagil 2y ago"stackoverflow please help me how do i fix memory fragmentation"
- loeg 2y agoSort of tl;dr: mimalloc doesn't actually free memory in a way that it can be reused on threads other than the one that allocated it; the free call marks regions for eventual delayed reclaim by the original thread. If the original thread calls malloc again, those regions are collected (1/N malloc calls). Or (C) you can explicitly invoke mi_collect[1] in the allocating thread (the Rust crate does not seem to expose this API). [1]: https://github.com/microsoft/mimalloc/blob/dev/src/heap.c#L180 https://github.com/microsoft/mimalloc/blob/dev/src/heap.c#L1...
- Arnavion 2y agoThe mimalloc crate just provides the GlobalAlloc impl that can be registered with libstd as the global allocator using the `#[global_allocator]` attr. The underlying sys crate provides the binding for mimalloc API like `mi_collect`: https://docs.rs/libmimalloc-sys/0.1.39/libmimalloc_sys/fn.mi_collect.html https://docs.rs/libmimalloc-sys/0.1.39/libmimalloc_sys/fn.mi...
- hinkley 2y agoWe had learned helplessness on a drag and drop bug in jquery UI. I had like three hours every second or third Friday and would just step through the code trying to find the bug. That code was so sketchy the jquery team was trying to rewrite it from scratch one component at a time, and wouldn’t entertain any bug discussions on the old code even though they were a year behind already. After almost six months, I finally found a spot where I could monkey patch a function to wrap it with a short circuit if the coordinates were out of bounds. Not only fixed the bug but made drag and drop several times faster. Couldn’t share this with the world because they weren’t accepting PRs against the old widgets. I’ve worked harder on bug fixes, but I think that’s the longest I’ve worked on one.
- giancarlostoro 2y agoOne of my favorite most elusive bugs was a one liner change. I didn't understand the problem because nobody could reproduce it, or show it. Months later, after my boss told his boss it was fixed, despite never being able to test that it was fixed, I figured it out and fixed it. We had a gift card form, and we stored it in localStorage, if for any reason the person left the tab, and came back months later, it would show the old gift card with its old dated balance, it was a client-side bug. The fix was to use sessionStorage.
- contingencies 2y agoIt seems in the context of your story the old adage that organizations reproduce software in their own architecture again rings true, with multilayered bureaucracy, lies and promises resulting in "client state".
- giancarlostoro 2y agoWhen I tried to explain to him that I fixed the thing he claimed to have fixed, I heard him hesitantly say it wasn't the same bug. Not sure what he told his boss this time around the fix was for, but I was able to fully reproduce the bug with this fix. If you can't reproduce a bug, you cannot in my opinion say that it is fixed. If you have to reproduce it via local debugging and changing a value, or hard coding a value, I think you're possibly close, but there's a chance it might not be the case!
- IceTDrinker 2y agoPSA: do not use floating point for monetary amounts
- SAI_Peregrinus 2y agoMS Excel uses floating point, and it's used a ton in finance. Don't use floating-point for monetary amounts if you don't know what rounding mode you've set.
- koverstreet 2y agoIt's somewhat acceptable with double precision floats - never single precision floats. But far better to just use integer cents.
- IX-103 2y agoInteger cents implies a specific rounding mode (truncation). That's probably not what you should be using. Floating point cents gets the best of both worlds (if you set the right rounding mode).
- nurettin 2y agoI have used single precision floats in my latest project just to disprove this baloney.
- smh 2y agoYou are using 32 bit floats to represent money? Does your project correctly calculate $300,000.00 + $0.01, (or even just correctly represent the value $300,000.01) and if so, how?
- nurettin 2y agoObviously you can't accumulate cent by cent. You can't even safely accumulate by quarter. Epsilon is too large to do that. I calculate cumulative pnl using std::fma, then multiply AUM with that and round to cents. It's good enough for backtesting, and it shaves a bunch of seconds off the clock.
- znpy 2y ago> Allocators have different characteristics for a reason - they do some things differently between each other. What do you think mimalloc does that could account for this behavior? Interestingly, it would seem that Java programmers play with garbage collectors while Rust programmers play with memory allocators.
- sbt567 2y ago> Rust programmers *system
- deleted 2y ago[deleted]
- Exuma 2y agoI really love the design of this blog
- zokier 2y agoI wonder if there is something that could be done on language design level to have better "sympathy" to memory allocation, i.e. built upon having mmap/munmap as primitives instead of malloc/free; where language patterns are built around allocating pages instead of arbitrarily sized objects. Probably not practical for general high-level languages, but for e.g. embedded or high-performance stuff might make sense?
- eschneider 2y agoIn general for embedded, you don't page memory even if you're running something like embedded linux. For high performance stuff where you need low, predictable latency, you're probably not going to want to use dynamic memory at all.
- loeg 2y agoNot exactly what you're getting at, but you could maybe imagine an explicit version of malloc where allocations are destined either for thread-local only use, or shared use. Then locally freeing remote thread-local memory is an invalid operation and these kinds of assume-locality optimizations are valid on many structures. I think you can imagine a version of mmap that allows for thread-local mappings to help detect accidental misuse of local allocation.
- dathinab 2y agomost modern memory allocators use internally mmap, this is why it most times makes sense to not use the system allocate for long running programs Generally given that page size isn't something you know at compiler (or even install size) and it can vary between each restart and it being between anything between ~4KiB and 1GiB and most natural memory objects being much less then 4KiB but some being potentially much more then 1GiB you kind don't want to leak anything related to page sizes into your business logic if it can be helper. If you still need to most languages have memory/allocation pools you can use to get a bit more control about memory allocation/free and reuse. Also the performance issues mentioned have not much to do with memory pages or anything like that _instead they are rooted in concurrency controls of a global resource (memory)_. I.e. thread local concurrency syncronization vs. process concurrency synchronization. mainly instead of using a fully general purpose allocator they used an allocator whiche is still general purpose but has a design bias which improves same-thread (de)allocation perf at cost of cross thread (de)allocation perf. And they where doing a ton of cross thread (de)allocations leading to noticeable performance degradation. The thing is even if you hypothetically only had allocations at sizes multiple of a memory page or use a ton of manual mmap you still would want to use a allocator and not always directly free freed memory back to the OS as doing so and doing a syscall on every allocation tends to lead to major performance degradation (in many use cases). So you still need concurrency controls but they come at a cost, especially for cross thread synchronization. Even just lock-free controls based on atomic have a cost over thread local controls caused often largely by cache invalidation/synchronization.
- akira2501 2y ago[flagged]
- Patryk27 2y agoNot sure why all the hostility here - you haven't seen the code, know nothing about the domain and yet you're certain that our performance requirements are false and that "code base got sacrificed" (apparently by adding two lines of rather self-explanatory code?) Feels like you've just read grugbrain.dev and decided to shoot your golden tips at everybody without actually trying to understand the situation. Anyway, there's one good point here: > why is your allocator in this path, then? Because those prices change 24/7/365, million times a day, and so refreshing happens pretty much all the time in the background, eating CPU time. What's more, calculating prices is much more complicated than a hashmap lookup - hotels can have dynamic number of discounts, taxes etc., and they can't all be precomputed (too many combinations). You know, not all complexity is made up, a little trust in others won't hurt.
- urbandw311er 2y agoOut of interest, is there no way to rearchitect the whole thing to be event-based, ie more like a producer-consumer situation? Or do you have to loop back through all the source data and poll every hotel to fetch its current prices?
- Patryk27 2y agoSure, it is producer-consumer - we use Postgres' LISTEN/NOTIFY mechanism (mostly because we have no other use cases for queueing, so "exploiting" an already existing feature in Pg was easier). The example in article shows all hotels getting refreshed, but that's just because that's the quickest way to reproduce the problem locally. In reality we refresh/reindex only those hotels which have changed since the last refresh - over day(s) this accumulates and the OOM was actually happening not immediately on the first refresh, but after a couple of days (which is part of what made it difficult to catch).
- deleted 2y ago[deleted]
- Arnavion 2y agojemalloc also has its own funny problem with threads - if you have a multi-threaded application that uses jemalloc on all threads except the main thread, then the cleanup that jemalloc runs on main thread exit will segfault. In $dayjob we use jemalloc as a sub-allocator in specific arenas. (*) The application itself is fine in production because it allocates from the main thread too, but the unit test framework only runs tests in spawned threads and the main thread of the test binary just orchestrates them. So the test binary triggers this segfault reliably. ( https://github.com/jemalloc/jemalloc/issues/1317 https://github.com/jemalloc/jemalloc/issues/1317 Unlike what the title says, it's not Windows-specific.) (*): The application uses libc malloc normally, but at some places it allocates pages using `mmap(non_anonymous_tempfile)` and then uses jemalloc to partition them. jemalloc has a feature called "extent hooks" where you can customize how jemalloc gets underlying pages for its allocations, which we use to give it pages via such mmap's. Then the higher layers of the code that just want to allocate don't have to care whether those allocations came from libc malloc or mmap-backed disk file.
- CraigJPerry 2y agoTangent: what’s the ideal data structure for this problem? If there were 20million rooms in the world with a price for each day of the year, we’d be looking at around 7billion prices per year. That’d be say 4Tb of storage without indexes. The problem space seems to have a bunch of options to partition - by locality, by date etc. I’m curious if there’s a commonly understood match for this problem? FWIW with that dataset size, my first experiments would be with SQL server because that data will fit in ram. I don’t know if that’s where I’d end up - but I’m pretty sure it’s where I’d start my performance testing grappling with this problem.
- jrpelkonen 2y agoI think your premise is somewhat off. There might be 20 million hotel rooms in a world, but surely they are not individually priced, e.g. all king bed rooms in a given hotel have the same price per given day.
- om8 2y agoTLDR: use shitty allocators, win shitty memory leaks
- malkia 2y agoIn C++, your https://en.cppreference.com/w/cpp/memory/new/new_handler https://en.cppreference.com/w/cpp/memory/new/new_handler should call mi_collect.
- PaulDavisThe1st 2y agoA perfect demonstration of how many of harder problems we face writing (especially non-browser-based) software are in fact not addressed by language changes. The concept of memory that is allocated by a thread and can only be deallocated by that thread is useful and valid, but as TFA demonstrates, can also cause problems if you're not careful with your overall architecture. If the language you're using even allows you to use this concept, it almost certainly will not protect you from having to get the architecture corect.
- the-smug-one 2y agoI think Rust's language design is in part to blame, as it does not force the programmer to think sufficiently of the layout of the memory, instead allowing them to defer to a "global allocator".
- PaulDavisThe1st 2y agoThis identical problem could easily occur in a C or C++ codebase.
- the-smug-one 2y agoI never said that C and C++ doesn't suffer from the same design problem? I'd say that Zig is the best in class here, typically forcing you to pass along an allocator to each data structure. C is a bit better than C++, as it uses an allocator explicitly, while C++ relies on new/delete with a default impl calling malloc/free. Still a language design issue: C++ and Rust doesn't put allocation concerns front and center, when they very much are. Not encouraging thinking about these things is very bad for systems languages.
- PaulDavisThe1st 2y agoThis issue isn't really about per-structure allocators at all. It's about the idea that you are using per-thread allocators, and one of your threads allocates a lot of memory, then goes to sleep for a long time. Per-thread allocators are orthogonal to per-structure allocators.
- bsder 2y agoWelcome to systems programming. Allocators are invisible--until they aren't.
- rurban 2y agoThe Annotated C++ Reference Manual: “C programmers think memory management is too important to be left to the computer. LISP programmers think memory management is too important to be left to the user.”