21 ms·
Why is D's garbage collection slower than Go's?
- silisili 4y agoI actually find it one of the biggest downfalls of D. IMO, you have to pick one, among other things for what's listed in the article. Riding the fence leads to a worse experience for both sides.
- masklinn 4y ago> Riding the fence leads to a worse experience for both sides. I don't entirely agree, I think having a garbage collected pointer (GCP) can make a lot of sense as a performance optimisation: manual memory management (MMM) and reference counting (RC) have stampeding characteristics similar to GC pauses when releasing large hierarchies e.g. persistent data structures, or large trees of widgets: all the tree is freed synchronously and recursively, if it's large it can be very sensible. This stampeding is more predictable than a GC pause, but it's no less problematic when it occurs, and mitigating it can be difficult (you have to manually shunt the objects you'd like to release off-thread). A GCP can do that shunting on its own, letting a GC thread perform the releases asynchronously, and piecemeal. Furthermore, GCP means you're not expecting to precisely track allocations at the system level, which means the usual tricks (arenas, bump allocators, ...) can be applied by default to GC pointers, where they usually can't be to MMM or RC pointers. This decreases allocation and deallocation overhead by reducing the amount of work the system has to do.
- pjmlp 4y agoHence why other languages do make this differentiation, only D refuses to go down this path.
- zozbot234 4y ago> you have to manually shunt the objects you'd like to release off-thread. A GCP can do that shunting on its own, letting a GC thread perform the releases asynchronously, and piecemeal. You could also use async programming to do the same thing in GC-free code. Or just use an arena that can be freed all at once.
- masklinn 4y ago> You could also use async programming to do the same thing in GC-free code. No? Allocations are blocking, "async programming" won't do anything. Unless you have an async dealloc which does the shunting implicitly, at which point you don't need an async dealloc, you can just have a sync one which shunts the actual dealloc and mandate a background thread. Which means now you're mandating a background thread for freeing memory. And you hope that the load is light enough that implicit thread which you've tasked with all deallocations (rather than just the ones which make sense) keeps up. > Or just use an arena that can be freed all at once. That is essentially the same thing, you now have a different allocation strategy for these, except it's a lot more limited and specific.
- SleepyMyroslav 4y ago> No? Allocations are blocking, Modern malloc/free implementations must satisfy requirements such as being scalable. It eliminates any chance to implement blocking deallocation from non-owner threads. A modern implementation already has to be async in some ways. You can check paper behind Mimalloc for example.
- kaba0 4y agoAn arena allocator presupposes a huge number of objects each having the same lifecycle - which is not all that common/general.
- generichuman 4y ago> which is not all that common/general. Your objects in trees and arrays don't have the same lifecycle? Quoting GP: > ... reference counting (RC) have stampeding characteristics similar to GC pauses when releasing large hierarchies e.g. persistent data structures, or large trees of widgets: all the tree is freed synchronously and recursively, if it's large it can be very sensible. This is mostly where RC & MMM fails. This is also where you are supposed to use arena allocators. If you have a tree you're constantly adding to / deleting from, you can use a Pool backed by an Arena. At that point you don't use pointers, only indices into the pool. When you want to remove something from the tree, you mark the objects in the pool "deleted" until you want something added back again.
- WalterBright 4y ago> I actually find it one of the biggest downfalls of D. Why? The gc is optional in D. You can use malloc/free, or your own custom allocator.
- skocznymroczny 4y agoI agree about riding the fence. The main reason D is so criticised for having a GC because it gives users a choice. I'd also rather have it either go full GC or full manual/semi-manual memory management, probably leaning more towards GC side personally. There are other aspects where D gets hit from both sides because it never picks a side, instead it tries to cater to both sides, never going full into any of them. I guess there is some benefit in a language that's general and doesn't force you into specific paradigms, but it also increases the surface area (maintenance area) of the language and adds complexity for developers when every step of the way you have to consider the alternatives.
- WalterBright 4y agoD obviates the need to use different languages for different parts of an application, by supporting multiple paradigms.
- ErikCorry 4y agoAgreed that riding the fence doesn't give a satisfactory result. In theory, GC is optional in Go, but I don't think anyone seriously running without a GC, or writing code that works well without a GC.
- DeathArrow 4y agoI love the garbage collector in Nim. Not only you can chose between different memory management strategies but you can tune the garbage collector to your liking and it's pretty fast by default.
- winrid 4y agoYes, it's amazing. Now we just need a Clion plugin that uses nimsuggest so I can actually use an IDE... :)
- DeathArrow 4y agoUntil then there's Visual Studio Code.
- winrid 4y agoI'm good
- deleted 4y ago[deleted]
- Tozen 4y agoNot at all professing any love of Nim, but do think their concept of various options for memory management was a good one. Not all programmers have the same goals and problems. Part of the issue can be that languages that do not provide any options for memory management, can try to make it seem that GC is more of a liability than it is. In the case of Nim, D, and other langauges... They are giving options, versus none. The lack of convenient choices, might be the greater issue, versus stigmatizing GC.
- grumpyprole 4y ago> Not all programmers have the same goals and problems. That's why we have different programming languages. The danger of making a language a jack-of-all-trades is that it will be master of none. Restrictions are more often than not a good thing, they give the power to reason, for both humans and machines (giving, for example, memory safety).
- alrlroipsp 4y ago
- dang 4y agoPlease don't cross into personal attack.
- alrlroipsp 4y ago
- jhhh 4y agoYou hate when the people most knowledgeable about something take the time to explain things?
- ziotom78 4y agoHow interesting! I know virtually nothing of GCs, but posts like this suggest that there is a huge amount of research behind something that looks do deceptively simple as calling a "free/delete" here and there. Kudos to whoever works on this stuff!
- nanolith 4y agoGC is a quite rich area of research. I've implemented novel GC algorithms in firmware. There are some algorithms that are just more compact when GC is available. Algorithms and heuristics that are graph heavy or require pruning nodes with possible cyclic references are just more elegant and code space efficient with GC. However, firmware is highly resource and timing constrained, which means that the GC is specialized.
- masklinn 4y ago> And since there are all kinds of pointers in D, one no longer can use a moving GC allocator, because it cannot know exactly where 100% of the GC pointers are. Go's GC is non-moving and it can't safely be made moving in the current state[0]. It also used to be partially conservative until 1.3[1][2][3]: the heap was precise, but the stack was conservative. [0] https://github.com/golang/go/issues/46787 https://github.com/golang/go/issues/46787 [1] https://go.dev/doc/go1.3#garbage_collector https://go.dev/doc/go1.3#garbage_collector [2] https://docs.google.com/document/d/1lyPIbmsYbXnpNj57a261hgOYVpNRcgydurVQIyZOz_o/pub https://docs.google.com/document/d/1lyPIbmsYbXnpNj57a261hgOY... [3] https://docs.google.com/document/d/13v_u3UrN2pgUtPnH4y-qfmlXwEEryikFu0SQiwk35SA/pub https://docs.google.com/document/d/13v_u3UrN2pgUtPnH4y-qfmlX...
- hayley-patton 4y agoIt's also worth mentioning that being conservative only on the stack, even when moving, is often good enough. SBCL still works that way, and the space and time overheads are pretty tiny [0]. [0] Fast conservative garbage collection https://users.cecs.anu.edu.au/~steveb/pubs/papers/consrc-oopsla-2014.pdf https://users.cecs.anu.edu.au/~steveb/pubs/papers/consrc-oop...
- titzer 4y agoAFAICT pinning is temporary? E.g. even Java's native interface allowed pinning objects for a limited amount of time, which had to be incorporated into the design of Java's long series of GCs.
- ErikCorry 4y agoA historic mistake. These days a Java GC doesn't have to support pinning, and I don't think they do. Ctrl-F pinning at https://docs.oracle.com/javase/7/docs/technotes/guides/jni/spec/design.html https://docs.oracle.com/javase/7/docs/technotes/guides/jni/s...
- titzer 4y agoErg, but soon the Virgil GC will have to support pinning, because you can directly call the kernel to do I/O with on-heap byte arrays. Currently I get away with that because Virgil is single-threaded only, but that's gotta change for Wizard to compete with more advanced VMs.
- moonchild 4y agoI've explained[0][1] in the past why this is nonsense, and seem to have eventually convinced[2] the one person who actually read the literature. Beyond that--crickets. 0. https://forum.dlang.org/post/mpczwoeykpcwjakuderd@forum.dlang.org https://forum.dlang.org/post/mpczwoeykpcwjakuderd@forum.dlan... 1. https://forum.dlang.org/post/eujlhvszhpbyoxemscxb@forum.dlang.org https://forum.dlang.org/post/eujlhvszhpbyoxemscxb@forum.dlan... 2. https://forum.dlang.org/post/ssq34d$2ur7$1@digitalmars.com https://forum.dlang.org/post/ssq34d$2ur7$1@digitalmars.com
- mhh__ 4y ago
- WalterBright 4y agoI've worked with languages with multiple fundamental pointer types (needed for 16 bit code). This kind of thing never worked well. For example, you'd need 4 versions of strcpy() to deal with 2 pointer types. What a mess DOS programming was with that.
- moonchild 4y agoIt's a good thing that's not the solution I proposed, then.
- WalterBright 4y ago> So instead, generate multiple definitions and rewrite callers:
- moonchild 4y agoYes, in much the same fashion as with template functions: multiple definitions are generated, and callers are rewritten to target the appropriate specialisation.
- 4y ago
- Tozen 4y agoI think many people look at GC in the wrong way, as if absolute maximum speed is always the goal. It depends on what the goal is, and GC performance can be as fast as needed or enough. GC is also a convenience, that many may find worthwhile, to not worry about memory management. Optional GC, also means you are free to go for it manually, if there is really and truly some bit of extra performance needed. In regards to language wars, it's arguably better to make the case of how to make manual memory management more convenient to use, when the optional GC is turned off.
- WhitneyLand 4y ago>GC is also a convenience “Convenience” makes it sound like syntax sugar, it’s not merely convenient. It eliminates certain types of bugs fundamentally, often it can simplify how the code is reasoned about.
- mpweiher 4y ago"Convenience" is exactly right and syntactic sugar is underrated. Just like static typing is a convenience. And dynamic typing is a (different) convenience. And unit tests are a convenience. The computer runs fine without any of these. And none of them are a magic wand that eliminates all bugs. Instead all of them make debugging easier by reducing the search space, usually eliminating fairly trivial bugs. Which is more useful than it sounds because a lot (the majority?) of bugs are pretty trivial and at the same time hard to find because they are often so trivial that they stare us in the face and we can't see them. But again, all these things are conveniences for the human programmer, they don't make one iota of a difference for the actual computer running the bits.
- pjmlp 4y agoI always discuss about SLAs when arguing about language X vs Y. Even if language X loses in microbencharks against language Y, if the application written in X is still within the project delivery SLA for acceptance testing, who cares about the microbencharks.
- 4y ago
- hayley-patton 4y agoThe Go collector isn't generational or moving, and the write barrier AIUI is only used for getting it to run concurrently. The barrier records interesting writes when the collector is running, to avoid "losing" objects whose liveness changed while the collector was running. Barriers are pretty cheap; [0] claims a 0.9% time overhead for a card-marking barrier common in generational collection and 1.6% for an object-logging barrier which is also useful for concurrent collection. Apparently they're cheaper nowadays, but the results aren't published yet [1]. That's not to say that the barriers are free, but it seems feasible that collector optimisations could still cover that ground. It can help throughput, still, by running the collector concurrently. I've felt that while working on a new parallel (but not yet concurrent) collector for SBCL; the program parallelises well, but a serial collector hurts worse than it should by stopping the program. [0] "Barriers reconsidered, friendlier still!" https://users.cecs.anu.edu.au/~steveb/pubs/papers/barrier-ismm-2012.pdf https://users.cecs.anu.edu.au/~steveb/pubs/papers/barrier-is... [1] https://twitter.com/stevemblackburn/status/1494240906006110209 https://twitter.com/stevemblackburn/status/14942409060061102...
- WalterBright 4y agoIn a pointer heavy language like D, a 1% loss is a big deal. I.e. it would no longer be a systems programming language.
- hayley-patton 4y agoI would expect to be able to find another 1% slowdown somewhere in an implementation of another "system programming language" - would that also disqualify it from being such a language?
- WalterBright 4y agoIt's a competitive business. If a compiler generates 1% faster code, you need 1% fewer servers in your server farm, which translates to enormous amounts of money. Or if your hedgefund trading software runs 1% faster, you can get your trades in faster than the other guy, eating his lunch. Or would you like to reduce your costs of rented servers in the cloud by 1%?
- shaburn 4y ago
- self_awareness 4y agoSince it's optional, why does it matter that much?
- forgotpwd16 4y agoBecause it's "optional". Some of D's features (https://dlang.org/spec/garbage.html#op_involving_gc https://dlang.org/spec/garbage.html#op_involving_gc) depend on it. It's optional in terms of you cannot use those. But if it's to never use them, why they even exist?
- WalterBright 4y agoThe GC is very nice for evaluating functions at compile time. It has zero runtime cost.
- tristanbvk 4y agoHow would one go about instantiating an object in D without GC if the new keyword relies on it. Accounting for memory allocation and object initialization (construction etc)? Legit question coming from a D enjoyer
- WalterBright 4y agoThe compile time function execution enables allocation with `new` and it handles initialization, etc.
- tristanbvk 4y agoThanks Walter!
- p0nce 4y agoIt doesn't. It just got posted here.
- noobermin 4y agoSomewhat on topic, but it feels like D evangelists and Walter Bright keep extoling everywhere about D but D has yet to really be used as widely as other modern languages. D forever was the "C++ replacement" and I feel like it still hasn't shaken that perception in people's minds even though D has come onto its own as a unique language.
- pjmlp 4y agoSince Andrei Alexandrescu decided to refocus on C++, I think it kind of settled the future of the language, when core members go back to the language that apparently is less capable. https://research.nvidia.com/person/andrei-alexandrescu https://research.nvidia.com/person/andrei-alexandrescu
- anta40 4y agoI guess D doesn't fits my use cases. I mostly write native mobiles apps (hello Kotlin), but ocasionally deal with backend codes in... Go. If I need "C++ replacement" for system programming, then the choice would be umm... Rust?
- habibur 4y agoI won't be worried about the speed of a Garbage collector. Rather focus on its efficiency. Actual work done per cycle per byte. Collectors can run on a separate core and show double the speed, but then you have to halve the total number of working processes on your machine. Thus total work done per cycle remains the same. Or it can just be lazy and run collection half the time, and thus show half processor use. But RAM use will then double for the same work load. At the end actual work done per byte of memory remains the same. So focus on efficiency. The user can utilize the spare cores and RAMs for paralyzing their work, or for more work.
- tsimionescu 4y agomalloc() and free() also do work. Depending on the workload, a GC can do more or less work than manual memory management.
- habibur 4y agoThat's beside the point.
- rtfeldman 4y ago> And since there are all kinds of pointers in D, one no longer can use a moving GC allocator, because it cannot know exactly where 100% of the GC pointers are. I was astonished to learn that researchers found a way to implement a compacting malloc (!!!) by using very clever virtual memory tricks - and which they were able to use to demonstrate memory usage improvements in a long-running Redis instance that used their drop-in malloc replacement: https://www.youtube.com/watch?v=xb0mVfnvkp0 https://www.youtube.com/watch?v=xb0mVfnvkp0
- zozbot234 4y agoThat's just trading a fragmented heap for fragmented page tables, though. Modern OS's even support "huge" pages to specifically avoid that page table overhead.
- chrisseaton 4y agoWhy do you say ‘just’? It’s an effective technique.
- _rlh 4y agoStill one of the best ideas in the field in recent years. I will note that it also works for non-moving GC collectors and if they are precise, like Go, they can also update pointers and eliminate the redundant page table entries.
- hiq 4y ago> 1. It enables a killer feature - CTFE that can allocate memory. C++ doesn't do that. Zig doesn't do that. As anything but a compiler expert, I don't understand how GC is a necessary condition of having memory allocation within CTFE[0], can somebody expand on that? [0]: CTFE seems to be an acronym for https://en.wikipedia.org/wiki/Compile-time_function_execution https://en.wikipedia.org/wiki/Compile-time_function_executio... (I was not familiar with it).
- WalterBright 4y agoThe CTFE interpreter knows what "new" does, and just calls the compiler's "new" function to allocate memory. If the user wrote their own custom allocator, then the CTFE would have to interpret that, which is kind of a big mess.
- eatonphil 4y agoDumb question: what's the difference between CTFE (I haven't heard of that term before) and Zig's comptime and C++'s constant expressions?
- WalterBright 4y ago1. Zig and C++ cannot allocate memory for CTFE. 2. Zig and C++ require any code to be run at compile time to be specially marked (comptime, constexpr). D will run any code that appears in a const-expression at compile time. 3. Zig and C++ require that if a function is to be used for CTFE, the entire function must be compatible with CTFE. D only requires the path taken through the function to be compatible with CTFE. D even has a `__ctfe` pseudo-variable that can be used to branch within a function to compile-time and run-time paths. https://dlang.org/spec/function.html#interpretation https://dlang.org/spec/function.html#interpretation
- leni536 4y agoC++20 has limited support for allocating memory at compile time. The allocations can't be leaked from the constant expression context, so that limits usefulness. There are proposals to extend it, but there are some non-trivial const correctness issues to be solved there, AFAIK.
- chrisseaton 4y ago> although it does do escape analysis to figure out what can be allocated on the stack instead. (Java does this as well.) Walter has a misconception here - Java does scalar replacement - it doesn’t currently do stack allocation.
- ErikCorry 4y agoInteresting distinction, but once the object is replaced by scalars, those scalars are placed on the stack. Do I understand it correctly that you are saying real stack allocation would involve allocating the whole object, including the header, on the stack, and passing references to such objects to (non-inlined) functions that worked on either stack allocated or heap allocated objects?
- chrisseaton 4y ago> Interesting distinction, but once the object is replaced by scalars, those scalars are placed on the stack. No, they become data flow edges, so could be in a register, or part of an addressing operation, or value-numbered, or nothing at all if they’re never used, so only on the stack as a worst-case fallback. > Do I understand it correctly that you are saying real stack allocation would involve allocating the whole object, including the header, on the stack Yes, which is useful in some cases, but generally a lot weaker than full scalar replacement.
- vips7L 4y agoAn interesting read on Java escape analysis/scalar replacement if anyone is interested: https://gist.github.com/JohnTortugo/c2607821202634a6509ec3c321ebf370 https://gist.github.com/JohnTortugo/c2607821202634a6509ec3c3...
- Kukumber 4y agoI don't think it is slower, the real answer is 'it depends'; D's GC will only run when the GC needs to grow its buffer, so D's GC can be actually much faster than Go's The problem however is, actually 2 problems: - stop the world, nobody wants that in a world with lot of cores and threads - it doesn't scale, the more pointers in your heap, it'll need to scan and traverse your WHOLE heap whenever it needs its buffer to grow, and that doesn't scale well So it's good when you don't have much in your heap, and it starts to loose its benefits the bigger your program become, i wouldn't use it for my servers But that's not the main problem of D, since the GC is optional, it's just not competitive with what's available in the market today The people who want to drag D into the Java/C# territory are the problem in my opinion D would be better if it focused being a system language, and took what C had to offer and put it to the next level, simplify the language, boost the existing features, allocators, pattern matching, tooling, compiler performance, hot-reload, binary patching That's the thing i want to hear when there is a new version, not the endless GC topics
- arunc 4y ago> The people who want to drag D into the Java/C# territory are the problem in my opinion Absolutely and that's the majority of D community. > took what C had to offer and put it to the next level ImportC is a fantastic stuff that Walter is working on. > simplify the language Sane defaults, but the ship has sailed. Reminds me of a talk Scott Meyers gave a while ago "The last thing D needs (is to hire him)". I think it's time they hire him. > compiler performance Rather focus on just LDC or GDC and drop DMD altogether. DMD is a good piece of software, but for such a small community, I find it alarming that they waste human effort across 3 different compilers and still complain lack of resources.
- Kukumber 4y ago> ImportC is a fantastic stuff that Walter is working on. I agree, it's one of the things that stands out when you decide to pick a system language: "how does it play with C? can i easily consume the ecosystem?" > Rather focus on just LDC or GDC and drop DMD altogether. DMD is a good piece of software, but for such a small community, I find it alarming that they waste human effort across 3 different compilers and still complain lack of resources. I disagree, there is value in having your own backend, DMD compiles so fast, it's a comparative advantage, they should never give that up GDC/LDC are great because that allows D to be highly portable, even if they are slower to compile than DMD Even Zig people decided to maintain their own self hosted backend for that reason, performance and independence They learnt from D, a real language has its own backend, if you don't then you are just LLVM sugar
- strictfp 4y agoI would love it if some language would implement a segmented heap, where each part could be GCed separately. Erlang has this model with it's lightweight processes. And it's a great model that helps not only with GC, but also guaranteeing no shared state between different parts of the code.
- zozbot234 4y ago> where each part could be GCed separately That would mean disallowing cyclical references across different GC heaps; non-cyclical references would just create additional GC roots. It would be a step away from totally automated memory management and towards something closer to "smart" reference counting.
- strictfp 4y agoI'm saying that references between heaps should be disallowed altogether.
- pjmlp 4y agoSee Pony.
- throaway53dh 4y ago
- qualudeheart 4y agoD has a long history, but D has always been a language in transition. It is going through a period of refinement where features are added and the underlying abstractions are made more efficient. So, I expect the performance problems to be solved reasonably soon. For now, D developers should look at the current status of D's garbage collection and how to improve it. When D was originally developed it was intended to be a language that was in between C and Java. D was to be a very low level language. So, it was not as low level as C and not as high level as Java. It still managed to outperform Java and rival C. Ten years from now it'll be even faster.
- yencabulator 4y ago> The GO GC can also take advantage of always knowing exactly where all the GC pointers are, because there are only GC pointers. You can call the mmap syscall (or a wrapper like C.malloc) from Go just fine, and obtain non-GC pointers. Not all Go pointers are pointers into the GC-managed region. This makes me question the rest of the opinions herein.
- yvdriess 4y agoCan that space contain Go objects?
- yencabulator 4y agoNon-pointers, yes. (EDIT: Pointers to non-Go memory are fine, though generally impossible to use in e.g. mmapped file data. Pointers to Go memory stored outside of the GC region will not be seen by the GC, nor guaranteed valid past a single CGo function call.) A common trick is to e.g. treat an mapped file as a large array of structs.
- Imperatorn 4y agohttps://dlang.org/blog/2017/06/16/life-in-the-fast-lane/ https://dlang.org/blog/2017/06/16/life-in-the-fast-lane/
- Imperatorn 4y agohttps://dlang.org/blog/2017/03/20/dont-fear-the-reaper/ https://dlang.org/blog/2017/03/20/dont-fear-the-reaper/ Basically, you should only use GC for small litter, i.e. some short-lived tiny objects, and for larger and longer living objects use deterministic memory management tools available in the language and libraries.
- heynowheynow 4y agoIf one wanted a fast GC, they could begin with code generation or implementing a VM to work with something like ORCA. It has more powerful sharing and exclusivity semantics than Rust. Pony being the example language implementation. In benchmarks, it stomps Azul C4 and BEAM/HiPE.