Most people's mental model of garbage collection is still just stop-the-world mark-and-sweep, which is extremely old. Game developers hate GC because of impossible to predict or the lack of ability to limit latency pauses eating into the frame budget - but there are tons of new, modern GC implementations that have no pauses, do collection and scanning incrementally on a seperate thread, have extremely low overhead for write barriers and only have to scan a minimal subset of changed objects to recompute liveness, etc. that are probably a great choice for a new programming language for a game engine. Game data models are famously highly interconnected and have non-trivial ownership, and games already use things like deferred freeing of objects (via area allocators) to avoid doing memory management on the hot path that it could do automatically for everything.
Part of it is also due to Unity3D sticked with the ancient stop-the-world mono GC for a very long time and discuraging developers from doing allocation in official talks.
Flipping this, in the C++ world, "garbage collection", is more than just freeing unused references. "Resource release is destruction"
Try-with-resources is ok, but not great.
Ultimately developers, especially those concerned with performance, prefer full control. And that means deterministic lifetime of memory and resources.
Another big cost of garbage collection is memory usage: it's not uncommon for a high-performance GC to require 2x-3x the memory for the same application compared to non-GC. For games this is not trivial given how heavy some of the assets are.
After being both a Java programmer for the greater part of a decade, then spending a bunch of years doing Objective C/Swift programming, I don't really understand why the Automatic Reference Counting approach hasn't won out over all others. In my opinion it's basically the best of all possible worlds:
1. Completely deterministic, so no worrying about pauses or endlessly tweaking GC settings.
2. Manual ref counting is a pain, but again ARC basically makes all that tedious and error prone bookkeeping go away. As a developer I felt like I had to think about memory management in Obj C about as much as I did for garbage collected languages, namely very little.
3. True, you need to worry about stuff like ref cycles, but in my experience I had to worry about potential memory leaks in Obj C about the same as I did in Java. I.e. I've seen Java programs crash with OOMs because weak references weren't used when they were needed, and I've seen similar in Obj C.
Compared to other, state of the art GC strategies, it's slow and has bad cache behavior esp. for multithreading. Also, not deterministic perf wise.
Can you explain these or point me to any links, I'd like to learn more.
> has bad cache behavior for multithreading
Why is this? Doesn't seem like that would be something inherent to ref counting.
> Also, not deterministic
I always thought that with ref counting, when you decrement a ref count that goes to 0 that it will then essentially call free on the object. Is this not the case?
Writing to the rc field even for read only access dirties cache lines and causes cache line ping ping pong in case of multi core access, where you also need to use slower, synchronised refcount updates so not to corrupt the count. Other GC strategies don't require dirtying cache lines when accessing the objects.
Determinism: the time taken is not deterministic because (1) malloc/free, which ARC uses but other GCs usually not, are no deterministic - both can do arbitrary amounts of work like coalescing or defragmenting allocation arenas, or performing system calls that reconfigure process virtual memory - and (2) cascading deallocations as rc 0 objects trigger rc decrements and deallocation of other objects.
> Also, not deterministic
> not deterministic perf wise.
was what parent wrote (emphasis added), I assume referring to the problem that when an object is destroyed, an arbitrarily large number of other objects -- a subset of the first object's members, recursively -- may need to be destroyed as a consequence.
Bad cache behavior: you're on core B, and the object is used by core A and in A's L2 cache. Just by getting a pointer to the object, you have to mutate it. Mutation invalidates A's cache entry for it and forces it to load into B's cache.
determinism: you reset a pointer variable. Once in a while, you're the last referent and now have to free the object. That takes more instructions and cache invalidation.
> there are tons of new, modern GC implementations that have no pauses
There are no GC implementations that have no pauses. You couldn't make one without having prohibitively punitive barriers on at least one of reads and writes.
There are no CPUs that have no pauses; who knows, you may have to wait a microsecond for that cache line to come in from RAM.
There are no operating systems that have no pauses; if you don't want to share the CPU, you're going to have to take responsibility for the whole thing yourself. Most people are not even using RTOS.
There are no reference counting implementations that have no pauses. Good luck crawling that object graph. Most people are not even using deferred forms of reference counting.
As far as I know, there is one malloc implementation which runs in constant time. No one uses it.
There are tracing GCs whose pause times are bounded to 1ms. That is enough for soft-real-time video and audio (which is what matters to video games). In general, you are not going to get a completely predictable environment unless you pick up an in-order CPU with no cache and write all your code in assembly.
I think the more accurate / useful statement is, there's no GC with bounded pause times that's also guaranteed to free memory fast enough you don't run out even though you shouldn't. In other words, GC can never fully replace thinking about your allocation patterns.
(I suspect this is obvious to a lot of programmers, but especially in working with Java programmers who just skim JVM release notes and then repeat "pauseless", it's also not obvious to many too.)
Right, but if you're allocating way too much, no lifetime management model will save you.
As a game dev I still will take a modern GC over refcounting or manual lifetime management any day. There are scenarios where you simply don't want to put up with the GC's interference, but those are rare - and modern stacks like .NET let you do manual memory management in cases where you are confident it's what you want to do. For the vast majority of the code I write, GC is the right approach - the best-case performance is good and the worst-case scenario is bad performance instead of crashes. The tooling is out there for me to do allocation and lifetime profiling and that's usually sufficient to find all the hotspots causing GC pressure and optimize them out without having to write any unsafe code.
The cost of having a GC walk a huge heap is a pain, though. I have had to go out of my way to ensure that large buffers don't have GC references in them so that the GC isn't stuck scanning hundreds of megabytes worth of static data. Similar to what I said above though, the same problem would occur if those datastructures had a smart pointer in them :/
GC has a lot of issues, especially in engine level code, but in practice every game I've worked on or shipped has had at least one garbage collector running for UI or gameplay code. Lua, Actionscript (Scaleform), Unreal Script, Javascript, managed C#. Every game had GC performance issues as well, and we wrote code to generate minimal or no garbage
Some game engines (C++) do allocations in multiple memory areas. If some allocated memory is known to be needed only for the current frame, then it is allocated from that memory area. Then at the end of the frame the whole region is freed. This is an explicit garbage collection at end of each frame. Memory allocation can further be split according to the CPU thread, thus avoiding global locking in the allocator. With double-buffering the next updated/prepared frame can use its own memory area, while the previous one finishes rendering.
The problem with GC or with reference counting is that it needs to operate on each allocated object separately. If the task for the GC can be reduced to operate on whole memory areas only, its overhead is greatly reduced.