30 ms·
For Better Computing, Liberate CPUs from Garbage Collection
- notacoward 7y agoReadable copy of the paper at Berkeley: https://people.eecs.berkeley.edu/~krste/papers/maas-isca18-hwgc.pdf https://people.eecs.berkeley.edu/~krste/papers/maas-isca18-h... ETA: this doesn't seem to be quite the paper that the story refers to, but undoubtedly describes the same work in enough detail for people to get the gist of it. Darn paywalls.
- kazkada 7y agoThe paper is available on Scihub: https://sci-hub.se/10.1109/MM.2019.2910509 https://sci-hub.se/10.1109/MM.2019.2910509
- azinman2 7y agoI’m probably wrong, but didn’t the Symbolics LISP machines have some kind of hardware support for GC? I think for platforms like Android this makes a lot of sense. Should help quite a bit with battery consumption and responsiveness. Also makes sense for server loads in Java or Go.
- ghjjjj 7y agoIt did.
- gwern 7y agoA https://en.wikipedia.org/wiki/Tagged_architecture https://en.wikipedia.org/wiki/Tagged_architecture yes. Tags make life a bit easier for the CPU when it accesses objects and is deciding what to do with them, but the CPU is still doing all the work and walk the RAM to do the GC. (I vaguely recall stories about Lisp machine users who would turn off GC while working, and then let it run when they left work and returned the next day.) The idea here seems to be to have an entire separate chip, a specialized CPU, which walks RAM independently of the 'main' CPUs, whose only task is freeing up memory.
- lispm 7y ago> I vaguely recall stories about Lisp machine users who would turn off GC while working That was more in the early days when the GC was more primitive. The later GC was actually several different ones. A huge impact had the so-called 'Ephemeral GC', where the machine tracks changed RAM memory (it uses virtual memory in general) and where the EGC focuses on those areas with objects which were only very short-lived. That means, given a sufficient amount of memory, the machine stayed fast and responsive for a long time. The main problem was the low amount of RAM (20MB were common after the mid 80s), because RAM was very expensive, and the large amount of virtual memory (possibly hundreds of MB on slow disks). Thus what made the machines really slow was the GC over virtual memory with relatively slow disks. Thus the later mainly used GC was incremental&generational©ing&with areas and with support from the EGC - a full Mark&Sweep GC would keep the machine possibly busy for 15-40 minutes.
- _old_dude_ 7y agoThere is a renaissance of that idea in ZGC [1]. You have a lot useless bits in a 64 bits pointers and thanks to the virtual memory, you can manage to have an untagged address and its corresponding tagged address referencing the same physical memory. This give you free bits that you can use to track the liveness/evacuation of a graph of objects. [1] https://wiki.openjdk.java.net/display/zgc/Main https://wiki.openjdk.java.net/display/zgc/Main
- p_l 7y agoLisp Machines didn't use tags for tracking liveness/evacuation of objects, though. They used them for safety, which automatically gave them precise, as opposed to conservative, GC which always knew whether it was dealing with a pointer. They also had special, CPU-handled type of forwarding pointers, which when accessed "normally" would transparently redirect you to forwarded location.
- _old_dude_ 7y agoThe forwarding pointer you are describing is equivalent to a ZGC colored pointer with the evacuation bit set that a GC barrier (a load barrier) will rewrite to the evacuation address. and yes ZGC doesn't use colored pointer to track if a value is an integer or a pointer because Java unlike Lisp is typed so the VM derives those information from the bytecode.
- p_l 7y agoSymbolics (and other systems like it) had typed memory with hw type checking (so you had safe memory by default), and had one important bit for efficient GC (one that is also utilized by Azul and Shenandoah) - forwarding pointers, in their case implemented "transparently" to actual code. With forwarding pointer, you can copy data and update a reference in a way that is transparent to concurrently running code, as when it access the data again the CPU (or software implementation) will notice the data had been moved and follow the forwarding pointer to new location. The rest of the support for GC was centered around structure of virtual memory, without anything very specific to GC in hw - we're talking about things like "memory is divided into X areas, inside each area the lower side is ephemeral... " etc. There was also some register use to keep track of GC status. I think some versions had MMU or page fault handler update the GC register data.
- etaioinshrdlu 7y agoIt does seem like just doing it in hardware may be a linear gain but isn't a fundamentally better algorithm. There's a proof that you do need to pause your program eventually, if you want to be sure you get all the garbage.
- gizmo686 7y agoHardware is fundamentally parallel, CPUs are fundamentally serial; it is possible for a hardware solution to have a super-linear speedup in time. As a simple example, what is the time complexity of zeroing out n bytes of memory. With a CPU, this is O(n). However, with proper hardware support, this can be done in O(1) time. For a simple garbage collecting example (no idea how their chip does it), consider a simple mark and sweep algorithm. Assume the chip has an internal object graph. At the begining of GC, only root nodes are tagged. At each step, the neighboor of every tagged node is tagged. With a CPU, this step takes at least Ω(n) time, where n is the number of tagged nodes. However, if this is done entirely by hardware, then (depending on how the hardware is designed) each node can independently look at its neighboors and complete a single step in just O(1) time. Moving stuff to hardware gives you a set of primitives that is asymptotically different than the primitives you have on a general purpose CPU.
- AnimalMuppet 7y agoBut hardware that has to interface with the main memory can't be fundamentally parallel, because of memory bandwidth limitations. If you want to make this part of the memory, then my objection does not apply. But if it's an external chip to the memory, you still are fundamentally serial.
- gizmo686 7y agoHow much does it need to interface with main memory? If the external chip has its own memory then it can maintain its own view of the object graph. The software can update it when references are made/deleted and query it when an allocation needs to be made. There is still the overhead of communicating when references are made/deleted, but you need that anyway, as something that only looks at memory doesn't know what is a pointer, or the size of objects. The communication overhead is then linear in with respect to the amount of updates to the object graph, not the size of the object graph. You could even go so far as to not put the chip on the memory bus at all.
- stcredzero 7y agoWhile we're at it, how about we liberate CPUs and caches from communication between threads and cores?
- garmaine 7y agoMany ARM chips have this.
- stcredzero 7y agoAny useful links? What are the terms I should search for?
- garmaine 7y agoIt’s called a “weak ordering memory model.” Synchronization requires explicit memory fence instructions. RISC-V, btw, supports both weak and strong memory models as an implementation choice.
- CoolGuySteve 7y agoFrom an environmental perspective, I wonder how much energy is consumed (and emissions generated) for garbage collection and interpreters. These things exist to make programming easier but are then duplicated across thousands of servers. If everyone used some compiled language that was just a little simpler, a little safer, had just a little better memory management/tooling, or like here, had better hardware support, how much would that reduce global emissions caused by data centers?
- Wowfunhappy 7y ago> These things exist to make programming easier but are then duplicated across thousands of servers. I think about this kind of thing, but more in regard to Electron et al. Garbage collection, by contrast, seems like a more worthwhile trade-off, because there's a great argument to be made that garbage collection isn't just easier for programmers but safer for users, in terms of security. Since it mostly removes an major class of vulnerabilities.
- exabrial 7y agoSo... Rust, basically
- CoolGuySteve 7y agoI didn't want to start a language war but yeah, basically. Swift and/or ObjC + ARC seem to be attacking the same problems on the embedded side out of necessity. It's too bad Apple is so apathetic about Linux servers. Even just replacing Python with Go might save a massive amount of KWh consumed.
- Scramblejams 7y agoI'm more interested in optimizing the client side, at least where the web is concerned. I keep wondering when Google is going to start generating energy ratings for the top few thousand sites on the web. It would shame the low performers into producing front ends that are no longer glacial. I was loading wunderground.com's 10 day forecast for the umpteenth time lately and reflecting on how horribly slow and inefficient it is... Surely website energy ratings must be in the works somewhere, right?
- exabrial 7y agoAzul Systems has asked Intel to do this once... but instead created their own processors with interesting memory barrier properties for awhile that greatly sped up JVMs beyond what was capable (at the time) on x86-32/ppc/sparc. Eventually they gave up and became a purely software company, but their "Java Mainframe" product was many times faster than the Intels of the age executing the same code despite much slower CPUs. Died a quick life despite the cool factor.
- aero142 7y agoI feel like Amazon is going to bring back custom hardware like this. Imagine if this was an instance type.
- ddingus 7y agoI think everyone is about to make custom hardware. A recent custom chip project I have been a part of for basically a decade is nearing production. Made on a fairly large, old process. Despite this, many custom features have been added. For many tasks, the performance will be competetive with much more complex, resource intensive devices. It is built in a way that allows for efficient, multi core computing, concurrent or parallel, sans an OS. People will write those, but many will just grab the pieces they want, put them on cores, then write their target app on top. The prior version, Propeller 1, made doing that a lot easier than one might think. What struck me was the combination of very well planned silicon, coupled with software, can really perform. It reminds me of the custom chips we saw in early computing. Amiga, SGI, many others made adapter cards to get things done. All of that stuff nailed the tasks cold, would always perform. As CPUs got quicker, more could be move to software of course. But now that is hard again Software plus purpose built silicon is going to deliver peak performance. Always has. In that project, it took years, but a great many use cases were considered, FPGA simulations ran, code written, and then augmented with hardware, special instructions and sub systems intended to maximize performance while retaining a lot of flexibility. The way I see it, general purpose computers may just end up back where they were before. In 8 bit times, an Apple 2 was an all software machine. People bought and made cards to do specific things very well. Other computers had custom chips that focused on games, etc... In the later era, Amiga, SGI made great hardware that was focused on specific things, while the PC was more like the Apple 2. We are currently leaving a long era where general purpose computing made sense for most cases, and software got refined, things got faster, and we saw many good cycles. Soon, more efficient CPU designs, often with well tuned instructions for given tasks may be directing a lot of purpose built silicon. Phones and tablets vs laptops give us a tiny look at one part of how it might go. GPU instances, specialized network maybe filesystem CPUs with highly optimized instruction sets already exist. More will come. Many use cases can benefit from real time or just faster performance per watt. Software plus custom silicon will nail that, and it is getting easier to do. RISC V has plans for specialized instructions baked in. It may well be that "one to bind them all", nudging out the more expensive ARM, for example. And maybe not. ARM is lean and mean and mature. Who knows? All I know is the drive to do more per watt as well as the drive to improve peak sequential compute, because there are still too many use cases where doing that makes sense, given either the nature of the problem space, or accumulates software warrant the effort. We are already seeing the GPU idea expanded on. Gonna be interesting and more difficult times ahead.
- mailslot 7y agoSeems like we could just as easily stop using garbage collection. ... or even go back to reference counting / smart pointers and just live with the “limitation” that we can’t have circular references.
- hnaccy 7y agoWhy limitation in quotes? Do you feel it is not one in practice?
- mailslot 7y agoI argue that some limitations, like clear resource ownership, are good... but there are still workarounds that aren’t too arduous. Weak pointers and the like are decent enough abstractions to handle the rare case... or just going manual. It’s not a limitation in practice, no. It hasn’t been a problem for me in over a decade in a multitude of languages. On Java codebases? I’ve witnessed some frightening levels of sloppiness that are only solvable with a profiler.
- Eric_WVGG 7y agodoesn’t seem cool to admire Apple technologies, but ARC seems to work amazingly well with zero CPU overhead
- seandougall 7y agoIt’s not zero — those -retain and -release calls still exist when necessary and have a small penalty — but it does minimize them well, and is lower overhead and vastly more predictable than GC.
- favorited 7y agoIt's a little better than that, because the overhead of ObjC method dispatch is avoided. There are no calls to -retain or -release (if the receiver hasn't overridden those methods), it's just a C function call to objc_release et al. No objc_msgSend involved.
- jhallenworld 7y agoSo my idea for GC is to offload it to a separate machine through a communications channel. The main CPU sends messages to the co-processor whenever it allocates memory, or whenever it mutates (whenever it writes a pointer to allocated memory or to the root set- there could be special versions of the move instructions which send these messages as a side-effect). There is a hardware queue for these messages and the main processor stalls if it's full (if it's getting ahead of the co-processor). The co-processor then maintains a reference graph any way it likes in its own memory. It determines when memory can be freed using any of the classic algorithms, and sends messages back to the main processor to indicate which memory regions can be freed. This has some nice characteristics: the co-processor does not necessarily disturb the cache of the main processor (it can have its own memory). Garbage collection is transparent as long as the co-processor can process mutations at a faster rate than the main processor can produce them. The queue handles cases where the mutation rate is temporarily faster than this.
- aidenn0 7y agoThat would seem to eliminate all moving GC algorithms though.
- xenadu02 7y agoNot really though you may want to. The main CPU could compact memory if it felt fragmentation was high enough to warrant it. By its nature the traffic is asynchronous so the CPU won't do a perfect compaction (it won't have processed all the inbound messages to mark garbage as free) but it would probably be close enough. The big win here is making it all async. No need to stop the world, issue write barriers to application activity, or otherwise synchronize for GC. You'd have a bigger highwater mark for used memory but otherwise the main CPU can pickup the "the memory from x to y is now free" messages on any thread and whenever it is convenient.
- aidenn0 7y agoCompacting without tracing can't really be done, and not all moving algorithms are collect and compact (though collect and compact seems to have won over other moving algorithms in the SMP erra)
- _bxg1 7y ago> globally this represents a large amount of computing resources. Much of which would just sit idle otherwise, on client machines. Of course, the energy savings still apply. > He also points out that many garbage collection mechanisms can result in unpredictable pauses, where the computer system stops for a brief moment to clean up its memory. This is more of a hard barrier that's being solved. All in all pretty cool idea, but I think the impact would be different from what's discussed here. Truly high-performance computing is already written in non-GC languages. This hardware would give medium-intensity GC programs (read: web servers on JVM, .NET, Node, Ruby) a boost, and could also allow some higher-but-not-peak intensity software to be written with GC where it might not be today (games come to mind), although that could actually encourage more energy usage than what it would save.
- p_l 7y agoWe wrote latency-sensitive and high-performance code in GCed languages back in 1980s - avoiding pauses or having predictable latencies (in fact, more predictable than usual manual memory management!) are more of "we don't teach people how to program" rather than issue with GC. As for energy savings, many garbage collectors have amortized energy use lower than malloc/free. Even pretty simple ones (some of the simplest I've seen beat it so hard it's not competition, but they are specific to their application)
- gdwatson 7y agoCould you explain further or give links to more information? I'd love to read about old-timey techniques for programming in GCed languages.
- jayd16 7y agoHis point is it's not rocket science. Preallocate and pool what you might need, don't call new in your tight loop. Change the GC algorithm to something that never runs unexpectedly. If you're still allocating such that you need GC eventually, manually GC at an appropriate time like a load screen.
- unictek 7y agoCould this be applied to Chrome V8 for Javascript memory garbage collection?
- ben509 7y agoThat should be an ideal application. Usually the trouble with 3rd-party garbage collection is that it has to discern pointer vs. any other machine word. That's more of a problem with C; this is why the Boehm GC library calls itself "conservative". A runtime like V8 can follow a spec when it allocates memory so everything is properly marked.
- didibus 7y agoMaybe I'm naive, but with multi-core CPUs, and parallel GCs, isn't it somehow the same? One core is mostly only used for GC, while the others do other things? Edit: I guess they mention their chip itself can do it at a high level of parallelism, so that's probably one more advantage. But CPUs with additional slower cores and a lot more cores are in the works as well.
- muxr 7y agoBut the GC has to do it in a thread safe way which involves locking/synchronization. Otherwise you get nasty race conditions.
- Causality1 7y agoOk, so saving 15 percent of 10 percent of power use by changing both how we build processors and how we write software. Doesn't seem worth it.
- tybit 7y agoSeems worth it if the software developed is a fraction of all software written. Sure your VM/runtime will have to be refactored, but not the millions of programs that run on it.
- arcticbull 7y agoIMO garbage collection is the epitome of sunk cost fallacy. Thirty years of good research thrown at a bad idea. The reality is we as developers choose not to give languages enough context to accurately infer the lifetime of objects. Instead of doing so we develop borderline self-aware programs to guess when we're done with objects. It wastes time, it wastes space, it wastes energy. If we'd spent that time developing smarter languages and compilers (Rust is a start, but not an end) we'd be better off as developers and as people. Garbage collection is just plain bad. I for one am glad we're finally ready to consider moving on. Think about it, instead of finding a way of expressing when we're done with instances, we have a giant for loop that iterates over all of memory over and over and over to guess when we're done with things. What a mess! If your co-worker proposed this as a solution you'd probably slap them. This article proposes hardware accelerating that for loop. It's like a horse-drawn carriage accelerated by rockets. It's the fastest horse.
- davedx 7y ago> It wastes time, it wastes space, it wastes energy. But all of these are much cheaper than developer labour and reputation damage caused by leaky/crashy software. The economics make sense. Anecdotally, I spent the first ~6 years of my career working with C++, and when I started using languages that did have GC, it made my job simpler and easier. I'm more productive and less stressed due to garbage collection. It's one less (significant) cognitive category for my brain to process. Long live garbage collection!
- arcticbull 7y ago> But all of these are much cheaper than developer labour and reputation damage caused by leaky/crashy software. The economics make sense. Of course, with traditional languages, that's the trade-off we're being asked to make. That's my point! We need to develop languages that accurately encapsulate lifetimes statically so that we can express that to the compiler. If we do, the compiler can just make instances disappear statically when we're done with them -- not dynamically! Like Rust but without the need for Rc and Arc boxes. Rust gets us 80% of the way there. The truth is with most of the Rust I write, I don't have to worry about allocation and deallocation of objects, and it happens. It's almost ubiquitous. We need to extend that to 100%. > Long live garbage collection! Long live the rocket powered horse!
- jondubois 7y ago>> consumes a lot of computational power—up to 10 percent or more of the total time a CPU spends on an application. I stopped reading there. 10% is nothing. For such a useful feature as automatic garbage collection, for the vast majority of applications, I'd gladly give away 50% of the CPU. In terms of ensuring code correctness and robustness, if I had to choose static typing or automatic garbage collection, I'd pick garbage collection every time. It adds a lot of value in terms of development efficiency and code simplicity.
- ddebernardy 7y agoThere are times when it counts... https://www.youtube.com/watch?v=JEpsKnWZrJ8 https://www.youtube.com/watch?v=JEpsKnWZrJ8
- lemagedurage 7y agoI can agree with that but also you missed the point of the article by not reading along there.
- gmueckl 7y ago10% is a lot. That's at least a few big cloud data center's worth of servers at this point. I'd gladly give the operators a reason to shut these down and save electricity.
- jillesvangurp 7y agoWorse, it's a tradeoff and not a constant overhead. As many Java based server products show, things are fine if you do things such that you avoid garbage collection. GC is only problematic if you are doing bad things like constantly creating lots of objects, or worse, keeping them around for too long. E.g. Elasticsearch uses a lot of memory mapped files these days instead of heap memory (which they used more heavily in the past). That has done wonders for stop the world garbage collect cycles, which used to be a source of cluster inconsistencies because nodes tend to drop out of the cluster when they stop responding due to garbage collection issues. These days that's much less of an issue. On servers, idling CPUs is the norm. We run our java servers on t2s in amazon. CPU throttling is not a problem for us. Our servers run out of IO before they run out of CPU. Garbage collects are not an issue. The little there is seems to have little or no impact on response times. I'd say the innovation in this space comes from new approaches like e.g. Rust with its borrowing mechanisms or functional programming where due to everything being stateless, there's no need to garbage collect that much. It would be interesting to see other languages with borrowing mechanisms; it doesn't sound like something that can easily be retrofitted to existing ones.
- qwsxyh 7y agoI don't care about being absolutely fast when writing code. The convienience of not having to care about memory management is far more important to me. That's why I like GCs.
- nottorp 7y agoHmm so there's a coprocessor that does the GC... doesn't it need to lock the memory away from the main CPU while it does that? And doesn't this lead back to unpredictable pauses and slowdowns?
- simen 7y agoThey're comparing to an in-order CPU. Given that most CPUs are out-of-order (at least of the non-embedded variety, and GC is less used in such applications anyway), it would be better and more intellectually honest to actually compare to a typical CPU that performs GC. They kind of address this in the paper but only in a short aside: "Note that previous research [1] showed that out-of-order CPUs, while moderately faster, are not the best trade-off point for GC (a result we confirmed in preliminary simulations)." So they don't quantify what any of this means. I think it's an interesting idea, but it doesn't bode well when they seemingly choose the wrong target for comparison and hand-wave away the difference as insignificant.
- reitzensteinm 7y agoThe comparison at least in the abstract is energy efficiency. It's quite likely that a small in order CPU is very good at chasing dependent pointers around the heap for its power consumption. Imagine a linked list. Each pointer access is likely to miss to main memory, and no concurrency is possible. Both the highest and lowest end cores will sit around making a single request every 80ns. They claim that the comparison was to the best alternative and I'd probably take them at their word barring any specific evidence.
- klodolph 7y agoI think an interesting comparison here is GPU cores, where a core will get blocked on a memory access and it will switch out its state for another. It looks like this is the approach here, which is a bit more aggressive than ordinary out-of-order approach. It's less ordered.
- daemin 7y agoOne talking point I'd like to ask is: For small short lived scripts and applications, do we even need to free any memory these days? For example you write a script which takes several seconds to execute, moves files, computes stuff with strings, etc. Should we really invest time and effort in the script interpreter to free the memory, where instead we can just exit normally and let the OS handle the clean up. I would imagine this kind of paradigm could be much faster to run because of less runtime work being performed. The allocator used could also be a simple linear allocator which just returns the next free address and increments the pointer. If using multiple threads there could be one per thread. What do people think of this?
- jjaredsimpson 7y agohttps://openjdk.java.net/jeps/318 https://openjdk.java.net/jeps/318 Java epsilon gc is a no op gc
- obruchez 7y agoThis reminds me a bit of Erlang's "let it crash" philosophy.
- bpicolo 7y agoFor short lived scripts and small programs you’re not expecting to need bleeding edge performance, so it doesn’t matter much either way.
- lioeters 7y ago> For small short lived scripts and applications, do we even need to free any memory these days? ..we can just exit normally and let the OS handle the clean up. That makes me think: why couldn't larger programs be composed of many such small, short-lived scripts/processes that give up all allocated memory upon exit? I suppose there could be accumulated overhead for starting many such processes, and also the issue of allowing shared memory spaces that are explicitly not automatically freed. I'm way out of my depth in this line of thinking though, so, just speculating.
- alexhutcheson 7y ago
- stephc_int13 7y agoI never understood the need for Garbage Collectors. In my opinion, the difficulties of memory management are extremely overrated. I write code in C/C++ for almost 20 years and I never encountered a difficult bug that would have been avoided with a Garbage Collector. If a coder really has a hard time with manual memory management it means he can't really code, this is a beginner problem...
- maaaats 7y agoI only work in GCed languages, so I don't know how manual memory management works, except segfaults in some courses at university. Guess I should quit my job, as you say I apparently can't really code :/ Thanks for letting me and others here know!
- ndepoel 7y agoIn my experience GC makes programmers sloppy in their resource usage. Just allocate a bunch of memory annnd... whatever, the GC will take care of that. But there are a lot of other resources besides memory that aren't automatically cleaned up like that. So what happens is people forget to close network sockets, forget to unsubscribe event handlers, forget to set certain references to null to actually allow the GC to do its work, etcetera... The existence and over-reliance on GCs has led to a mindset where many programmers are just not aware that what you create must also be destroyed at some point.
- stephc_int13 7y agoLet me rephrase this. In the CS / Programming space, we have quite a lot of difficult problems to solve, in my opinion, memory management is really easy compared for example to multithreading. In fact the so called Garbage Collector do not really solve the problem of memory management, there are still a lot of potential for leaks and poor memory management when GC is used. And it adds a lot of complexity to the language runtime, and it gets in the way, I've seen a lot of discussions about how to avoid triggering the GC, in the end, the solutions are more complex than good old manual memory management. I am not saying that memory management is trivial if it was it could efficiently be handled by the compiler or the runtime. There is no magic bullet solution for easy memory management, one has to choose the right policy considering the context, which is most of the time out of reach for the compiler. The context is mostly the expected lifetime of the memory allocation.
- deleted 7y ago[deleted]
- sasdsd 7y agosss
- thekingofh 7y agoYou can usually precisely control garbage collection by turning it off or forcing it to run. The cognitive load to handle memory manually is not insignificant. If you control memory manually, you eventually end up designing some kind of mechanism like ref counting or something else to handle memory cleanup automatically. And there's significant reasoning that ref counting might not be the most desired solution for all use cases. Best is a combination of the ability to handle memory manually, with some more automated garbage collection when there's a need to write stuff that doesn't necessarily have to be the absolute fastest. Kitchen sink languages like C++ tend to have both and don't force the developer in either direction. Best would be to #define out 'new' and make manual memory handling explicit.
- wolfspider 7y agoThe Kiwi scientific accelerator uses a similar approach with FPGAs I believe: https://www.cl.cam.ac.uk/~djg11/kiwi/ https://www.cl.cam.ac.uk/~djg11/kiwi/
- DenisM 7y agoObjective-C ARC (automatic reference counting) solved the problem neatly for my iOS apps. Is there some overhead? Maybe, but it's neatly spread out through the entire application life time, so there is rarely[1] a UI-freezing stutter associated with GC. To reduce the overhead I turned off thread-safety and simply never access the same objects from more than one thread (object has to be "handed off" first if it comes to that). One wart on the body of ARC is KVO, which I avoid like a plague for many other reasons anyway. The other wart is strong reference loops. This can be solved by the app developer by designing architecture around the "ownership" concept (owners use strong references to their ownees, all other links are weak references). This is a good idea in itself as it increases clarity of the program. I do make an occasional slip, which is where I need to rely on Instruments, and I do wish I had better tools than that, something more automatic that would catch me in the act. Maybe a crawler that looks for loops in strong references during the development process but is quiet in release builds. Or at least give me a pattern to follow that makes it easy to catch my errors. For example, we could assign a sequential number to each allocated object, and only higher-ranked object could strongly refer to lower-ranked object. This won't work for everyone but I wouldn't mind fitting my app to this mold if that gave me immediate error when I slip. [1] if you release a few million objects all at once it may stutter for a second. Could be handed off to a parallel thread maybe.
- fwip 7y agoARC is just garbage collection that doesn't always work (circular references).
- klodolph 7y agoTrue! And ARC in Rust suffers from the same problem (note that while ARC in Rust and ARC in Swift are the same thing, the "A" happens to stand for different words in each case).
- steveklabnik 7y agoThey’re not the same thing; Swift is “automatic” and Rust is “atomic”; Swift’s ARC is implemented the same way as Rust’s Arc, but Rust’s is manual.
- jasonhansel 7y agoDidn't the old Lisp machines also do this?
- SlipperySlope 7y agoIn Java, I created thread local resource pools including strings which eliminate garbage collection in sensitive routines. Of course it’s much faster in java to perform pooled string comparison with ==. Likewise I always use the indexed version of a for loop to avoid the iterator otherwise allocated. GC in java is great for non-priority code which is most of the application.
- ekianjo 7y ago> but the automated process that CPUs are tasked with consumes a lot of computational power—up to 10 percent or more of the total time a CPU spends on an application. Is that even a problem when most CPUs are idle 90% of the time even when doing typical daily tasks?
- rkrzr 7y agoIt is a problem if you have a server process running that you want to be extremely responsive (latency <100ms) and that then suddenly decides to do garbage collection for a minute or two before answering incoming requests.
- jakeinspace 7y agoMy day job is writing C for on an embedded real-time system. No manual memory management necessary... because we're forced to declare all struct and array sizes at compile time! Not a malloc or free in sight. Obviously, it's extremely limiting - pretty limiting as far as algorithms beyond "read data off bus, store in fixed array, perform numeric calculation, write back to bus." But I've gotta say, it's pretty freeing to write C in such a limited environment.
- kazinator 7y agoI'm skeptical; GC is closely tied to programming language run-times. How is some accelerator going to know which pointers in an object are references to other GC objects and which are non-GC-domain pointers (like handles to foreign objects and whatnot)? How does the accelerator handle weak references and finalization? People aren't going to massively rewrite their language run-times to target a boutique GC accelerator.
- mar77i 7y agoI remember ruby had some approach with reusing previously allocated yet out of scope objects. I can very well imagine taking this concept above and beyond, and having virtually separated stacks for each type...
- jokoon 7y agoIt reminds me of the days I was reading Knuth's quote "97% of the time premature optimisation blabla" every time someone was trying to make something faster. CPUs are not getting faster, yet it seems using tools that makes things run faster are somehow taboo. Wirth's law: Wirth's law is an adage on computer performance which states that software is getting slower more rapidly than hardware becomes faster. Why is java taught in university, and why is this language considered like some kind of standard? Most OSes are written in C, yet most of silicon valley frowns upon writing C because of arrays. Even C++ is getting a bad reputation.
- beersigns 7y agoThink most companies resist using lower level languages due to their ultimate purpose being to build products that provide value and sell them. They are generally far less concerned with the technical details and conciseness of the implementations. "Good enough" is a very squishy term but for most companies, for better or for worse, that bar is pretty low. There are plenty of industries that focus on lower level langs and use them pretty well but it's not the norm for big corporations who value rapid turn around over all other factors.
- pkulak 7y ago> frowns upon writing C because of arrays Do you mean horrifying security flaws?
- jokoon 7y agoYou're right, but you cannot accuse the entire language of being insecure, ultimately it's the responsibility of the developer. Also you can't always say we don't use C only for security reasons, as other languages also have their security issues. There are many modern ways to avoid those flaws. Like someone answered, it boils down to a matter of development cost. I'm also quite skeptical when people always rise the objection of security when writing software. Security is not so simple, and so far it's its own specialty, and pretending that it's worth it to make things slower and that the security gain is actually there, is not really completely accurate. Security is almost a post 9/11 paranoia knee jerk reaction.
- kilon 7y agoThe irony of the thing is that in manual memory management languages you end up doing your own garbage collectors and in garbage collector languages you end up doing your own manual management. Unfortunately if you look in a language to solve such complex problems you are heading straight to severe disappointment land. Same shit different package. I still prefer dynamic languages by a long margin because of their ability to do decent metaprogramming and reflection which is essential for managing any form of data. Pick your poison and enjoy the hype while it lasts.
- k__ 7y agoHasn't Rust basically solved that problem? But yeah, legacy stuff could profit from this.
- klodolph 7y agoNo, because it is in general intractable to figure out object lifetime at compile time. Rust has just solved the problem for more cases than, say, C++ does, or perhaps just with more rigor.
- Aardappel 7y agoWant to get away from garbage collection, retain safety, but think Rust is too invasive? Try compile time reference counting: http://aardappel.github.io/lobster/memory_management.html http://aardappel.github.io/lobster/memory_management.html