4 ms·
In your example, the objects are all the same size. That would certainly be easy. If you have three local objects that are 8 bytes, 16 bytes, and 32 bytes… if
by coder543 5y ago
In your example, the objects are all the same size. That would certainly be easy.
If you have three local objects that are 8 bytes, 16 bytes, and 32 bytes… if you do a single 48 byte allocation on the TLAB, how can the GC possibly know that there are three distinct objects, when it comes time to collect the garbage? I can think of a few ways to kind of make it work in a single buffer, but they would all require more than the 48 bytes that the objects themselves need. Separate TLAB arenas per size class seem like the best approach, but it would still require three allocations because each object is a different size.
I understand you’re some researcher related to Truffle… this is just the first I’m hearing of multiple object allocation being done in a single block with GC expected to do something useful.
- chrisseaton 5y ago> If you have three local objects that are 8 bytes, 16 bytes, and 32 bytes… if you do a single 48 byte allocation on the TLAB, how can the GC possibly know that there are three distinct objects, when it comes time to collect the garbage? Because the objects are self-describing - they have a class which tells you their size. object_a = tlab object_a.class = ClassA object_b = tlab + 8 object_b.class = ClassB object_c = tlab + 16 object_c.class = ClassC tlab += 48 check tlab limit
- coder543 5y agoOk, so it’s not as simple as bumping the TLAB pointer by 48. Which was my point. You see how that’s multiple times as expensive as stack allocating that many variables? Even something as simple as assigning the class to each object still costs something per object. The stack doesn’t need self describing values because the compiler knows ahead of time exactly what every chunk means. Then the garbage collector has to scan each object’s self description… which way more expensive than stack deallocation, by definition. You’re extremely knowledgeable on all this, so I’m sure that nothing I’m saying is surprising to you. I don’t understand why you seem to be arguing that heap allocating everything is a good thing. It is certainly more expensive than stack allocation, even if it is impressively optimized. Heap allocating as little as necessary is still beneficial.
- chrisseaton 5y agoEvery stack allocation scheme I've seen creates a fully reified object though, with a class pointer as normal. You may be confusing with scalar replacement of aggregates, which is a separate concept. https://chrisseaton.com/truffleruby/seeing-escape-analysis/ https://chrisseaton.com/truffleruby/seeing-escape-analysis/
- coder543 5y agoGo does not put class pointers on stack variables. Neither does Rust or C. The objects are the objects on the stack. No additional metadata is needed. The only time Go has anything like a class pointer for any object on the stack is in the case of something cast to an interface, because interface objects carry metadata around with them. These days, Go doesn’t even stack allocate all non-escaping local variables… sometimes they will exist only within registers! Even better than the stack.
- chrisseaton 5y ago> These days, Go doesn’t even stack allocate all non-escaping local variables… sometimes they will exist only within registers! Did you read the article I linked? That's what it says - and this isn't stack allocation, it's SRA. Even 'registers' is overly constrained - they're dataflow edges.
- coder543 5y agoI was addressing your statement that seemed to say class metadata is included in every language when dealing with stack variables. I trusted your earlier statements that this was definitely the case with Java/Truffle. Misunderstandings on my part are entirely possible. Sorry I haven’t had time to read your article. It’s on my todo list for later.
- coder543 5y agoGo isn't just doing SRA, as far as I understood it from your article, though it is certainly doing that too. Go will happily allocate objects on the stack with their full in-memory representation, which does not include any kind of class pointer. Here is a contrived example that creates a 424 byte struct: https://godbolt.org/z/5cKh7xTzq https://godbolt.org/z/5cKh7xTzq As can be seen in the disassembly, the object created in "Process" does not leave the stack until it is copied to the heap in "ProcessOuter" because ProcessOuter is sending the value to a global variable. The on-stack representation is the full representation of that object, as you can also see by the disassembly in ProcessOuter simply copying that directly to the heap. (The way the escape analysis plays into it, the copying to the heap happens on the first line of ProcessOuter, which can be confusing, but it is only being done there because the value is known to escape to the heap later in the function on the second line. It would happily remain on the stack indefinitely if not for that.) It's cool that Graal does SRA, but Go will actually let you do a lot of work entirely on the stack (SRA'd into registers when possible), even crossing function boundaries. In your SRA blog post example, when the Vector is returned from the function, it has to be heap allocated into a "real" object at that point. Go doesn't have to do that heap allocation, and won’t have to collect that garbage later. Most of the time, objects are much smaller than this contrived example, so they will often SRA into registers and avoid the stack entirely… and this applies across function boundaries too, from what I’ve seen, but I haven’t put as much effort into verifying this.