7 ms·
I'm convinced there was another style of C API where the callee would malloc a struct, populate it, and free it immediately before returning a pointer to it. Of
by smarks 4y ago
I'm convinced there was another style of C API where the callee would malloc a struct, populate it, and free it immediately before returning a pointer to it. Of course, the only way the caller can use the result is after it has been freed.
Naturally there was a dependency on the exact behavior of the allocator, specifically, that it had to leave a freed block of memory untouched sufficiently long for the caller to be able to use the results. I seem to recall the stipulation was that freed memory was left untouched until the next memory allocation operation. The caller also had to be careful about using or copying the results immediately, before doing too much work.
I have dim memories of people talking about this sort in thing at university in the early 1980s; we were a 4.2 BSD shop. I also recall debugging some old C source code (srogue, which also has BSD heritage) decades later, and encountering use-after-free crashes. There were several instances of this. There were too many to be accidental; it seemed deliberate to me.
I suspect the reason for this "technique" was to relieve the caller of the burden of freeing the memory. It allowed the caller to return variable-length data easily, which couldn't be done if the pointer was to a static data area. And finally it relieved the callee of defining an explicit "free" API.
Frankly I think this is a terrible API style. However, code that used it properly and that was sufficiently careful would actually function properly. But it seems like an incredibly fragile and sloppy way to design a system.
- giomasce 4y agoIt also fails on multithdreaded programs, unless you add even more assumptions on the allocator.
- smarks 4y ago"Multidreaded programs" :-) This was at least a decade before multithreading in C and Unix. But yes this "technique" would have failed miserably in a multithreaded environment.
- Jweb_Guru 4y agoI mean, it wouldn't even work with interrupts, technically (though an interrupt that plays "nice" shouldn't allocate).
- dragontamer 4y agoSounds like just a common mistake where people used the "stack" accidentally. Ex: struct someStruct* badfunc(){ struct someStruct toReturn; toReturn.a = foo(); toReturn.b = bar(); return &toReturn; } In most people's code, this would probably work... struct someStruct a = *badfunc(); func2(); // This will overwrite the "toReturn" // from the last call, but as long as the // struct was copied before any other function // call, you're probably fine though in // undefined-behavior land ----------------- Either that, or you're talking about strtok (and other non-reentrant functions).
- smarks 4y agoIt was definitely malloc'd memory, as I remember removing free() calls from the callee and adding them to the caller.
- dragontamer 4y agoThat's the old "source/sink" pattern. "Source" functions return a malloc'd pointer. You have to manually call free on it, or pass it to a "sink" function (which internally calls free). In C, this is seen in strdup. C++ took this pattern and formalized it into auto_ptr<>, and later unique_ptr<> when RValue references became a thing in C++11.
- pitched 4y agoI’ve used this before in embedded code to save rom space by not having to include malloc. The stack _is_ the heap! Just have to be very careful about when you can call another function. The best version of this is where you allocate a block on your stack then pass that as a pointer up to the next function to use. The one who owns the memory is the one who allocates it (Rust style?). Or, have the linker allocate global blocks works too.
- drran 4y agoJust bump the stack pointer after call to function and use the allocated memory without risk of overwrite, like alloca() does.
- Jach 4y agoIn GC languages a common approach is to have "finalizers" to make something like this a possible and sometimes convenient way of dealing with a foreign API, I wonder if what you saw was something similar? The idea is to allocate the foreign memory, then make a finalizer which is just a hook that will (eventually) call the foreign free only when some object (like a wrapper for the foreign memory) is collected. Something similar could be hidden behind some preprocessor macros, with guarantees only until the next OUR_MALLOC... The problems for the GC languages tend to be fewer but if the wrapper is out of scope but someone grabbed and maintains a hold on the foreign memory directly, they're playing with fire as for when the GC will execute the finalizer hook and make that memory invalid. It's also a frustrating technique when foreign APIs -- particularly in certain graphics contexts -- require allocation threads to be the same as freeing threads, and of course depending on the implementation of free and the GC it might be an expensive operation to have a bunch of them suddenly happen at once when all you were expecting was a new native object and not a bunch of GC work behind the scenes.
- smarks 4y agoDefinitely not GC. This was K&R C on BSD Unix around 1984.
- pjmlp 4y agoThat is why finalizers are yesterday solution, most modern GC based languages have eventually catched up with Common Lisp and offer region based resource management (try-with-resources, use, using, defer, with,...), and in some cases trailing lambdas, which completly hide the resource management from the consumer. For scenarios like you're describing, .NET has SafeHandles for example.
- mtlmtlmtlmtl 4y agoMan this sounds like something some students came up with after partaking in the ganja. And it coming out of Berkeley in the 80s certainly tracks with that... Not a good way of doing things. I mean, have fun using Valgrind. Or switching out libc, etc. And what about key material? There you would still have to do a second step of zeroing or junking the memory when done with it anyway.
- smarks 4y agoYes, grad students at Berkeley in the early 1980s. For some reason I associate this technique with Bill Joy (who obviously was a major influence on a lot of what went into the 4bsd releases). However I have no evidence of this, nor whether any or what kinds of substances might have been involved.
- derefr 4y agoSounds like it was just using the heap to badly imitate returning variable-length data on the stack under a callee-preserved calling convention. (Callee writes the variable-length data inside their own stack frame, pops the stack frame in the function epilogue, and “leaks” the pointer and size of the data in caller-expected return registers. Caller uses the dangling data — carefully not pushing to the stack until it has finished. Everything works out.)
- bitwize 4y ago> I'm convinced there was another style of C API where the callee would malloc a struct, populate it, and free it immediately before returning a pointer to it. Of course, the only way the caller can use the result is after it has been freed. That's more than a bad API design, that's undefined behavior -- squarely in nasal-demons territory. Depending on the compiler, the callee, the caller, or the entire observable universe can be optimized away into a no-op.
- Athas 4y agoWell, this sounds like it was before ANSI C, so there was no defined notion of undefined behaviour - I think that term of art came with the later standardisations. And if it was written to run on a specific OS or compiler (4BSD), one can argue that it was a really bad design, but it worked reliably on what was essentially the implementation-defined platform it was targeting.
- smarks 4y agoOh yes, totally undefined. But consider the time frame, early 1980s K&R C on 4bsd Unix on a VAX. This predates ANSI/ISO C and Posix. It even predates “nasal demons.” There was no specification; or perhaps the implementation was the specification. The fact was that at some point the bsd allocator did leave freed memory untouched until the next memory allocation operation, and so people wrote programs that relied on this. Again, I’m not defending this, but this seemed to be the way that some people thought about things. I even remember questioning some code that used memory after having freed it. It was explained to me that this was “safe” because the memory wouldn’t be modified until the next malloc! Also, remember that BSD was the system where if you did printf("%s", NULL); it would print “(null)” instead of getting SIGSEGV. And in general, deferencing a null pointer would return zero. The rationale for this was that it “made programs more robust.” (Again, I disagree, don’t argue with me about this!) One more common technique from the BSD era (srogue again, but other programs did this too). To save the state of a program, write to a file everything between the base of the data segment to the “break” at the top of the data segment. To restore, just sbrk() to the right size and read it all back in, overwriting everything starting at the base of the data segment. I always found it surprising that this worked, but it worked often enough that people did sh!t like this.