13 ms·
I see where you're coming from but IMO this "Cheshire cat" idiom to hide the implementation details is not exactly like private, it fact it can do things that p
by simias 3y ago
I see where you're coming from but IMO this "Cheshire cat" idiom to hide the implementation details is not exactly like private, it fact it can do things that private can't do, and doesn't do things private does.
The advantage of hiding your state behind an opaque struct with builders and accessors is that you can change the size and layout of said struct without it being a breaking API change. The code remains binary compatible even, no need for a recompile if you're shipping a shared lib. This is something just using private members doesn't achieve since with private members the compiler still knows and uses the layout of the struct, it just forbids access to it.
That's why you can even find C++ libraries use this idiom even though C++ obviously has `private`. It's about having a stable, opaque API.
On the other hand because of this added indirection, there's usually a greater performance hit to accessing these opaque structs since code can't be inlined. With private since the compiler can still see inside the struct, it's able to more aggressively optimize the code. You can also store the objects directly on the stack without requiring malloc.
IMO the right way to have private members in C structs is... to document that members shouldn't be touched directly, perhaps using a special naming convention or embedding the publicly-accessible members in a dedicated sub-struct to prevent confusion.
- zokier 3y ago> The code remains binary compatible even, no need for a recompile if you're shipping a shared lib. This is something just using private members doesn't achieve since with private members the compiler still knows and uses the layout of the struct, it just forbids access to it. There is somewhat common PIMPL idiom to work around the binary compat issue. Iirc there were some macros floating around to make it easier to manage.
- maleldil 3y agoWhat do the macros do? Isn't this as easy as forward declaring the impl class, adding std::unique_ptr<Impl> as a private field and have public methods refer to the field? I'm struggling to understand why macros would help here.
- saghm 3y agoI'm also only familiar with this idiom in C++, but based on the description in the parent comment, I suspect that this is sometimes used in C too, in which case you obviously can't use unique_ptr or private fields; maybe macros might be a way to avoid having to write a bunch of boilerplate to achieve a similar effect?
- comex 3y agoPerhaps the macros help you forward methods on the outer class to methods on the impl class? While your approach of having public methods refer to the field also works, it’s nice to have public and private methods in the same place (the impl class’s definition) and using the same syntax (neither having to go through the impl field).
- maleldil 3y ago> it’s nice to have public and private methods in the same place (the impl class’s definition) In my experience, the impl class is usually defined in the main class's cpp file.
- Joker_vD 3y ago> IMO the right way to have private members in C structs is... to document that members shouldn't be touched directly, perhaps using a special naming convention or embedding the publicly-accessible members in a dedicated sub-struct to prevent confusion. Reminds me of that one time when glibc broke the whole of Debian for s390 architecture by changing the fields in the jmp_buf struct (which is public): [0]. [0] https://lwn.net/Articles/605607/ https://lwn.net/Articles/605607/
- bluetomcat 3y agoTo achieve a reasonable level of encapsulation in C, a header file must be seen as a public-only interface. It should declare only the structs that are relevant for the user of the module. If that's "struct my_module_handle { ... }", declare it and document the corresponding accessor and modifier functions. Everything else must reside in the C source file with internal linkage (static storage class). The whole source file is your implementation. There is an anti-pattern where header files are used for all the declarations needed internally by the source file. Including (pasting verbatim with the preprocessor) that file from another module would bring in all the unnecessary declarations.
- simias 3y agoI think what you say makes complete sense at module-level (as in, for a standalone lib for instance) but I never bother segregating things internally within a lib/module/exe and rely on good documentation and coding practices to avoid having member mutations all over the place. If I code in Rust or C++ I can use namespacing and public/private to give every single object in the codebase a clean interface, but in C doing that is just frustrating, not to mention potentially inefficient.
- hbossy 3y agoThis is how it's supposed to be done but you always end-up moving them to header just to make writing unit tests less painful.
- 10000truths 3y agoThis is a smell. Your unit tests should not have to rely on internal implementation details.
- Dwedit 3y agoLink-time optimization means that you're probably not going to take that much of a performance hit. But yes, opaque structs do enforce that it will be treated as a plain pointer, and the compiler (usually) cannot treat it as an aggregate of variables.
- Athas 3y agoIf a linker did that, changing the layout would be an ABI-breaking change. I think this opaque struct design is most common for dynamically loaded libraries, where link time optimisation does not occur (unless dynamic linkers got a lot more fancy recently).
- c-linkage 3y agoI like the way that Windows does it, where they have as the first element of the struct a double-word size (dwSize) element that records in 32-bits the size of the structure. The size essentially acts as a version identifier, as long as you never rearrange the fields and only append fields for new versions. The opaque functions test the value of the dwSize element to see what actions can be performed on the object. The code that you develop can still access the member fields directly, and those accesses can be inlined and optimized aggressively by the compiler.
- bluejekyll 3y agoThis implies a branch statement in every function call, doesn’t it?
- veltas 3y agoThere's zero cost abstractions and then there's zero features abstractions.
- speed_spread 3y agoThe cost of that branch is insignificant next to that of the syscall you opted to make.
- zik 3y ago> is not exactly like private That seems like a straw man to me. He never said it was exactly equivalent - he said it provided encapsulation and isolation. 'private' is another language's mechanism which does that but they're obviously not identical.
- juunpp 3y ago> On the other hand because of this added indirection, there's usually a greater performance hit to accessing these opaque structs since code can't be inlined. Is this still true? Link-time optimization, which includes inlining, seems to be all the rage these days.
- simias 3y agoFor static linking quite possibly, although you could still have un-optimizeable sources of overhead because of having to indirect through pointers everywhere, for instance in nested structs. Something like a->b->c->d is going to be more expensive than a.b.c.d, for instance. I'm also not sure that dynamic linkers are typically smart enough to do this type of optimization, but I admit that I'm not familiar with the state of the art in dynamic linking technology.
- matheusmoreira 3y agoThe Java analogue of opaque structures is factory methods. Their only purpose is to hide the new keyword. That's necessary precisely because new introduces a hard dependency on the constructed type at the binary interface level, it gets literally emitted into the byte code. The supertypes returned by those methods are just there to serve as compatible pointers to the real types.
- rewmie 3y ago> That's why you can even find C++ libraries use this idiom even though C++ obviously has `private`. It's about having a stable, opaque API. In C++ this is a popular idiom and a standard technique covered by multiple references under the name pointer to implementation (pimpl). This is not a C thing beyond the point that C has pointers.