4 ms·
A modern alternative to C's FFI needs to be developed, perhaps a standard like gRPC but as a compiler extension. C's FFI contains so little information due to i
by ampdepolymerase 6y ago
A modern alternative to C's FFI needs to be developed, perhaps a standard like gRPC but as a compiler extension. C's FFI contains so little information due to its weak type system that it is wholly unsuitable for higher level languages. Yet we use it because it is the lowest common denominator. The C FFI has cost hundreds of hours of lost productivity due to having to re-annotate all the types in the non-C caller. This is before we even take into account the mess that is function pointers and types generated using the pre-processor. And no, tools like bindgen and c2rust are stop gaps, not a real solution. A lot of low level code cannot be effectively called safely in a higher level language at all without hundreds of hours of manual work wrapping and annotating each individual function. We need a new standardized, zero cost, cross language interop format that functions on the ABI/memory level. Rust's Cxx is a good start but it doesn't scale m:n with regards to new languages. Every new language has to implement a custom wrapper for calling another high level language if they do not want to lose 80% of the benefits of the respective language's type system when passing through the primitive C FFI. For example, calling Rust from Julia, calling Julia from V lang. The lack of a common interop format is hurting the developer ecosystem as more proliferation of programming languages simply leads to greater fragmentation. The web solved this a long time ago with JSON APIs, generated OpenAPI clients (and now Graphql), and to a lesser extent, gRPC. We need to do better when it comes to systems programming.
- simias 6y ago>We need a new standardized, zero cost, cross language interop format functions on the ABI/memory level I don't think that's doable unless the underlying language already has the same ownership and lifetime guarantees as Rust does, otherwise these bindings have to protect against memory issues by effectively taking ownership of the data to garantee its lifetime, which is not zero cost. I simply don't think that you can auto-generate zero-cost safe C or C++ bindings automatically.
- estebank 6y ago> I simply don't think that you can auto-generate zero-cost safe C or C++ bindings automatically. As stated, that's likely an accurate statement, but if you separate the requirements for the bindings the problem looks way more tractable: - auto-generate: I'm pretty positive that this point is only hard given the following, and I feel you would agree with that assessment, so I'll just focus on the other ones. - zero-cost: doing zero-cost ffi is the best option, but a lot of advanced constructs that people want to use cannot be well represented in that way because they rely on the internal consistency of the host language. That means that dynamically loading a library written in Rust might make it so that you can only interact with ADTs, trait objects and whitelisted provided traits on those ADTs, potentially only through trait objects. That's not zero-cost, but it is super powerful. ObjectiveC has lived its entire existence going through vtables and that has even enabled runtime introspection that was super useful. - safe: safety can be accomplished easily enough if you don't constrain yourself to zero-cost or if you put constraints on what is supported for zero-cost ffi. You could say that method calls on trait objects is safe, but type parameters are out of scope because ffi and monomorphization are not compatible. You could restrict yourself to either no borrows crossing ffi, or no mutable data (only internal mutability through trait methods) allowed. - C or C++: I would focus on other languages, personally. Having easy interop with Java, C#, Python and JavaScript would open opportunities that do not have anything to do with systems programming per-se, but that would make it more likely that the next numpy isn't written in pure C. - automatically: the need for extra annotations in either end of the interop simplifies part of the problem and might even help with documentation. Adding some light ownership information to APIs in non-Rust languages would be great for non-Rust users of those APIs.
- bee_rider 6y agoHey! Surely there's some Fortran left in some of the libraries that back NUMPY.
- chubot 6y agoThis is a very real problem, particularly the M:N issue. But I think it's inevitable you have to do things on both sides of the language boundary, i.e. pushing them towards a lowest common denominator. (e.g. Microsoft's COM, which actually worked pretty well) Really what it pushes you toward is wire protocols and NOT relying on the type system. That is the C and C++ ABIs are related to but different than the APIs (type system). My experience tells me that transparent interop is kind of a pipe dream. The problems always pop up somewhere. By "transparent", I meant "Rust function calls C++ function" and "C++ function calls Rust function" without other metadata/bindings. The codegen becomes a big problem in practice. Fundamentally a lot of languages LOOK the same but they ACT completely differently. A Rust function is not a C++ function is not a Python function is not a Go function. (And funny thing -- as of C++ 11, C++ now has many different notions of "function", because of move semantics). And functions are actually "easy" compared to types (e.g. inheritance vs. typeclasses vs. interfaces) ---- Basically I would say the problem is that you either spend time manually wrapping or annotating your code (as you say), OR you put an ever growing list of heuristics in the code generator, a la SWIG. Those heuristics have bugs, and will make your program unreliable. They're also "someone else's problem", which leads app developers to come up with horrific workarounds. So I go with the simple manual wrapping, and reducing the number of things to wrap by changing the structure of your program. This also has other benefits like efficiency, i.e. crossing the language boundary less often. (Related: I think better build systems can go along way toward solving this problem. Unfortuately there seem to be a lot of language-specific build systems and package managers now, which only exacerbates the interop problem. If you have one language-neutral build system, it's not that bad.)
- 10000truths 6y agoThe whole point of FFI is that it's a lowest common denominator - it's a foreign function interface, meant for interop between two potentially wildly different runtimes, up to and including assembly language. That means it can't rely on quirks like language-specific type information or bounds checking. The only thing that all runtimes on a machine are guaranteed to have in common is that they use the same set of memory and CPU registers, so that is all that FFI can use to define subroutine interfaces. Hence terms like 'ABI', 'calling convention' and 'memory layout'. The good news is that whatever issue you have can probably be solved without resorting to FFI. There are plenty of other interop options - for example, if you control the library, you can use your operating system's IPC mechanisms with an agreed-upon serialization format. Or, if your library and application are within the same ecosystem, you can use language-specific library management features such as Python's import or Rust's crates. FFI is a bit like C itself - if you're reaching for it as a solution to your problem, then your finger is already on the trigger of the footgun.
- ampdepolymerase 6y agoAnd most people used to believe that ergonomic, zero cost memory safety was impossible too. And yet it moves. There is no reason why we have to settle for a poor FFI solution just because it was the choice of our predecessors.
- 10000truths 6y agoMemory safety and FFI operate on two different layers of abstraction entirely. What does it mean to enforce 'memory safety' across the boundaries of two different applications/libraries that may have entirely different ways of keeping track of memory allocations and deallocations? Again, FFI has to be agnostic to such implementation details in order to work, and it does so by forcing you to define interfaces in terms of the lowest possible level of abstraction common to all applications. That means explicitly specifying expected memory layout and calling convention on both the application and library sides, i.e. an ABI.
- jrumbut 6y agoIPC is, I think, the concept that has been neglected and led to the desire for some kind of really intelligent FFI. If you have code in Ruby and in Julia that you want to mix together it's probably because both pieces of code leverage the distinct capabilities of those two languages, which have nice capabilities because they accepted very tradeoffs. Would you really want Julia with Ruby's Number class? It would entirely kill performance. You want it to be an integer maybe, or a double, or sometimes something else. It's going to be application specific is the point, because another user want some different subset of Number's behavior. I don't think there's a shortcut here besides making a very limited or opinionated FFI system.