4 ms·
This is a good rundown of the basics of ML modules, the syntactic aspects, if you will. But it doesn't really give you the important insight: Abstraction. is
by Drup 9y ago
This is a good rundown of the basics of ML modules, the syntactic aspects, if you will. But it doesn't really give you the important insight:
Abstraction. is. powerful.
if you consider the IntMap module provided in the article, you notice that the type `t` is not defined. It is kept "abstract". One defining property of ML modules is that you can indeed keep the definition of types hidden and nothing can break abstraction. The rest of the program can only manipulate the type `t` through the functions provided by the module.
This allows, in practice, to encode pretty much any property you would like. In the case of maps, you can use BST and never have to worry about users breaking the invariants: the user can't even see it's a BST! You can provide any form of validation and be sure that the data you manipulate will stay validated (as long as the module is correct, of course).
Functors allows you to rely on this by giving you a way to abstract away a whole module. This gives you excellent guarantees: if two modules behaves exactly the same, swapping them is semantic-preserving. Said in another way: if I can prove that sets-as-lists and sets-as-bst are functionally equivalents, then I can swap one for the other, and the rest of the program will behave the same!
Refactoring becomes a breeze.
Functors (and modules) makes all the Javaesque reflections and dependency injections completely redundant: modules are (in OCaml) first class objects that you can manipulate and apply as you wish.
It also gives you separate compilation! A lot of people say that the OCaml compiler is blazing fast: this is precisely because OCaml's module system ensure that each module can be typechecked and compiled incrementally.
It pains me that modules get so little appreciation. Many ML programmers get a sort of selective blindness: just like fish in the sea, they don't realize the power of what they're swimming in. Proper modules with actual abstraction are, in my opinion, the single most essential feature for "large scale" programming.
- skybrian 9y agoI agree this is great stuff, but let's not get carried away about types. They don't guarantee compatibility. Two functions with the same type signature can compute different things, and code can advertently rely on this difference. For example, changing a hash function can cause code to break that relied on iteration order of a hashmap. (And with in a large enough codebase and a commonly used, rarely changed hash function, this is probably inevitable.) Also, even if the values returned by the function are the same, you can still observe differences between different implementations by comparing performance, and code can inadvertently depend on performance differences. This can cause breakage or security holes. Still, if you can avoid type errors, you're probably doing well.
- Drup 9y agoI didn't say anything about types. :) I said functionally equivalent. What it means depends on the definition of "functionally equivalent" of course. The first version, which is types, is the most immediate and obvious, but as you pointed out, not that useful. Abstractions, however also works at the semantics level: if A behaves like B trough the interface S, then you can confidently say that for any F (of type S -> S'), F(A) behaves like F(B). So you can just test that A and B are not distinguishable (property testing is very good at this) and feel confident about your refactoring! The definition of behaves like depends on your language. Usually, this doesn't account for performances, as you point out, but only for the observable semantics. This is known as parametricity[1]. If you really want to get into the deep end, [2] demonstrates all that for a (very rich) ML module system. [1]: https://en.wikipedia.org/wiki/Parametricity https://en.wikipedia.org/wiki/Parametricity [2]: http://www.cs.cmu.edu/~crary/papers/2017/mapp.pdf http://www.cs.cmu.edu/~crary/papers/2017/mapp.pdf
- skybrian 9y agoGood point! But leaving types aside, the same argument applies. "Observable semantics" seems to mean "observable according to our semantic model" which is an agreement to pretend that machine behavior outside the model isn't observable and avoid relying on it. This agreement is normally a good thing since it's what allows for the same program to run on different machines, or to compile the same program with different optimizations, or swap in a newer version of a library with a better hash function and claim it doesn't break backward compatibility. Nevertheless, it can be a blind spot as we saw with Meltdown and Spectre, so I wanted to emphasize that this is a useful myth but the world isn't obliged to go along with it. It's often important to observe program behavior that isn't specified by the language.
- Drup 9y agoAh, yes, This is where safety and security differs! Abstraction is about safety: preventing errors and giving additional guarantees, not preventing attacks. :) Nothing I said hold when people are being malicious: Even in OCaml, you have escape hooks that allows you to break past abstraction boundaries and do whatever you want.
- zvrba 9y agoIt seems to me that you can get the same thing in C++ with classes <> modules ; templated classes <> functors. What am I missing?
- choeger 9y agoAn interface and safety. In C++ the template's arguments are directly copied into it's body. A module can only be accessed via its interface. This in turn allows for a typechecker to guarantee the absence of certain errors. It also enables separate compilation that you do not have for C++ template's. The downside is of course that a C++ compiler can apply more optimizations. As a side note: the argument, I can do X with Y so why use Z is somewhat misleading, when Y and Z are both Turing complete ;).
- Drup 9y agoThere is no abstraction in C++, you can always poke around the memory layout of what you are given, and change your behavior depending on that. For example, let us say you have a module satisfying this signature: module M : sig type sorted val import : int array -> sorted end = struct type sorted = int array let import = Array.sort end In OCaml, if you have a value of type `sorted`, you know it's indeed sorted. In C++, as soon any code external to the module had an handle on it, you don't know! It could have modified the array behind your back, since it can look directly at the definition, or worse, poke in the memory layout.
- rapala 9y agoYes, you can access raw memory and flip bits to modify a private member of an object (and it migth even be defined, not too sure about that thou). But I don't find that an valid argument for the statement that you can't do abstractions in c++. That's just not something people do. The c++ version of your example would be, if I understood your code correctly, to take an std::vector as an constructor parameter, copy it to an private field and sort it.
- Drup 9y agoI don't know all that much about C++'s object system, so I can't give you a concrete example on how it breaks down. However, you say that it is "not something people do" ... well maybe not in C++ (I highly doubt that), but it's very common in many languages. In C, it's common to look inside structs directly and change things. Javascript libraries do it all the time: They inspect their arguments, look at the types and change their behavior depending on it. It's a common programming practice to poke deep into the data-structures and do things. In Java, they made it an art with reflection and monstrosity such as Spring. Abstraction is a bit like immutability: Sure, you can try to fake it in languages that don't have it, but then you are just praying that everyone plays by your rules. :)
- pjmlp 9y agoWhile ML has lots of cool features, that kind of abstraction capabilities and compilation speed is also possible in other languages with modules. Even if their modules aren't exactly ML-like, they offer similar features. For example CLU also explored hidden types and generic interface definitions. Modula-3 does offer opaque types, generic module interface definitions with multiple implementations and class inheritance. However it does lack closures/lambdas, forcing the use of function pointers, which aren't so developer friendly, Similarly the way traits work in Scala are quite close to ML modules. So there are ways of achieving ML module capabilities, even if not 1:1 as they are done in Standard ML.