4 ms·
Disclaimer: I don't speak for OCaml or OCaml Labs but I am a contributor to multicore. It is indeed that it is hard to retrofit parallelism on to an existing l
by sadiq 7y ago
Disclaimer: I don't speak for OCaml or OCaml Labs but I am a contributor to multicore.
It is indeed that it is hard to retrofit parallelism on to an existing language while trying to retain backwards compatibility _and_ performance.
Backwards compatibility is tricky because there's lots of C code using the C API.
Performance is hard because OCaml users are used to well performing code, with low and predictable pause times from the GC (<10ms).
The community is small and it seems like there wasn't appetite for maintaining two distinct runtimes with very different performance characteristics.
The current implementation for multicore GC (https://github.com/ocaml-multicore/ocaml-multicore https://github.com/ocaml-multicore/ocaml-multicore) is reasonably close to upstream in performance on single threaded code and yet will scale up to multiple threads. It requires a change to the C API though.
There's a modified multicore GC (https://github.com/ctk21/ocaml-multicore/tree/stw_minor_gc https://github.com/ctk21/ocaml-multicore/tree/stw_minor_gc) that doesn't require the C API change and we're currently writing up a paper that contrasts the two wit a fairly substantial amount of benchmarking (https://github.com/ocamllabs/sandmark https://github.com/ocamllabs/sandmark).
Happy to answer any questions.
- throwaway894345 7y ago> Performance is hard because OCaml users are used to well performing code, with low and predictable pause times from the GC (<10ms). I would expect that low pause times are easy enough (or relatively easy considering the difficult domain of GC optimization), but keeping the allocations cheap is (one of) the key constraints, no?
- sadiq 7y agoYes, it's very easy to make pause times low if you're willing to make allocation expensive - it's all trade-offs. In multicore's case it's about keeping low pause times while also keeping allocation cheap _and_ maintaining throughput.
- chrisseaton 7y agoWhy would allocation become more expensive in parallel? Surely you’re allocating in thread-local space? It’s like two machine instructions. Where does the extra overhead come from?
- throwaway894345 7y agoIf you want conventional shared memory parallelism, then your allocations can't assume thread-locality.
- chrisseaton 7y ago> If you want conventional shared memory parallelism, then your allocations can't assume thread-locality. I don't really understand why not. Can you expand on it? Doesn't the JVM for example do thread-local allocations in a conventional shared-memory parallel environment?
- throwaway894345 7y agoI was mistaken; wasn't thinking clearly as I responded. JVM can get away with this because it's generational/moving; bump allocator works in the young generation and then objects are subsequently moved. Generational/moving is tricky, especially with C APIs (if the GC moves an object that C has a pointer to, the C code dereferencing the pointer will find potentially garbage data) and I believe they make it difficult to get good STW times, but at this point I'm pretty well out of my depth.
- sadiq 7y agoBoth current multicore GCs do exactly that to keep allocation cheap but there are different design choices within the constraints we could have gone with. To avoid frequent synchronisation we could have gone with a (potentially generational) non-moving collector for the minor GC which would still have preserved the C API, probably allowed for very low pauses but would have made allocation more expensive.
- pdimitar 7y agoIn what form will Multicore OCaml support the, you know, multicore stuff? Will it be native OS threads? CSP? Actors? Something else? I am very excited to use OCaml with multicore abilities but lately I realised that I have no clue how will the initial support even look like in terms of a programming API.
- Ono-Sendai 7y agoWhy don't you use reference counting GC?
- the_why_of_y 7y agoBecause atomic reference counting (needed for multi-threaded runtime) is well known to have terrible performance.