3 ms·
In a few use cases, such as big-data processing and things like that, a common pattern is to haul a tiny data-processing kernel to several machines to do an ope
by luizfelberti 6y ago
In a few use cases, such as big-data processing and things like that, a common pattern is to haul a tiny data-processing kernel to several machines to do an operation on some data they hold, rather than shuffle the data around.
This pattern can be seen in several places and frameworks like Spark & company, and it's also pretty much the foundational idea behind processing data with Joyent's Manta[0]. It always pretty much boils down to "serialize a function, and send it somewhere, to run on the data locally".
Some of these frameworks do it only through native functions (Spark runs serialized JVM procedures, and for Python functions I believe it shells out to an actualy Python interpretes), others are completely agnostic to this and work with binaries (Manta).
It'd be pretty hard to do, and have a lot of overhead, if all of those "functions" that I'm shuffling around are, at the lower bound, 130MB in size.
I understand that a lot of this can be done directly at the Julia layer for Julia things (much like Spark does with JVM things), but it kinda limits the applications of this nonetheless.
[0] https://github.com/joyent/manta https://github.com/joyent/manta
- ViralBShah 6y agoJulia can already do this in its `remotecall` API (actually has done this since v0.1), which serializes a closure and sends it to a remote node for computing. https://docs.julialang.org/en/v1/manual/distributed-computing/ https://docs.julialang.org/en/v1/manual/distributed-computin... However, this is not related to the system image (sys.so) that is used to start the Julia process, which has a compiled version of base, stdlibs and whatever else you choose to build in.