3 ms·
We're using Haskell to produce an entire compute platform for building, composing, and running type-safe containerised micro-services for data science. This cer
by mands 9y ago
We're using Haskell to produce an entire compute platform for building, composing, and running type-safe containerised micro-services for data science. This certainly touches on a lot of the areas that Haskell is traditionally known for being good at, e.g. building DSLs, type systems, interpreters, etc.
However, the work also includes a runtime platform that is more low-level, including building our own container system, talking to the Linux kernel, using cgroups/resource management, and distributed message passing - areas where languages such as Go have found a niche, and may be classified as Real World.
However for us, Haskell's type-safety, runtime performance, and extensive ecosystem has been a boon even in this domain. We effectively use it as an (incredibly powerful) general-purpose language here, and it's worked more than fine.
We're currently at around 15,000 lines of code with a team of 5 Haskellers, and it hasn't really been a problem regarding performance, understanding the codebase, or with newcomers to the team.
(plug - we're at https://www.github.com/nstack/nstack https://www.github.com/nstack/nstack and are hiring more Haskellers)
- nickpsecurity 9y agoThis work might give you some ideas even though it's a bit dated: http://programatica.cs.pdx.edu/House/ http://programatica.cs.pdx.edu/House/ Also, COGENT for lowest-level stuff being wrapped for use in Haskell somehow might be interesting. Used on ext2 filesystem already. https://ts.data61.csiro.au/projects/TS/cogent.pml https://ts.data61.csiro.au/projects/TS/cogent.pml
- mands 9y agoThanks for the pointers - house is great, I remember reading some of the papers a long time ago. I haven't seen COGENT before - will take a look over the weekend - thanks!
- seabrookmx 9y agoAny particular reason you're building your own container system instead of leveraging LXC or Docker? For the massively parallel workloads you find in data science, it seems like you'd benefit a lot from the wealth of container orchestration tools around Docker (swarm, Rancher/Cattle, Kubernetes) in order to easily scale out your functions. Especially when many companies already have these set up for their more vanilla applications. This is an example I've seen that can leverage a docker swarm for invoking functions, loosely modeled after AWS Lambda: https://github.com/alexellis/faas https://github.com/alexellis/faas
- mands 9y agoHi there - we've actually built a lot of our container ecosystem around existing Linux tools, including `systemd-nspawn`, `btrfs` and more rather than creating the whole stack from scratch - and again this is all controlled from Haskell. We experimented with Docker, Kubernetes and more, but found they they made lots of assumptions about what was running inside a container that didn't mesh with our compute model, so using lower-level primitives worked better for us. We're really lucky also to have one of the main `rkt` developers joining us soon to work on the container side.