7 ms·
Haskell in the Datacentre
- lmm 10y agoI think the bigger question is why they're using a C++ Thrift server. Haskell is good at transforming data - they should be able to use an OS socket and do the message decoding/encoding in pure Haskell.
- paulajohnson 10y agoAt a guess, the server probably provides a unified front-end to services written in multiple languages.
- newmana 10y agoThis was covered in a previous post: "Haskell is sandwiched between two layers of C++ in Sigma. At the top, we use the C++ thrift server. In principle, Haskell can act as a thrift server, but the C++ thrift server is more mature and performant. It also supports more features. Furthermore, it can work seamlessly with the Haskell layers below because we can call into Haskell from C++. For these reasons, it made sense to use C++ for the server layer." https://code.facebook.com/posts/745068642270222/fighting-spam-with-haskell/ https://code.facebook.com/posts/745068642270222/fighting-spa...
- tome 10y agohttps://www.reddit.com/r/haskell/comments/5i9qmc/haskell_in_the_datacentre_simon_marlow/db6ifv2/ https://www.reddit.com/r/haskell/comments/5i9qmc/haskell_in_...
- chongli 10y agothe bigger question is why they're using a C++ Thrift server Because it's already there and it already works just fine. It's so frustrating to see this sort of comment on nearly every post about a language outside the mainstream. It's a religious argument and not a productive one.
- lmm 10y ago> Because it's already there and it already works just fine. Well clearly it didn't work just fine if its threading model was interacting poorly with that of their Haskell code. If there's a mismatch between those two components then that will continue to bite them in the future, and it's legitimate to ask whether moving the boundary would be a better way to solve the problem.
- easytiger 10y agoyour right, makes more sense to have both in c++
- massudaw 10y agoThat was the last tried solution, didn't work out well. And what is holding is nothing haskell specific. Is just wasting effort re implementing good broad used code at Facebook.
- mattnewton 10y agoThere is an old saying: "We did what we did, not because it was easy; But because we thought it was going to be easy" When they realized the bug the path forward probably seemed like patching and not a Haskell re-write, and that's a reasonable mistake to make.
- thinkpad20 10y agoI think the "why not write it all in Haskell" argument is a bit spurious. Only the engineers on the team knew best the advantages and disadvantages of the various approaches. Presumably there were very good reasons to integrate it the way that they did, whether driven by architecture, performance, knowledge, library, smallest-effective-change, or other reasons. It's clearly worked out very well in the most basic sense (it does the job). And on top of that, we now have bug fixes and other improvements submitted to ghc. Seems like a win-win.
- hellofunk 10y agoDon't confuse the features of a language with its actual run-time productivity.
- zhte415 10y agoSurely a bigger question is how a site with a basis in PHP transformed to making PHP compilers, .js frameworks, Haskell, C++ etc. They seemed to have succeeded in creating a culture of coders/hackers that were opinionated/empowered sufficiently to use their tool of choice, and to persuade others of the advantages. This doesn't really happen in large-cap companies (though does happen to empowered teams, but even then, only somewhat - thinking finance). And speaks to advantages.
- user5994461 10y agoIt's a big company. Thinkg just go to all directions when you've got thousands of devs.
- fusiongyro 10y agoIf you hire developers at scale, you probably get quite a bit of variety.
- deleted 10y ago[deleted]
- bojo 10y agoHappy to see they pushed the changes to the upstream GHC!
- 4ad 10y ago> GHC’s runtime is based around an M:N threading model which is designed to map a large number (M) of lightweight Haskell threads onto a small number (N) of heavyweight OS threads. [...] To cut to the chase, we ended up increasing N to be the same as M (or close to it), and this bought us an extra 10-20% throughput per machine. Ah, yes. As a Go developer, really wish Go moved to an 1:1 threading model.
- crawshaw 10y ago1:1 in Go would mean moving servers from blocking-style code to callbacks/futures. (I believe the story is slightly different in Haskell.) Nice sequential blocking-style code is my favorite thing about Go. I realize it's not C, and it makes runtime work much harder, but the payoff up in complex servers is completely worth it.
- 4ad 10y ago> 1:1 in Go would mean moving servers from blocking-style code to callbacks/futures. No, it would not mean such a thing at all. Where did you get this idea? In fact, on some operating systems/linker combinations, gccgo uses a 1:1 threading model. > Nice sequential blocking-style code is my favorite thing about Go. Mine too.
- crawshaw 10y agoIt would in practice. A typical server under load has more outstanding requests to answer than it can OS threads. It is not uncommon in a Go server at Google to find a million or so goroutines. If you want to have an OS thread that can also run C code, it needs a large enough stack to run C code. A million such stacks is not practical. On an OS that lets you avoid creating stacks with threads like Linux, and which has quite light-weight kernel-side accounting for threads, it would be possible to give each goroutine a thread. You would still need an N:M pool for cgo. As an added problem, 1:1 is slower for certain kinds of common channel operations. Anywhere you have an unbuffered channel and block waiting for a value, a 1:1 model requires a context switch to pass the value, whereas M:N means the userland scheduler can switch goroutines on the OS thread. It is precisely these benchmarks that led Ian to implement M:N in gccgo. If there are combinations where it is 1:1, that is either a new decision he has made in the last year, or (more likely) OS/ARCH combinations that he hasn't moved to M:N yet. I have seen a similar attempt at 1:1 lightweight tasks in C++ that ran into this. Without the ability to preemptively move tasks between OS threads, it ran into performance problems. Programs that needed the speed in that model had to switch to a futures-style of programming.
- greenspot 10y ago> At Facebook we run Haskell on thousands of servers Wow, I didn't know this. Anyone knows for which programs they use Haskell? The article doesn't say anything.
- kornakiewicz 10y agohttps://code.facebook.com/posts/745068642270222/fighting-spam-with-haskell/ https://code.facebook.com/posts/745068642270222/fighting-spa...
- DigitalJack 10y agoits for spam detection. Everything that gets posted on Facebook is processed by this code. There is a talk on it from a year or two ago at a functional programming conference. I'll see if I can find it and edit this post.
- noir_lord 10y agoI remember this one https://m.youtube.com/watch?v=UMbc6iyH-xQ https://m.youtube.com/watch?v=UMbc6iyH-xQ is that it?
- DigitalJack 10y agoNo, but I apparently didn't bookmark it and I'm having trouble finding it again. This might be it, don't have time to verify right now: http://cufp.org/2015/fighting-spam-with-haskell-at-facebook.html http://cufp.org/2015/fighting-spam-with-haskell-at-facebook....
- thedufer 10y agoSimon Marlow gave a talk at icfp in 2014 or 2015 that might be what you're referring to
- thedufer 10y agoSorry, couldn't look it up very well from my phone. I'm fairly certain you're referring to Simon Marlow's "Fighting Spam with Haskell" from ICFP 2015. Video: https://www.youtube.com/watch?v=IOEGfnjh92o https://www.youtube.com/watch?v=IOEGfnjh92o Abstract: http://cufp.org/2015/fighting-spam-with-haskell-at-facebook.html http://cufp.org/2015/fighting-spam-with-haskell-at-facebook....
- rvdm 10y agoReally great to read more about real world large scale Haskell use. I'd be curious to find out if the success FB has had with Haskell for spam filtering means they might consider Haskell for different parts of their stack too? Does anyone have any insight on this?