5 ms·
Something the site didn't answer, what high-level languages are suitable for writing highly-parallel network code? I need to write some code that does stuff lik
by adbge 15y ago
Something the site didn't answer, what high-level languages are suitable for writing highly-parallel network code? I need to write some code that does stuff like: given a couple million URLs, download each, extract an element, gather all extracted elements in a list.
- sinope 15y agoErlang
- gtrak 15y agosounds like a job for clojure, lazy seqs and java's threading and network libraries make this sort of thing not too terrible
- DanWaterworth 15y ago`Highly-parallel network code` was the problem erlang was designed for. Haskell would also be a good fit.
- Eventh 15y ago"Erlang was designed to program fault-tolerant systems" according to Joe Armstrong [1]. [1] http://www.infoq.com/presentations/Systems-that-Never-Stop-Joe-Armstrong http://www.infoq.com/presentations/Systems-that-Never-Stop-J...
- DanWaterworth 15y agoHow do you make a system fault-tolerant? Make it distributed. How do distributed systems work? By communicating via a network.
- jemeshsu 15y agoI suppose if you use zeromq, any high level language that has zeromq bindings will be able to do the job. That is over 30+ languages including Python, Ruby, Lua, Perl, Node.js, Objective-c, Java, Scala etc etc
- exDM69 15y agoThe choice of language is not so important here, but the choice of a language implementation is. If you're doing a lot of network I/O in parallel, you need a language and a framework that can take advantage of hardware concurrency (multiple threads and/or processes), but should be able to do fast switches between your tasks (not requiring a full OS supported context switch, which is slow). This pretty much rules out e.g. Python and Ruby, because the CPython/CRuby interpreters suck at concurrency (there are other implementations available, which are a little better). You can still do many Python processes, which do lightweight switching between I/O tasks somehow (e.g. with Twisted), but then load balancing between processes becomes a PITA. Node.js is similar, there's only one thread running at once. With Node.js (and twisted) you also have to manually "compile" your code for asynchronous style continuation passing style, something humans suck at but compilers do very well. Now given the task you said, you could take C and epoll/kqueue and write a small framework yourself, but processing the data you got with C might not be too nice. Or you could use an interpreter or a compiler that does this automatically for you. This is where the choice of language kicks in. In order for the language implementation to be able to effectively organize parallel execution, it needs some information from the language. C doesn't have any helpful information, which is why it often relies on full OS context switches where all machine state is stored and restored (or a non-portable hack with the stack). Something lisp-y with first class continuations might be helpful, but many lisp implementations don't really do concurrent execution well. So, for a language that concurrently executes network I/O efficiently while still being high level and fast, my recommendation for your task would be: Haskell. There's years of research and engineering work done on just what you're looking for in Haskell. Just forkIO as much as you wish, write your code in regular sequential imperative style, compile for multithreaded execution (ghci --threaded) and let the compiler and runtime do all the hard work for you. Erlang might do the job too. If anyone knows of other languages with smart I/O multiplexing and co-operative userland threads or fibers (with concurrent execution!), please reply here! NOTE: my assumption was that you're actually processing the data you're receiving and you're more or less CPU limited. If you're I/O bound, you don't necessarily need concurrent execution and may get away with Twisted or Ruby fibers or Node.js (without using multiple processes and load balancing or task queues).