3 ms·
Haxl: Making Concurrency Unreasonably Easy [video]
- aban 9y agoGreat talk. For more on concurrency and parallelism in Haskell, check out Parallel and Concurrent Programming in Haskell [0], deemed as the best book on the subject, also written by Simon Marlow. [0]: http://chimera.labs.oreilly.com/books/1230000000929 http://chimera.labs.oreilly.com/books/1230000000929
- jankotek 9y agoI have not fully digested yet, but seems very similar to Scala Parallel Collections and Java8 Streams. There are databases which implements such interfaces.
- willtim 9y agoIt isn't the slightest bit similar to those! Haxl is a high-level library for specifying (and optimising) concurrent data retrieval.
- jack9 9y agoI'm not sure that's an apt description either. Haxl's features include auto batching of IO through parallelism and a temporal cache...for error logging when failures occur. It's spelled out in the first minute he talks about it. In the IO batching code (the boilerplate) you specify how and what to output in your logging. Haxl can be used without performing optimization (write an inefficient batcher) or doing data retrieval (the I/O is fire and forget and never cache anything). Maybe I am misunderstanding the basic usage.
- willtim 9y agoIt's not easy summarising it in one sentence, but 'auto batching' and 'caching' is what I meant by optimising the data retrieval. EDIT: the Haxl docs describe it as "a library and EDSL for efficient scheduling of concurrent data accesses with a concise applicative API"
- solidsnack9000 9y agoWhich databases are those and what are the relevant features? Are you talking about FILESTREAM? https://technet.microsoft.com/en-us/library/bb933993(v=sql.105).aspx https://technet.microsoft.com/en-us/library/bb933993(v=sql.1...
- jankotek 9y agoHazelcast Jet, Spark, my project... they work similar way. But as I wrote, I need to read the presentation first.
- solidsnack9000 9y agoJet and Spark are not what most people would call databases, since they have such a heavy stream processing focus and are rarely stores of record.
- daxfohl 9y agoI wonder how long it will be before compilers/interpreters of async-aware languages just do this by default. CPUs and low-level language compilers jump through all kinds of hoops of out-of-order execution, branch prediction, caching, parallel execution, etc. I picture a day maybe 10 years from now where developers in most languages don't even have to think about these things. All the old-timers will still be structuring their code "as though it didn't exist" whereas the new kids will fly along without even thinking about it. Kind of like garbage collection the first few years.
- eru 9y agoEven in Haskell you still have to write some boilerplate to get this by default. Languages that don't separate pure computation from IO (and other side-effects) make it even harder on the compiler. So to answer your question: implementations will do this by default only after pure/constant will become the default for all functions/variables, with side-effecting/mutable clearly marked.
- hughperkins 9y agoIf one looks at it from a 'what a today's compiler can do', then sure, one needs to statically declare everything. If one looks at things from a 'any technical task can be handled by machine learning, sooner or later' point of view, there seems to be no obvious reason why parallelization, which is a purely technical task, not like say, writing music, could not sooner or later be handled by machine learning algorithms?
- eru 9y agoWriting music is a pretty technical task, actually. Especially if you optimize for something as tangible as: please human ears, or "fit within the actions in the movie / computer game". The idea here is that we want to find a way to make automatic parallelization of computer programmes happen as much as possible long before we've solved general AI. So even in a language like Haskell you still have to pay some attention to make your code amenable to parallelism. See Guy Steele's advice (see https://vimeo.com/6624203 https://vimeo.com/6624203).
- deleted 9y ago[deleted]
- sedachv 9y agoI looked through the slides but not the video and the slides ignore the hard problem: how do you schedule these requests? How do you know how many parallel requests you can issue without hammering the database or service? How do you batch queries so that you get acceptable latency and a query size that will not choke the database? The last question is probably easy for most use cases where you have independent requests coming in (typical web application) - in the context of a single request you can usually get away with batching as much as is possible. But the scheduling problem is very similar to the promises of "free parallelism because Church-Rosser" - actually taking advantage is an open problem. Even when you know how much time each job takes in advance, multiprocessor scheduling is NP-hard. Anyway, if someone watched the video and the question is addressed there, please let me know so I can watch it.
- eru 9y agoI think the idea is that with the Haxl approach the scheduling can be dealt with independently from your business logic. By the way, I don't see how Church-Rosser would give you any free parallelism---even in theory. You'd still have to heed Guy Steele's advice (see https://vimeo.com/6624203 https://vimeo.com/6624203).
- sedachv 9y ago> By the way, I don't see how Church-Rosser would give you any free parallelism---even in theory. Church-Rosser theorem means that all possible reduction sequences lead to the same normal form, so you can β reduce the subterms in parallel. If you look at Paul Hudak's and SPJ's publications from the late 1980s and early 1990s many of them are actually about trying to exploit this implicit parallelism: http://sunsite.informatik.rwth-aachen.de/dblp/db/indices/a-tree/j/Jones:Simon_L=_Peyton.html http://sunsite.informatik.rwth-aachen.de/dblp/db/indices/a-t... http://dblp.uni-trier.de/pers/hd/h/Hudak:Paul http://dblp.uni-trier.de/pers/hd/h/Hudak:Paul https://pdfs.semanticscholar.org/8912/5be7f9c222793c18b99d0644c16ff19d9f06.pdf https://pdfs.semanticscholar.org/8912/5be7f9c222793c18b99d06... This turns out to be a hard problem and fell out of fashion as a research topic. The big problem is managing overhead. There has been some recent work to try to incorporate automatic profiling feedback to adjust the parallelism granularity: http://dominic-mulligan.co.uk/wp-content/uploads/2015/05/S-REPLS_Calderon.pdf http://dominic-mulligan.co.uk/wp-content/uploads/2015/05/S-R...
- pmarreck 9y agoHow does this differ from BEAM langs which already make concurrency "unreasonably easy"?
- kornish 9y agoFor one, they serve different purposes. Haxl is specifically for concurrent data retrieval while BEAM is a general purpose platform for fault-tolerant computation. For another, they operate via different interfaces. BEAM languages communicate concurrently only via a message passing interface. Haxl lets the author write declarative code specifying what to retrieve, then the library executes it concurrently and in parallel under-the-hood. Both are properly described as "unreasonably easy" - just for different sorts of things, on different platforms, with different interfaces.
- vdijkbas 9y agoHaxl is a powerful abstraction with IMHO a beatifuly simple implementation. However for our use case at LumiGuide (reading and writing registers of modbus devices) it wasn't simple enough. We just needed an abstraction for batching and did not need caching and the other features Haxl provides. So I wrote monad-batcher which as the name implies only provides a batching abstraction (which can also be used to execute commands concurrently). All the other features can be build on top of monad-batcher as separate layers (separation of concerns). The library is available on Hackage but needs a bit more documentation (a tutorial would be nice): http://hackage.haskell.org/package/monad-batcher http://hackage.haskell.org/package/monad-batcher