5 ms·
It is not "parallel programming", which is hard. Concurrency and synchonization is.
by beza1e1 15y ago
It is not "parallel programming", which is hard. Concurrency and synchonization is.
- wladimir 15y agoWhat other side is there to "parallel programming"? How many cases of parallel programming are there, in which you need no concurrency and synchronization at all?
- yvdriess 15y agoDataflow architectures and languages for example. Or even vanilla SIMD instructions. The heart of the issue is that von-Neumann architectures are really not well suited to doing parallel programming: global PC in a single random access read/write memory. Any modification you make to that model to duplicate one module will introduce some heavy concurrency issues for you to deal with. For example multi-threading gives you multiple PC in the same memory space, leading to races, deadlocks, starvation etc. Compare this to simple SIMD. You do a parallel operation float4 + float4 without any need for concurrency or synchronization.
- wladimir 15y agoBut even in dataflow architectures (for example, the float4+float4 example) there are places where you want the different paths to meet. That's where synchronization (a barrier) is needed, as both results need to be available before the operation can be stared. Of course in the case of SIMD this nicely happens internally in the hardware so nothing can go wrong, but in more complicated cases, for example if you're programming CUDA you need to care about it sometimes. I agree that an alternative hardware architecture could probably solve this, but that is taking it a bit far and doesn't help solving any immediate problems.
- scott_s 15y agoEven if you have merged paths in dataflow and you don't need a barrier, it can be difficult to reason about. Specifically, it's difficult to figure out what a legal result even is.
- kd0amg 15y agoBut even in dataflow architectures (for example, the float4+float4 example) there are places where you want the different paths to meet. That's where synchronization (a barrier) is needed, as both results need to be available before the operation can be stared. A dataflow architecture should be doing this in hardware -- don't issue an instruction for execution until all of its operands have been reported. The point is that it's not something the programmer needs to be explicitly concerned about.
- Daniel_Newby 15y agoA dataflow architecture should be doing this in hardware -- don't issue an instruction for execution until all of its operands have been reported. There is still a need for application-level synchronization. For example, to keep the same money from being withdrawn from a bank account twice.
- wladimir 15y ago"should be", yes, let's move our problems to the hardware guys. I'm all for it. But hardware takes long to develop (if practical at all; a hw implementation might become to slow and expensive), and even longer to be mainstream, so I don't really see changing the hardware as a solution.
- scott_s 15y agoSomeone had to do the hard work of reasoning about concurrency and enforcing synchronization; it doesn't come for free. In the case of SIMD instructions, it was the processor designers. With well designed interfaces, parallel programming can be easier. But such interfaces abstract away the need to consider concurrency and synchronization - mostly. If you use the constructions outside of the bounds where safety is promised, then all bets are off. For example, parallelizing for loops with independent iterations with OpenMP is trivial, and you don't have to consider concurrency and synchronization. But once you provide non-independent loops, everything blows up and now those things are very important.
- yvdriess 15y agoAgreed, the cost is shifted to the back-end. But sometimes it is worth to pay that cost up-front, cfr the memory management debate. Synchronization and concurrency are much simpler in a system where you do have guarantees. No amount of interfaces or libraries will indeed make OpenMP in C safe, but no amount of hacks are going to make a fine-grained acyclic data flow graph deadlock or share state. The backend of the latter can pay the upfront cost of optimizing the shit away, for example no-copy optimizations, in a safe environment. One of the biggest research effort in dataflow at MIT came in the aftermath of Multics; the ambitious SMP time-sharing OS research project that later spawned UNIX. citing: http://en.wikipedia.org/wiki/Jack_Dennis http://en.wikipedia.org/wiki/Jack_Dennis
- scott_s 15y agoSynchronization and concurrency are much simpler in a system where you do have guarantees. I agree wholeheartedly, but there is a consequence that cannot be ignored: the resulting programming model is less expressive. The consequence of providing those guarantees is that there are something programmers just can't do. It's a trade-off, and I think we're still exploring how to provide a programming model that both abstracts away the complexity while still providing an expressive enough programming model to be useful in most circumstances.
- yvdriess 15y ago
- beza1e1 15y agoTake for example problems you can solve with map-reduce. Very easy to parallelize. From the application developers point of view there is no concurrency or synchronization necessary, because it was already solved in general by the framework. Unfortunately, there are lots of problems, where this is not possible. For example, for a distributed concensus problem the conflict resolution is application-specific.
- roel_v 15y agoIt is not "running fast" which is hard. Moving your legs up and down very fast is.
- maurycy 15y agoIt is a bad example. "Running fast" is a subset of "moving your legs up and down."