4 ms·
Since it's a really big, technical book here are some quick quotes I thought were useful in general: "This book is a handbook of widely applicable and heavily
by aout 12y ago
Since it's a really big, technical book here are some quick quotes I thought were useful in general:
"This book is a handbook of widely applicable and heavily
used design techniques, rather than a collection of
optimal algorithms with tiny areas of applicability."
"As a rough rule of thumb, use the simplest tool that will
get the job done. If you can, simply program sequentially.
If that is insufficient, try using a shell script to mediate parallelism."
"[...] parallelism is but one potential
optimization of many. A successful design needs to
focus on the most important optimization. Much though I
might wish to claim otherwise, that optimization might or
might not be parallelism."
It's a very good handbook for advanced developers who use native languages. It basically adresses any problem you might encounter when dealing with sync, locks etc...
Near page 300 the author writes about higher level languages and clearly states the shortcomings of these new types of programming. Not sure what to think about it. Are we really discussing the usage of high level "scripts" to handle heavy computations in the 10 next years? I'd love to hear a math or physics Ph.D about this.
EDIT: I mean, Wolfram is kind of a "ultra high level scripted platform" and seems to work pretty well. Isn't that the future?
- rprospero 12y agoYou have summoned a physics Ph.D. I do quite a bit of parallel numeric work, but I'm usually trying to take a calculation measured in hours and change it to one measured in hours, so I'll defer to my colleagues with the multi-year computations when they inevitably arrive to correct my misunderstandings. While I haven't used the Wolfram language specifically, I did quite a bit of my thesis in Mathematica 8 and became quite familiar with its abilities and limitations. On the one hand, the high level scripted platform did do an amazing job of parallelizing certain classes of computation with no outside effort. However, when you left those classes of problem, it wasn't really any help. In my own work, I had a particularly nasty integral that I needed evaluated numerically. The integral itself only took about an hour to run, but I'd need to run it with multiple sets of parameters that could easily eat up a week of computer time. Now, since I had multiple parameters, Mathematica would intelligently run multiple integrals at the same time, so I could have eight calculations done every hours, instead of one, and it would take a day to run, instead of a week. However, I ran into a situation where the parameter for each integral would depend on the results of the previous one, so running them in parallel was no longer an option. I had to go down to a lower level and speed up the integral itself. On one level, this should have been trivial. It was a simple, one dimensional integration, evaluated via Monte-Carlo. Let each worker thread pick it's own integration points and reduce the results at the end. It's trivially parallelizable. Except Mathematica wouldn't do it. It was brutally resistant to any attempt at speeding up the integration. I eventually wrote my own parallel Monte-Carlo integrator in another language and used that, because the high level scripted platform was too high level for what I wanted to do. My colleagues ran into similar issues. They often had large, time consuming calculations that had a recursion somewhere in the process that caused Mathematica to jump back to serial processing. Sometimes I could find a different representation of the problem that the system could skip. More often, I didn't have the time or domain expertise to force Mathematica to see that the task could be run in parallel. So, on the question of whether an "ultra high level scripted platform" is the future, my answer is "the future of what?" You can only get to high level by choosing a domain to model and the higher your level, the more specific the domain. Now, to prevent misunderstanding, I should mention that these are some massive domains that could benefit greatly from a high level, parallel model. I cannot count the number of times I asked "How do I parallelize this trivially parallelizable program" and not been able to find any answer more specific that "trivially". Yet, even with the advances, the moment that you leave that problem domain, you're still going to be going back to lower level techniques. To steal from the author, we can get big wins by shaving the Mandelbrot set, but there will still be people who need those little side valleys.
- kd0amg 12y agoNow, to prevent misunderstanding, I should mention that these are some massive domains that could benefit greatly from a high level, parallel model. I cannot count the number of times I asked "How do I parallelize this trivially parallelizable program" and not been able to find any answer more specific that "trivially". This is a large part of my motivation for working on data-parallel languages -- the parallel programming tools we have are generally good at making easy problems hard. I agree that the domain of highly regular data-parallel computation is large (given the attention going to GPGPU), enough so to warrant its own high-level language. APL already has a convenient and flexible (user-level) programming model for this but carries along enough ad hoc limitations and weird corner cases to make a good parallelizing compiler infeasible. Maybe someone will eventually come up with a general-purpose parallel language that offers a smooth transition from the easy parallelism to the hard (concurrent) parallelism, but so far everything I've seen that accommodates concurrency complicates the non-concurrent cases in doing so.
- gh02t 12y agoWell, scripting languages are actually pretty popular even now. I work in uncertainty quantification for nuclear engineering. One common task is to run the same simulation many times with slightly different parameters. Scripting languages like Python are popular for this and they give you easy parallelism. The current pattern in research is usually to hammer out a prototype in a scripting language, profile, then migrate hot code paths to Fortran or C++. That's only if necessary though, you can get remarkable performance most of the time just using NumPy and similar. I'm a long time user/lover of Mathematica and now what I guess is the Wolfram language, definitely an expert on it. There's one thing I know about it, which is that it isn't the future, at least not in my industry. Open source is surprisingly popular in the NE research community and as long as WL stays closed, it won't be significant.
- acaloiar 12y agoBioinformatics largely takes a similar approach. Because the field is rooted in Statistics, statistical theories are often prototyped in GNU R by Ph.Ds without formal computer science backgrounds. While statistically sound, R implementations often lack the performance properties necessary to crunch large datasets. This is where those of us with backgrounds in computation will optimize and parallelize code using lower level languages. There is no shame in starting from a scripting language. I often prototype ideas using R, Perl or Python before finalizing them in C, Java, OpenCL, etc.
- deleted 12y ago[deleted]