7 ms·
No, a supercomputer won’t make your code run faster (2017)
- pinewurst 5y agoThis is so true - so many bad Perl, Python and R jobs “requiring” supercomputer resources that a decent programmer would allow to run at better speed on a single modern server!
- dhosek 5y agoThere was a post on HN back in 2016 pointing out that some tasks in Hadoop could be executed more quickly by just using standard Unix tools and pipes. It annoys me to no end that very few languages provide out of the box mechanisms to be able to emulate that sort of simple parallelism. Although I did find a Google proposal for the C++ standard library that, at the very least, made it clear how to roll this functionality myself should I so choose.
- lmm 5y agoEverything's quicker if it doesn't have to work. Hadoop doesn't just run the code, it gives you job management, log collection, error reporting...
- dhosek 5y agoYes, but it was more than that in the difference. It's that for many (maybe most) tasks, a simple pipeline-style parallelization makes much more sense than going through all the pain of map-reduce. If you've got five steps and five threads and a big pile of data, it can make a lot more sense to try to have each step running in its own thread and passing a processed chunk of data to the next step rather than having five threads trying to each take some chunk of data and process it through the five steps itself. Factories have done this since the 19th century, but most programming languages make it hard.
- lmm 5y ago> If you've got five steps and five threads and a big pile of data, it can make a lot more sense to try to have each step running in its own thread and passing a processed chunk of data to the next step rather than having five threads trying to each take some chunk of data and process it through the five steps itself. Factories have done this since the 19th century, but most programming languages make it hard. You can do that with an actor model by giving each actor thread affinity, but it's rarely what you want since it's normally a lot more efficient to move the code to the data than move the data to the code. In a factory the machines are heavier than the products they're working on so it makes sense to move the machines to the materials; for most code tasks it's much more like harvesting a field or something where it's much more efficient to move the code to the data.
- wging 5y ago"Scalability! But at what COST?" is another classic in the category (though it's not precisely about on-box parallelism if I recall; the authors' acronym "COST" is the "Configuration that Outperforms a Single Thread"). https://www.usenix.org/system/files/conference/hotos15/hotos15-paper-mcsherry.pdf https://www.usenix.org/system/files/conference/hotos15/hotos... http://www.frankmcsherry.org/graph/scalability/cost/2015/01/15/COST.html http://www.frankmcsherry.org/graph/scalability/cost/2015/01/...
- azalemeth 5y agoPart of the reason behind this is that it's really, really hard to hire good programmers in academia -- salaries are difficult to put into grants and equipment is easier. It may well be the case that an übernerd could get your algo a significant runtime improvement, but sometimes you're massively unlikely to actually be able to get an übernerd. On the other hand, saying "I and doing big science, give me time on this cluster" is often more of a positive thing to write.
- ta988 5y agoAnd at the same time we give indecent salaries to PIs, Deans and basketball coaches. We waste large amount of money on inefficient processes and we take 5 years to develop projects that would have taken 6 months with a good dev and in the end would have cost less in maintenance and improvement. I needed to hire devs for a project, it was impossible to do.
- pinewurst 5y agoWhat I've seen is that low-salaried ops people are viewed as a substitute for capital expenditure in this environment. It strongly discourages purchasing software and even non-white box hardware. "Just have IT cobble together something open source", is the message from administration and PIs. The result is ramshackle, inefficient systems running on low bidder hardware from often dubious "HPC" resellers.
- chunkyks 5y agoI, too, am a software engineer dealing with researcher code on a daily basis, and I'm also the person that oversees most of our research compute hardware [sysadmin in a past life]. Some of my other favorite gems: * "I need a terabyte of RAM" * "50k samples ran in the blink of an eye. 100k ran in an hour. 1 million took a week. I'm about to do the full 50 million observations, I expect it'll take ten days" * "This code is always slow... that's a lot of math it does" * "Profiler? Never heard of such a thing" * "Big-O? What's that?" * "I only have ten days' money, and it took me that long. Afraid I can't afford your time to dig into this" * "AWS will fix this" Generally I'm dealing with the intersection of two or more of these. I do think, though, that the last paragraph is a bit unfair. Literally everyone I work with is incredibly smart; just their background isn't the same as mine. Things that I find incredibly tough are idiotic to them ["Can't I just take the average and be done with it?" "no"]. So, for most of the people I work with, it's not "common sense"; they literally lack the frame of reference to even know what's awry. I never hold it against anyone, and nor should you. My favorite example is from many years ago, a guy who had some code that took a month, "because that's just complicated math, it's how long that takes". I ran it through a profiler, and after about ten minutes work, it ran in seconds. I came out looking like a hero.
- arthurcolle 5y agoTerabyte of RAM is my favorite because I choose to believe he had a perfect data model and was literally telling you exactly what he needed :D In my recent work doing financial modeling (options, cryptos, other nonsense), I pretty much just try to cache absolutely everything with some arbitrary 1 week TTL and then continue to cache new data every day. Cache hits cause a 10% bump in the current TTL time delta. Kinda heuristic but whatever, works in most cases. Any other funny stories? I always love these engineering horror stories. Have you looked at Julia at all? Would be interesting to hear if there's been an uptick in usage from the research community from a veteran research code surgeon!
- chunkyks 5y agoThe mathematician that designed the model I referred to at the end ended up passing it on to me. I've now built something of an empire of people and projects using it, "just" by making better tools to manage its inputs, outputs, visualisation and automation. I was stupidly excited to recently do more work on that same model - I took one of its internal datastructures and turned it from a simple queue into a priority queue using a couple different queueing metrics. In both cases, I had run cases out to a scale where I was able to actually show going from O(1) to O(log(n)) performance - but got beautiful wins because the fastest work is work you never have to do. Not often I can empirically show traditional CS101 stuff in a real-world chart.
- maddyboo 5y agoYou know what would be fun? A website like Project Euler, but you're given programs that are less than optimally efficient and your job is to make them as fast as possible.
- okrad 5y agoTook a secure programming course in college that was basically just that for each assignment except the provided code wasn’t just inefficient but broken in interesting ways that prevented you from googling an error message. Incredibly fun course!
- optimalsolver 5y agoThis CS guy looked at some scientist's code and sped it up 14,000x: http://james.hiebert.name/blog/work/2015/09/14/CS-FTW.html http://james.hiebert.name/blog/work/2015/09/14/CS-FTW.html
- alex_smart 5y ago0.1 seconds is still too many. There is probably another 1000x speed up possible there. We just have to do a merge on two rather small arrays. The most expensive operation here would actually be writing out to the output array.
- France_is_bacon 5y agoI remember one time, there was code that took 3 hours to run. Then printing the output took another 3 hours. A huge problem was that it printed out a stack of output about 3 feet high, and if the printout jammed, as was often the case, you had to start all over, even if it was 10 pages short of completion of the print job. Another 6 hours. Another printing error, yet another 6 hours. So, as you probably can see, it could take 18 hours or more to get a correct printout. I took one look at the code, and changed it. It then took about literally 5 seconds to run the code. Then I added some code that would allow the entry of a page number, and the users could enter that page number and print from that page. So, it would then take always a maximum of 3 hours to print out the output. Then I told the night operator to run this program for the operations department at midnight. So, as far as the operations department was concerned, that brought the time down to zero, and the printouts were on everyone's desk when they walked in the door in the morning. The manager of that specific department literally came up to me after it was rolled out and was crying hard, and hugged me for about a minute, crying the whole time, as her job was on the line, no fault of her own, because she couldn't get the job done, because she couldn't get the printouts. There was no way that the most powerful computer in the world that would have got that code to run any faster, because of the horrible way it was written. Going from maybe 18 hours to zero time, on the same exact computer....what is that improvement in percentage anyways? What's 18 hours/0 hours? haha...
- tomjen3 5y agoInteresting. My first thought was to print to a PDF, then print the PDF, since you can restart it from whatever page you want, and you don't have to restart it. You solve the actual problem, but this would have been quicker.
- France_is_bacon 5y agoWell, it was more complicated than that, and situational, as things usually are.