4 ms·
I wonder what portion of time on those supercomputers is used in optimizing the code itself; and how much could be done. With those massive clusters you could
by darkmighty 8y ago
I wonder what portion of time on those supercomputers is used in optimizing the code itself; and how much could be done.
With those massive clusters you could afford testing in parallel trillions of mutations to your simulation kernels and prove the correctness of the fastest ones -- pruning should be extremely fast by finding counterexamples. Or even higher level architectural optimizations. Surely they could afford at least a few % of the total time on this pre-optimization (although the tools to achieve this automatically would need to be quite sophisticated!).
- inteleng 8y agoWhat do you do when these mutations end up crashing constantly?
- hedora 8y agohttp://fftw.org http://fftw.org is optimized by doing massive parameter sweeps on each architecture (it only considers correct implementations). There are also a few “software synthesis” and “sketching” approaches that use a constraint solver to find all correct implementations of a high level spec, subject to some implementation pattern. Then they either try them all with brute force or pick the one that optimizes some objective function.
- stochastic_monk 8y agofftw isn't exactly developing an optimal kernel from scratch. It's testing a range of different methods and simply choosing the best for one's parameters. Facebook's Tensor Comprehensions framework, which generates CUDA kernels through a genetic algorithm, is closer to the sort of approach which would take greatest advantage of the hardware it's on.