4 ms·
I work professionally in computational optimization, so here are some of my observations coming from the world of mechanics. In my experience, it's because the
by kxyvr 3y ago
I work professionally in computational optimization, so here are some of my observations coming from the world of mechanics.
In my experience, it's because the client is trying to reuse an existing simulation code for something that requires optimization. This could be something like imaging (parameter estimation), calibration, or control. For example, say someone has something like a PDE solver that can simulate sound waves going through a media. If you can match the output of this simulation to some collected data, the parameters for the simulation will tell you what material is where.
Anyway, the issue is that these existing codes are rarely instrumented or derived in a way to make it easy to get derivatives. Management doesn't want pay for a rewrite, so they demand that existing equity (the simulator) be used. Then, a derivative free optimizer (DFO) gets slapped on as a first attempt. To be clear, this doesn't work particularly well, but it tends to be the place where I most often see these algorithms used.
As to why it doesn't work very well, derivative based optimizers can scale to the order of hundreds of millions of variables without too many issues (modulo the kind of problem, but it works very well for a lot of problems of value). DFOs tend to have issues past a few dozen variables.
Now, for anyone not wanting to end up in this conundrum, when you write your original simulation code parameterize on the floating point type. Basically, don't write the code with double, use some kind of generic, template parameter, or whatever on something we'll call MyFloat. The purpose here is to allow the code to be more easily compatible with automatic differentiation (AD) tools, which will give a real and machine precision exact derivative. Don't do finite differences. Don't think reinstrumenting later is easy. It's not. It's a pain. It's expensive. And, to be clear, AD can be done poorly and it can be slow. However, I've never seen a code instrumented with AD and a good gradient based optimizer do more poorly than a derivative free optimization code. They've always done dramatically better. Beyond just giving the derivative, I'll also mention that even if the client wants to recode the simulator to give an analytic derivative for speed, it is vastly easier to do so if the results can be checked with a simulator instrumented with AD. There's also other tricks of value that can be used in a properly instrumented code such as interval arithmetic for some kinds of sensitivity analysis.
Anyway, there are other reasons to use derivative free optimization. Some functions just flat out don't have derivatives. However, in the world of mechanics, most do and DFOs gets abused due to poor prior planning. In either case, please, for the love of math, just parameterize on your floating point type. It would make my job much easier.
- domcoltd 3y agoAgreed. If (approximate) 1st-order information can be obtained in any way (even by *carefully* deployed finite difference), then gradient-based methods should be used. It is wrong to use DFO in such cases. > DFOs tend to have issues past a few dozens variables. It seems that the PRIMA git repo https://github.com/libprima/prima https://github.com/libprima/prima shows results on problems of hundreds of variables. Not sure how to interpret the results.
- domcoltd 3y ago> DFOs tend to have issues past a few dozens variables. It is interesting to check the results in section 5.1 of the following paper: https://arxiv.org/pdf/2302.13246.pdf https://arxiv.org/pdf/2302.13246.pdf where these algorithms are compared with finite-difference CG and BFGS included in SciPy.