3 ms·
This article inspired us to test whether EML trees can be trained using gradient descent for symbolic regression: instead of searching for the right tree, is it
by jesustabares 6mo ago
This article inspired us to test whether EML trees can be trained using gradient descent for symbolic regression: instead of searching for the right tree, is it possible to optimize one from start to finish?
We implemented the EML trees as a differentiable module of PyTorch. Each leaf is a softmax mixture over {0, 1, x}, and the tree is evaluated from the bottom up. The entire system can be trained with Adam.
Results: 7 of the 7 elementary functions (exp, ln, sqrt, x², x³, 1/x, sin) converge with ≤24 parameters at depth 3. Three (exp, ln, sqrt) achieve an RMSE < 0.005.
The main challenge is depth scaling. Random initialization at depth 4 always diverges: the exp() strings create towering exponential growth that cancels out the gradients. We tested 12 initialization strategies; only hierarchical hot-starting (training depth n-1 first) works. sin(x²) gets 12.9x better at depth 4 vs depth 3.
Two honest negative results: (1) The trained trees use continuous softmax mixtures, not discrete leaf assignments; therefore, numerical approximations, not exact formulas, are obtained for everything except exp(x). (2) A 49-parameter MLP and PySR outperform it in MSE by orders of magnitude. It's not a practical tool — it just shows that gradient descent can work on S → 1 | eml(S, S) without needing sin, exp, etc. as primitives.
Paper: https://doi.org/10.5281/zenodo.19592926 https://doi.org/10.5281/zenodo.19592926
Code: https://github.com/seetrex-ai/monolith https://github.com/seetrex-ai/monolith