3 ms·
Except their process isn't actually differentiable, as they admit near the end of the post, they just sort hand-wavily suggest that approximately differentiable
by D-Machine 7mo ago
Except their process isn't actually differentiable, as they admit near the end of the post, they just sort hand-wavily suggest that approximately differentiable methods "should" work. Also no mention at all of what the training data would be, where it would come from, or how a loss function could be constructed to continuously score "partially correct" programs (of what that would even mean, or if that idea is even coherent).
What was a good point, mentioned by @hedgehog in this thread (https://news.ycombinator.com/item?id=47367986 https://news.ycombinator.com/item?id=47367986), is that tool-calls break batching a lot, so there could be huge efficiency gains at scale if you can just pass through a computation sub-network (even if that sub-net is frozen and can't be updated, and is programmed in manually rather than trained in).
Why on Earth you'd want that sub-net to be a clunky transformer rather than just an efficient, GPU-accelerated custom non-trainable layer, though, is unclear to me.