5 ms·
I was actually thinking of that paper when I wrote that comment, hence the frustration that we don't actually know the intermediates. Still, perhaps the steppe
by devmor 1y ago
I was actually thinking of that paper when I wrote that comment, hence the frustration that we don't actually know the intermediates.
Still, perhaps the stepped output we get may hint at that kind of "cheating" and can be used in reinforcement... or perhaps that kind of reinforcement will just make the LLMs better at cheating. The problem is definitely a lot more complex than the trivial way I referenced it, at least.