3 ms·
KIND OF if you've tried to get o1 to give you outputs in a specific format, it often just tells you to take a hike. It's a stubborn model, which implies a lot
by yo-cuddles 2y ago
KIND OF
if you've tried to get o1 to give you outputs in a specific format, it often just tells you to take a hike. It's a stubborn model, which implies a lot
This is speculation, but it seems that the main benefit of reasoning models is that they provide a dimension along which RL can be applied to make them better at math and maybe coding, things with verifiable outputs.
Reasoning models likely don't learn better reasoning from their hidden reasoning tokens, they're 1) trying to find a magic token which when raised to its attention make it more effective (basically give it room to say something that jogs its memory) or 2) it is trying to find a series of steps which do a better job of solving a specific class of problem than a single pass does, making it more flexible in some senses but more stubborn along others
Reasoning data as training data is a poison pill, in all likelihood, and just makes a small window of RL vulnerable problems easier to answer (when we have systems that don't better). It doesn't really plan well, doesn't truly learn reasoning, etc
Maybe seeing the actual output of o3 will change my mind but I'm horrifically bearish on reasoning models