4 ms·
surely the LLM can do that. It is RL'd against some reward, if the known strategies are clearly suboptimal with easy improvement, it'll find it most likely
by Davidzheng 13d ago
surely the LLM can do that. It is RL'd against some reward, if the known strategies are clearly suboptimal with easy improvement, it'll find it most likely