5 ms·
It’s great, but it only works for problems where there is exactly one correct solution and it’s possible to automatically verify the solution - like math and pr
by thorum 2y ago
It’s great, but it only works for problems where there is exactly one correct solution and it’s possible to automatically verify the solution - like math and programming. So far these reasoning models have not shown much transfer learning of reasoning to other domains, and are often worse at non-math/code tasks than standard models.
- datameta 2y agoNot sure if this a trivial or naive thought - is that perhaps because in non-discrete ideas there is more granularity of information encoded compared to a numerical solution? Separatelt, but on a related note - do we need analog or quantum computing to "truly" scale?
- cluckindan 2y agoTo ”truly” scale we need a system directly doing calculation on some fundamental property of matter. Quantum computing is _close_ but suffers from a lack of interpretability: how would we know if our quantum simulation of the universe would include the quantum computer doing the simulation, as well as another recursive copy of the simulation, and another, and another…
- onlyrealcuzzo 2y agoThere are infinite solutions to coding problems. You could have lots of comments and pass statements and unnecessary conditionals. I don't think it matters that there is only one correct answer. It matters that you can verify reliably if the answer is correct enough.
- JKCalhoun 2y agoSo, we have a very good left-brain model.
- vintermann 2y agoWe can probably make it work at more nebulous problems by using a big LLM to judge the quality of the answers as well. It should be easier to recognise e.g. a really good poetic translation than to make one, and as long as that's true it could benefit from internal monologue "reasoning" in those domains as well.
- timr 2y agoThis has been done, and doesn't work especially well. See, for example, GANs, which are difficult to train and vulnerable to mode collapse. Not saying that it's impossible, but reinforcement learning of the form shown by DeepSeek was particularly well-established and robust. It's sort of ironic, actually...IIRC, OpenAI started in the world of this kind of reinforcement learning.
- deleted 2y ago[deleted]
- aurareturn 2y agoLLMs can verify math problems if you give it a calculator tool. Coding problems if you give it a Linux environment. So why can’t we give LLMs other virtual environments to verify if the solution is correct? For example, a stock simulator, a physics simulator, a driving simulator, etc.
- cluckindan 2y agoBecause simulations are not perfect models. A particle physics simulation simply cannot be accurate enough to provide the 5-sigma (99.99994%) confidence typically required in particle physics. Not even measurement is simple when the requirements are so strict, that’s why they spend billions building those huge accelerators.
- onlypassingthru 2y agoOnce the computer learns to apply the logic of Tic-Tac-Toe to Global Thermonuclear War, we will finally have the W.O.P.R.[0] [0]https://en.wikipedia.org/wiki/WarGames https://en.wikipedia.org/wiki/WarGames