5 ms·
Math and coding competition problems are easier to train because of strict rules and cheap verification. But once you go beyond that to less defined things su
by snemvalts 6mo ago
Math and coding competition problems are easier to train because of strict rules and cheap verification.
But once you go beyond that to less defined things such as code quality, where even humans have hard time putting down concrete axioms, they start to hallucinate more and become less useful.
We are missing the value function that allowed AlphaGo to go from mid range player trained on human moves to superhuman by playing itself.
As we have only made progress on unsupervised learning, and RL is constrained as above, I don't see this getting better.
- charcircuit 6mo agoLLMs already do unsupervised learning to get better at creative things. This is possible since LLMs can judge the quality of what is being produced.
- zozbot234 6mo agoThis is not formally verified math so there is no real verifiable-feedback aspect here. The best models for formalized math are still specialized ones. although general purpose models can assist formalization somewhat.
- zar1048576 6mo ago[dead]
- otabdeveloper4 6mo agoLLMs can often guess the final answer, but the intermediate proof steps are always total bunk. When doing math you only ever care about the proof, not the answer itself.
- eru 6mo agoOnce you have a working proof, no matter how bad, you can work towards making it nicer. It's like refactoring in programming. If your proof is machine checkable, that's even easier.
- prmoustache 6mo agoThat is also how humans work mostly. Once every full moon we may get an "intuition" but most of the time we lean on collective knowledge, biases and behavior patterns to take decisions, write and talk.
- otabdeveloper4 6mo agoI haven't had success in getting AI's to output working proofs. You'd need a completely different post-training and agent stack for that.
- jamesfinlayson 6mo agoYep, I remember a friend saying they did a maths course at university that had the correct answer given for each question - this was so that if you made some silly arithmetic mistake you could go back and fix it and all the marks were for the steps to actually solve the problem.
- number6 6mo agoThis would have greatly helped me. I always was at a loss which trick I had to apply to solve this exam problem, while knowing the mathematics behind it. Just at some point you had to add a zero that was actually a part of a binomial that then collapsed the whole fromula
- datsci_est_2015 6mo agoWhat’s funny is that there are total cranks in human form that do the same thing. Lots of unsolicited “proofs” being submitted by “amateur mathematicians” where the content is utter nonsense, but like a monkey with a typewriter, there’s the possibility that they stumble upon an incredible insight.
- dash2 6mo agoNot in this case: the LLM wrote the entire paper, and anyway the proof was the answer.
- jack_pp 6mo agoMaybe to get a real breakthrough we have to make programming languages / tools better suited for LLM strengths not fuss so much about making it write code we like. What we need is correct code not nice looking code.
- kuerbel 6mo agoYes yes Let it write a black box no human understands. Give the means of production away.
- eru 6mo agoLean might be a step in that direction.
- kube-system 6mo agoIf you can’t validate the code, you can’t tell if it’s correct.
- 3836293648 6mo agoNo? That's literally the thing they suggested to move away from. That is just an issue when using tools designed for us. Make them write in formal verification languages and we only have to understand the types. To be clear, I don't think this is a good idea, at least not yet, but we do not have to always understand the code.
- bloppe 6mo ago> programming languages / tools better suited for LLM strengths The bitter lesson is that the best languages / tools are the ones for which the most quality training data exists, and that's pretty much necessarily the same languages / tools most commonly used by humans. > Correct code not nice looking code "Nice looking" is subjective, but simple, clear, readable code is just as important as ever for projects to be long-term successful. Arguably even more so. The aphorism about code being read much more often than it's written applies to LLMs "reading" code as well. They can go over the complexity cliff very fast. Just look at OpenClaw.
- raincole 6mo agoExcept it's not how this specific instance works. In this case the problem isn't written in a formal language and the AI's solution is not something one can automatically verify.
- NitpickLawyer 6mo ago> I don't see this getting better. We went from 2 + 7 = 11 to "solved a frontier math problem" in 3 years, yet people don't think this will improve?
- number6 6mo agoBut can it count the R's in strawberry?
- Aditya_Garg 6mo agoyes its ridiculously good at stuff like that now. I dare you to try and trick it.
- frizlab 6mo agohttps://news.ycombinator.com/item?id=47495568 https://news.ycombinator.com/item?id=47495568
- thedatamonger 6mo agowhat bothers me is not that this issue will certainly disappear now that it has been identified, but that that we have yet to identify the category of these "stupid" bugs ...
- sigmoid10 6mo agoWe already know exactly what causes these bugs. They are not a fundamental problem of LLMs, they are a problem of tokenizers. The actual model simply doesn't get to see the same text that you see. It can only infer this stuff from related info it was trained on. It's as if someone asked you how many 1s there are in the binary representation of this text. You'd also need to convert it first to think it through, or use some external tool, even though your computer never saw anything else.
- 6mo ago
- typs 6mo agoI mean, this is why everyone is making bank selling RL environments in different domains to frontier labs.
- eptcyka 6mo agoDo we need all that if we can apply AI to solve practical problems today?
- fmbb 6mo agoDepends on the cost.
- computably 6mo agoWhat is possible today is one thing. Sure people debate the details, but at this point it's pretty uncontroversial that AI tooling is beneficial in certain use cases. Whether or not selling access to massive frontier models is a viable business model, or trillion-dollar valuations for AI companies can be justified... These questions are of a completely different scale, with near-term implications for the global economy.
- pjerem 6mo agoI mean, even if the technology stopped to improve immediately forever (which is unlikely), LLMs are already better than most humans at most tasks. Including code quality. Not because they are exceptionally good (you are right that they aren’t superhuman like AlphaGo) but because most humans are rather not that good at it anyway and also somehow « hallucinate » because of tiredness. Even today’s models are far from being exploited at their full potential because we actually developed pretty much no tools around it except tooling to generate code. I’m also a long time « doubter » but as a curious person I used the tool anyway with all its flaws in the latest 3 years. And I’m forced to admit that hallucinations are pretty rare nowadays. Errors still happen but they are very rare and it’s easier than ever to get it back in track. I think I’m also a « believer » now and believe me, I really don’t want to because as much as I’m excited by this, I’m also pretty much frightened of all the bad things that this tech could to the world in the wrong hands and I don’t feel like it’s particularly in the right hands.
- anabis 6mo ago> But once you go beyond that to less defined things such as code quality I think they have a good optimization target with SWE-Bench-CI. You are tested for continuous changes to a repository, spanning multiple years in the original repository. Cumulative edits needs to be kept maintainable and composable. If there are something missing with the definition of "can be maintained for multiple years incorporating bugfixes and feature additions" for code quality, then more work is needed, but I think it's a good starting point.