3 ms·
It's also extremely hard to nail down in much of mathematics or computer science! - is such-and-such theorem deep or shallow? - is this definition/axiom usefu
by aithrowawaycomm 2y ago
It's also extremely hard to nail down in much of mathematics or computer science!
- is such-and-such theorem deep or shallow?
- is this definition/axiom useful? (there's a big difference between doing compass-straightedge proofs vs. wondering about the parallel postulate)
- more generally, discovering theorems is generally not amenable to verifiable rewards, except in domains where simpler deterministic tools exist (in which case LLMs can likely help reduce the amount of brute forcing)
- is this a good mathematical / software model of a given real-world system?
- is the flexibility of dynamic/gradual typing worth the risk of type errors? is static typing more or less confusing for developers?
- what features should be part of a programming language's syntax? should we opt for lean-and-extensible or batteries-included?
- are we prematurely optimizing this function?
- will this program's memory needs play nicely with Rust's memory model? What architectural decisions do we need to make now to avoid headaches 6 months down the line?
- Davidzheng 2y agoNot clear to me that theorem discovery is not amenable to verifiable rewards. I think most important theorems probably are recovered automatically by asking AI systems to proof increasing complicated human conjectures. Along the way I expect emergent behaviors of creating conjectures and recognizing important self-breakthroughs. Much like regret emergence
- youoy 2y agoTheorems discovery is amenable to verifiable rewards. But is meaningful theorems discovery too? Is the ability to discern between meaningful theorems and bad ones an emergent behaviour? You can check for yourself examples of automatic proofs, and the huge amount of intermediate theorems that they can generate which are not very meaningful.
- janalsncm 2y agoUnless you can quantify what you mean by “meaningful” then it won’t be possible. It can’t read your mind.
- janalsncm 2y agoFor questions with a correct answer, you don’t need to verify the reasoning process. RL training will discover it. That’s R1-Zero. The point of R1 was to fix problems with the reasoning tokens and expand to subjective domains like creative writing.