3 ms·
Example?
by sudokuist 3y ago
Example?
- Nevermark 3y agoI was playing around with prime numbers, and simple made up relationships between them, such as between the square of a prime N vs. the set of primes smaller than N, etc. It caught me out with specific examples that violated my conjectures. In one case the conjecture held for all but one case, another conjecture was generally true but not for 2 and 3. In one case it thought a conjecture I made was wrong, and I had to push it to think through why it thought it was wrong until it realized the conjecture was right. As soon as it had its epiphany, it corrected all its logic around that concept. It was very simple stuff, but an interesting exercise. The part I enjoyed the most was seeing GPT-4's understanding move and change as we pushed back on each other's views. You miss out on that impressive aspect of GPT-4 in simpler sessions.
- sudokuist 3y agoHave you tried formalizing your ideas with Isabelle? It has a constraint solver and will often find counterexamples to false arithmetical propositions[1]. 1: https://isabelle.in.tum.de/overview.html https://isabelle.in.tum.de/overview.html
- Nevermark 3y agoI have not, thanks for the tip.
- riku_iki 3y agocurious why you referred specifically on isabelle, which looks ancient and over engineered, there are many other tools and langs in this area. I am not criticizing, but curious about your opinion.
- c-cube 3y agoIsabelle is good at counter examples in ways few other proof assistants are. In general its automation is excellent, partly because it uses a less powerful logic (HOL instead of CIC; more expressive logics are harder to write automation for). It's not obsolete.
- mannykannot 3y agoI have not been able to figure out how that would help in the context of this discussion. As I see it, what’s very interesting here is that an LLM is able to do this.
- riku_iki 3y agoI think the point is that LLM is not right tool for deep reasoning, and isabelle and others are much better such tools, even community trying to apply LLM in this area following current wave of hype.
- riku_iki 3y agoits hard to judge how deep and unique your conjectures were. I did similar testing of GPT4, and my observation is that it starts failing after 3-4 levels of reasoning depth.
- tux1968 3y ago> failing after 3-4 levels of reasoning depth. That sounds more like an implementation or resource limitation, rather than an inherent limitation of the technique in general.
- riku_iki 3y agoit is not obvious to me how you came to such conclusion. LLMs got lots of investments: 10s of billions of dollars and tons of compute, maybe more than any other tech in history, and can't crack 3 steps reasoning. It sounds like tech limitation..
- naasking 3y agoNone of these systems or their training sets have been specifically tailored to tackle abstract reasoning or math, so that seems like a premature conclusion. The fact that they're decent at programming despite that is interesting.
- vikramkr 3y agoThey're also brand new and at some undetermined part of the sigmoid curve. Trying to predict where you are on the curve while in the middle of a sigmoid is a fools errand, the best you can do is make random predictions and hope you are accidentally correct so you can become a pundit later
- golol 3y agoNice to see number of levels of reasoning depth mentioned. I personally believe the size of a (well-trained) LLM determines how many steps of reasoning in sequence it can approximate. Newer models get deeper and deeper, giving them deeper reasoning context windows. My hypothesis is that you don't need infinite reasoning depth, just a bit more than GPT-4 has. I think once you can tie your output together with thinking in terms of ~10+ reasoning steps you'll be very close to hunan performance.