4 ms·
This is more like solving programming than doing math, it either outsources by writing a sympy program or generates a brute-force or monte carlo simulation base
by Vetch 5y ago
This is more like solving programming than doing math, it either outsources by writing a sympy program or generates a brute-force or monte carlo simulation based answer.
My biggest gripe with this paper is how unclear it is on its methods. How many rewrites per question? How did they select the solved questions? They state they used a model trained on language and fine-tuned on code, but fine-tuning could describe either their or codex's process.
The biggest hint in favor of fine-tuning is that AFAICT, their dataset adds up to 205 questions but they've solved 235 questions. In which case I'd suspect overfitting on question form. In intro level math, problems are usually in template form and solving them boils down to matching on the correct template and slot filling the answer.
To prove whether it's been fine-tuned, people with davinci codex access should try to see if the can replicate this.
To prove it's not overfit, authors should release their training dataset if there is one and allow people to test with it.
How many parameters does davinci codex have? The original codex was 12 biliion IIRC and certainly wasn't this capable.
---
Some of the answers look wrong or incomplete.
Table 221 seems to be missing a factorial, such leniency suggests they must not be automatically scoring this.
Not sure what's going on in Table 144. Table 44 too.
In 42 and 47, particularly 42, the solution program seems incomplete.
211 is impressive, it also writes code for LU decomposition, permutations, runge kutta, card combinatorics and other clever stuff.
Even though for most of the more challenging problems, hard part was pre-digested or worse, it was hand fed the answer's algorithm, there were a few where the network had to have had a solid understanding of the math and code requirements. A program like they claim would revolutionize and hugely simplify mathematically and algorithmically involved programming.
The biggest cause for worry that it might have overfit is that it works with no sampling at all and gets a perfect score (they claim).
- Isinlor 5y agoThey are using OpenAI Davinci-Codex, comparable to GPT3-175B, without any further fine-tuning as far as I can tell. Over-fitting, in the traditional sense, specifically to these questions is very unlikely. What is problematic is at the bottom of page 6: > Prompts that result in correct solutions, whether from original or modified prompts, are used for evaluation metrics. In other words, they reject all results without correct solutions and claim perfect accuracy... And as you noticed even their evaluation is really not good enough to establish perfect accuracy.
- Vetch 5y ago> as far as I can tell I would prefer if we could have more certainty than that. I have a list of questions. 1) Does no one else have access to this codex? Has anyone tried to replicate this? It should be easy. 2) If it was not fine-tuned, there is still the issue of an over-fit question set, as you say. How easy is it to break? How general is it? 3) Are there really so many extensive examples on the use of sympy? Some of my googling (I know the math but not sympy) did not support this but it could be my lack of familiarity of the library ecosystem. 4) And the biggest issue where when it gave answers to certain limit problems without any computation, which should not have been possible unless it computed it internally, complexity of which kinda contradicts the entire exercise. Depending on how generalizable this result is, it could be anything from merely impressive to mind blowing and revolutionary. Imagine if the language is Coq or Lean instead of sympy!
- Isinlor 5y ago1) I have access to OpenAI Davinci-Codex. I can and I did recreate their results on few prompts. 2) It is not really an issue of how easy it is to break. If it is able to solve the problem it will do it somewhat reliably. The issue is that questions provide solutions. So, "Solve each equation for x: ln(x*2-1)=3" will just keep on generating more similar problems, but "Using Sympy, solve Eq ln(x*2-1)=3 for x." will work. For example, you will never get this prompt to work: "Find the limits as x → ∞ and as x → −∞ . Use this information, together with intercepts, to give a rough sketch of the graph as in Example 12. y = x 2 (x 2 − 1) 2 (x + 2)" but their step by step what to do prompt does work. If you don't know Python ecosystem and you don't know how the solution should look like then you are out of luck. 3) Sympy does not seem hard to use. Codex is actually pretty good at writing "template" code. It struggles with anything that requires some level of reasoning instead of matching template to a problem. 4) Could you point me to the exact questions? I can try to recreate it. OpenAI Codex is actually quite good. It is groundbreaking in comparison to anything before it, but it is far, far, far from perfect. It's like really bad stackoverflow in the sense that if you probe it couple of times it will provide you reasonably looking solution that may or may not work. I would also not focus on the current state. Trajectory of improvement is more important. 5 years ago you could not get anything even comparable to Codex, in next 5 years it may just be far from perfect, and in 10 years it may actually work. On many metrics, on average over years, Deep Learning is improving 2x every year or so.