3 ms·
> as far as I can tell I would prefer if we could have more certainty than that. I have a list of questions. 1) Does no one else have access to this codex? Ha
by Vetch 5y ago
> as far as I can tell
I would prefer if we could have more certainty than that. I have a list of questions.
1) Does no one else have access to this codex? Has anyone tried to replicate this? It should be easy.
2) If it was not fine-tuned, there is still the issue of an over-fit question set, as you say. How easy is it to break? How general is it?
3) Are there really so many extensive examples on the use of sympy? Some of my googling (I know the math but not sympy) did not support this but it could be my lack of familiarity of the library ecosystem.
4) And the biggest issue where when it gave answers to certain limit problems without any computation, which should not have been possible unless it computed it internally, complexity of which kinda contradicts the entire exercise.
Depending on how generalizable this result is, it could be anything from merely impressive to mind blowing and revolutionary. Imagine if the language is Coq or Lean instead of sympy!
- Isinlor 5y ago1) I have access to OpenAI Davinci-Codex. I can and I did recreate their results on few prompts. 2) It is not really an issue of how easy it is to break. If it is able to solve the problem it will do it somewhat reliably. The issue is that questions provide solutions. So, "Solve each equation for x: ln(x*2-1)=3" will just keep on generating more similar problems, but "Using Sympy, solve Eq ln(x*2-1)=3 for x." will work. For example, you will never get this prompt to work: "Find the limits as x → ∞ and as x → −∞ . Use this information, together with intercepts, to give a rough sketch of the graph as in Example 12. y = x 2 (x 2 − 1) 2 (x + 2)" but their step by step what to do prompt does work. If you don't know Python ecosystem and you don't know how the solution should look like then you are out of luck. 3) Sympy does not seem hard to use. Codex is actually pretty good at writing "template" code. It struggles with anything that requires some level of reasoning instead of matching template to a problem. 4) Could you point me to the exact questions? I can try to recreate it. OpenAI Codex is actually quite good. It is groundbreaking in comparison to anything before it, but it is far, far, far from perfect. It's like really bad stackoverflow in the sense that if you probe it couple of times it will provide you reasonably looking solution that may or may not work. I would also not focus on the current state. Trajectory of improvement is more important. 5 years ago you could not get anything even comparable to Codex, in next 5 years it may just be far from perfect, and in 10 years it may actually work. On many metrics, on average over years, Deep Learning is improving 2x every year or so.