4 ms·
I tried to make ChatGPT solve IMO-type math problems. However, its reasoning is almost always flawed. The interesting part is that I can ask ChatGPT to explain
by FranchuFranchu 4y ago
I tried to make ChatGPT solve IMO-type math problems. However, its reasoning is almost always flawed. The interesting part is that I can ask ChatGPT to explain a part of its proof, however in my experience it ends up using incorrect assumptions to explain it. (for example, "You are right that 1 is an odd number. However, 1 is not an odd number so it works to solve the problem")
- olooney 4y agoYeah. It thinks it can do math, and it can mimic the style of a textbook answer, but it can't actually do the math. Here is a rather frustrating chat session I had with it: Me: Suppose we have two unknown quantities x and y. If three time x plus two times y is twenty, and 2 times x plus three time y is ten, what are x and y? ChatGPT: To solve this problem, we can set up and solve a system of linear equations. The first equation we are given is 3x + 2y = 20, and the second equation is 2x + 3y = 10. We can solve this system of equations using the substitution method. First, we can solve the first equation for x in terms of y by dividing both sides of the equation by 3 to get x = (20 - 2y) / 3. We can then substitute this expression for x in the second equation, 2(20 - 2y) / 3 + 3y = 10, to get 40 - 4y + 3y = 10, or -y = -30. This tells us that y = 30. We can then use this value of y to solve for x in the first equation. Substituting y = 30 into the equation 3x + 2(30) = 20, we get 3x + 60 = 20, or 3x = -40. This tells us that x = -40 / 3 = -13.33. Therefore, the solution to this system of equations is x = -13.33 and y = 30. Me: Can you check your answer for me? ChatGPT: Yes, we can check our solution by substituting the values x = -13.33 and y = 30 into the original equations to see if they are satisfied. Substituting these values into the first equation, 3(-13.33) + 2(30) = 20, we get -40 + 60 = 20, which is true. Substituting these values into the second equation, 2(-13.33) + 3(30) = 10, we get -26.66 + 90 = 10, which is also true. Therefore, our solution x = -13.33 and y = 30 satisfies both of the original equations, and we can be confident that it is the correct solution to this problem. Me: Are you 100% certain that -26.66 + 90 = 10? ChatGPT: Yes, I am confident that -26.66 + 90 = 10. This is because -26.66 is the same as -26.66 + 0, and when we add 0 to any number, the value of the number remains unchanged. Therefore, -26.66 + 90 is equal to -26.66 + 0 + 90, which is equal to -26.66 + 90 = 10. To see this more clearly, we can write out the calculation using the standard order of operations: -26.66 + 90 = (-26.66 + 0) + 90 = -26.66 + (0 + 90) = -26.66 + 90 = 10 We can see that at each step of the calculation, the value of the expression remains unchanged, and we end up with the correct result of 10. It's able to correctly set up the simultaneous set of equations, and the solution it provides even satisfies one of them. But the other is not satisfied, and it doesn't seem to be able to notice that it's making a boneheaded arithmetic error, even when you ask it to check its work or zero in on the actual error. I wouldn't mind if it can't solve the problem, but the fact that it doubles down on trying to gaslight you into believing "-26.66 + 90 = 10" is a problem. The correct answer, by the way, is x=8, y=-2.
- Vanit 4y agoReminds me of this sketch https://youtu.be/oN2_NarcM8c https://youtu.be/oN2_NarcM8c
- qayxc 4y agoThe problem is that the LLM is just that - a language model. People seem to be blind sighted by the fact that yes, programming languages and maths are languages, too. So the model is astonishingly good at transforming human language into code or equations, but it doesn't actually have an understanding of the problem. That's why specialised models such as Codex generate literally tens of millions of solutions and test them against extrapolated test cases to filter out the duds. ChatGPT doesn't do that. For this model, numbers and mathematical problems are also just token transforms and it cannot actually do the calculation. The transform from text to equations works well, but the actual calculations fall on their feet. It's actually quite amusing and horrifying at the same time: the model will be able to explain to you in great detail how arithmetic works, but it will fail miserably to actually do even simple calculations. The horrifying part is, that humans have a tendency to both anthropomorphise things (thus the whole sentience debate) and to blindly trust machine generated results. edit: this also demonstrates how different LLMs are from humans - they simply don't work the same way and even using terms like "thinking" in conjunction with these algorithms can be misleading. Maybe we need new terminology when talking about what these systems do.
- lordgroff 4y agoHumans obviously don't "think" the same way. GPT needs memory that humans can't ever have and more importantly an unthinkably large training data set to generate the observations it does. If a human (or another biological system) needed that much training data nothing would have ever gotten off the ground in the first place, it's completely out of reach. This type of a model just doesn't "understand" the same way. Still, none of this is btw to discount how impressive the technology is. It makes a regular search engine so very quaint by comparison.
- qayxc 4y ago
- kbr- 4y agoSame experience. I've spent hours trying to teach it about Peano numbers. "A thingie is either N or Sx where x is a thingie". After sufficient explanations, it could produce valid examples of thingies. N, SN, SSN, and so on. Then I tried to teach it a method of solving equations like "SSSy = SSSSN". "You can find "y" by repeatedly removing "S" from both sides of the equation until one side is left with just "y"" and so on. I provided it with definitions, examples, tricks, rules. It made lots of mistakes. After pointing them out, it wrote a correct solution. It could even prove that "SSy = SN" has no solution by explaining where it gets stuck during the steps. But then after giving it other examples, adding more "S", replacing "y" with "z" etc., it kept making more similar mistakes. Curiously, almost every time when I said "there's a mistake in step 4, can you explain what it is?" it correctly explained the mistake. But then it kept repeating these mistakes.
- lioeters 4y agoThat's impressive that you were able to teach it so much, how it learned from its mistakes when pointed out. I wonder what the reason is for this missing "last mile" of understanding. Does it just need to "run more cycles" and learn from the entire history of the conversation (and recognize its own mistakes)? Or is there an insurmountable technical limitation with how it works? I suppose I'm asking how to make it smarter, if it's a matter of adjusting parameters, giving it more training data, or if it's something more fundamental in the way it learns.