3 ms·
What would an LLM have to do to convince you it was good at math? Check out this recent post by OpenAI where one of their models is solving 60%+ of problems fro
by pgspaintbrush 3y ago
What would an LLM have to do to convince you it was good at math? Check out this recent post by OpenAI where one of their models is solving 60%+ of problems from a high school math competition dataset: https://openai.com/research/improving-mathematical-reasoning-with-process-supervision#fn-2 https://openai.com/research/improving-mathematical-reasoning...
- wrs 3y agoIt’s actually better at math than it is at arithmetic, and I think this discussion has been about arithmetic. I could make up something about how math is more like language than arithmetic is. I suspect the hypothesis that math tests tend to have a lot of stereotypical problem structures from a shared curriculum is also relevant. But who knows at this point? Anyway, to convince me it’s good at arithmetic is not complicated…just be good at arithmetic! That is do it correctly, every time, for any size number.
- famouswaffles 3y ago>That is do it correctly, every time, for any size number. Then no human is good at arithmetic.
- wrs 3y agoFair enough, I’ll allow a 1% error rate per 10 addend digits.
- SkyPuncher 3y agoI suspect most people on this forum can do arithmetic for any "reasonable" size number. It might take weeks to complete, but most people on this forum can calculate large numbers by hand.
- famouswaffles 3y agoPost moving. "Reasonable" is just an arbitrary line. Especially since most if not all would make some mistake somewhere along the line. You can greatly increase GPT's arithmetic capabilities tackling it like a problem to solve "on paper" in context. And this was done on 3.5 not 4. https://arxiv.org/abs/2211.09066 https://arxiv.org/abs/2211.09066
- 8note 3y agoIf its going to take weeks, most people will get it wrong. That's a lot of calculations to never get wrong and never misinterpret some prior note you left
- Tainnor 3y agoOkay, but we have since invented machines that can do arithmetic correctly, every time. When we try to do maths via an LLM, we're just throwing all of that away.
- famouswaffles 3y agoSo ? I didn't tell you to use GPT-4 for arithmetic over a calculator. I simply pointed out that the only standard where GPT-4 is not good at arithmetic is a standard humans wouldn't fit the bill either. Especially since zero shot "mental" arithmetic is not even close to GPT-4 at its most accurate.
- Tainnor 3y agoThe discussion started "what would it take to convince people that [insert favourite LLM] is good at maths", and the response to that IMHO is that we have much better tools to do arithmetic (I don't even want to say maths), even if humans themselves are also poor at arithmetic. What's the point of building a system to be equally bad as humans at something that we know humans are bad at? LLMs have their uses but (at least at the current stage) performing arithmetic calculations is not one of them (to say nothing of more advanced mathematics).
- Tainnor 3y agoChatGPT, and probably GPT-4 too, is also hilariously bad at "more advanced" mathematics, including trying to come up with even slightly original proofs.