4 ms·
I'm not sure about its math, but GPT-4 fails miserably at simple arithmetic questions like 897*394=? The GPT-3.5 turbo is fined-tuned for arithmetic according
by dongping 3y ago
I'm not sure about its math, but GPT-4 fails miserably at simple arithmetic questions like 897*394=?
The GPT-3.5 turbo is fined-tuned for arithmetic according to ClosedAI (noted in one of the change logs), so it is sometimes slightly better, but nevertheless always fails equations like 4897*394=?
- Filligree 3y ago> I'm not sure about its math, but GPT-4 fails miserably at simple arithmetic questions like 897*394=? That's, um, about 300,000? ... 353,418 actually. But I'm not going to blame the AI too much for failing at something I can't do either.
- dongping 3y agoOne can resort to traditional vertical multiplication (which requires patience), or do 897*394 = (900-3) * (400-6) = 900*400 - 6*900 - 400*3 + 3*6 = 360,000 - (5,400 + 1,200) + 18 = 360,018 - 6,600 = 353,418
- 6510 3y ago8*3=24 and 800*300 =240000 8*9=72 and 800* 90 = 72000 8*4=32 and 800* 4 = 3200 9*3=27 and 90*300 = 27000 9*9=81 and 90* 90 = 8100 9*4=36 and 90* 4 = 360 7*3=21 and 7*300 = 2100 7*9=63 and 7* 90 = 630 7*4=28 and 7* 4 = 28 -------------------------- 353418
- dash2 3y agoBut you are smart enough to use a computer or calculator. And AI is a computer. So the naive expectation would be that it would be capable of doing as well as a computer. Also, you probably could do long multiplication with paper and pencil if you needed to. So a reasoning AI (which has read many many descriptions of how to do long multiplication) should be able to also.
- Filligree 3y agoI am indeed smart enough to do that. And so is the AI, if you use the right AI. (I.e, code interpreter.)
- ambrozk 3y ago> And AI is a computer. So the naive expectation would be that it would be capable of doing as well as a computer. Why would you judge an AI against the expectations of a naive person who doesn't understand capabilities AIs are likely to have? If an alien came down to earth and concluded humans weren't intelligent because the first person it met couldn't simulate quantum systems in their head, would that be fair?
- walleeee 3y agoI dunno, I simulate quantum systems (you, myself, my friends) in my head all the time
- dash2 3y agoThe original question was whether LLM's are "smart" in a human-like way. I think that if you gave a human a computer, he'd be able to solve 3-digit multiplications. If LLM's were human-like smart, they could do this too.
- Timon3 3y agoDid someone train LLMs with "access" to a computer? If not, why would you expect them to be able to use something they have never seen?
- mabster 3y agoI've got an engineers style mindset for these kind of calculations. 897 is about 900. 394 is about 400. 900×400 = 360,000. Only 2% error!
- mistercow 3y agoArithmetic is a pretty pathological case for ChatGPT because of BPE. Digits just tokenize in a way that makes arithmetic way more complicated. That said, I just fed both of your examples into GPT-4 and it answered them correctly without using CoT.
- dongping 3y agoThis was the response that I got from the GPT-4 API yesterday: {'id': 'chatcmpl-7uVF5xGqR1oEzXITw3WZYsnB4Yzt8', 'object': 'chat.completion', 'created': 1693700819, 'model': 'gpt-4-0613', 'choices': [{'index': 0, 'message': {'role': 'assistant', 'content': '353538'}, 'finish_reason': 'stop'}], 'usage': {'prompt_tokens': 11, 'completion_tokens': 2, 'total_tokens': 13}} Maybe they fine-tuned the ChatGPT version better, or fed it to an calculator.
- mistercow 3y agoI would guess that they have a later model on web than on API (I also see worse results on API with 0613). Further testing shows that it loses the plot after a few more digits, which wouldn't make sense if they were injecting calculations.
- hboon 3y agoI think I understand your logic, but ChatGPT+GPT-4 gave me the correct answer for "What is 897*394?" https://chat.openai.com/share/00f94e43-c353-400a-858a-50c10cf8eb4e https://chat.openai.com/share/00f94e43-c353-400a-858a-50c10c... (GPT-3 gave the wrong numeric answer though)
- dongping 3y agoThanks for testing it. I canceled my ChatGPT Plus a few months ago (when they changed the color from black to purple IIRC). So I only tested that with the GPT-4 API, with the following results: {'id': 'chatcmpl-7uVF5xGqR1oEzXITw3WZYsnB4Yzt8', 'object': 'chat.completion', 'created': 1693700819, 'model': 'gpt-4-0613', 'choices': [{'index': 0, 'message': {'role': 'assistant', 'content': '353538'}, 'finish_reason': 'stop'}], 'usage': {'prompt_tokens': 11, 'completion_tokens': 2, 'total_tokens': 13}}
- chaxor 3y agoIt absolutely FAILS for the simple problem of 1+1 in what I like to call 'bubble math'. 1+1=1 Or actually, 1+1=1 and 1+1=2, with some probability for each outcome. Because bubbles can be put together and either merge into one, or stay as two bubbles with a shared wall. Obviously this can be extended and formalized, but hopefully it also displays that mathematics isn't even guaranteed to provide the same answer for 1+1, since it depends on the context and rules you set up (mod, etc). I should also mention that GPT-4 does quite astoundingly good at this type of problem wherein new rules are made up on the fly. So in-context learning is powerful, and the idea that it 'just regurgitates training data' for simple problems is quite false.
- Closi 3y ago> [GPT] always fails equations like 4897 x 394=? In some ways, I think we should treat GPT like a human without access to a calculator. If you ask a human what 4897 x 394 is, they will struggle.
- isaacfung 3y agoSometimes chatgpt fails at tasks like counting number of a letter in a short string or checking if two nodes are connected in a simple forest of 7 nodes, even with chain of thoughts prompt. Human can solve those pretty easily.