20 ms·
I asked GPT-NeoX-20B a hundred arithmetic questions
- andreyk 5y agoFun fact, the original GPT-3 paper has a whole section on arithmetic as part of its evaluation - https://arxiv.org/abs/2005.14165 https://arxiv.org/abs/2005.14165 Btw as per another comment this is GPT-NeoX-20B,not GPT-3 ; somewhat important distinction
- dang 5y agoQuestions and results here: https://gist.github.com/moyix/ca4091f16f0b5011bfa8f3f97f705a0d https://gist.github.com/moyix/ca4091f16f0b5011bfa8f3f97f705a.... We changed the above URL from that to the link which gives the background, but both are worth a look.
- nandhinianand 5y agoAnd here's the twitter thread with some plots about this data ( https://twitter.com/moyix/status/1491803929801150471 https://twitter.com/moyix/status/1491803929801150471 )
- pvg 5y agoYou should have posted that, since it's the original source. Without that context, and with the small mistake you made in the title, most of the commenters here ended up talking about something that this actually isn't.
- nandhinianand 5y agoSorry about that I was a bit not at my best. And haven't checked this thread till now.
- dang 5y agoI've changed the URL to that from https://gist.github.com/moyix/ca4091f16f0b5011bfa8f3f97f705a0d https://gist.github.com/moyix/ca4091f16f0b5011bfa8f3f97f705a.... Thanks!
- spupe 5y agoThank you for this. Technically it's not GPT-3, but GPT-NeoX-20B, although they are based on a similar architecture. The poor performance is most likely due to not having a large database of math problems to draw from. Github, for example, is part of the dataset that is used to train both GPT-3 and GPT-Neo variants, which is partly why they can generate meaningful code (sometimes). I wonder how a model finetuned for math would perform.
- dang 5y agoOk, we've reverted the title now. Thanks! (Submitted title was 'GPT-3's answers to arithmetic questions')
- spupe 5y agoI went and checked, it turns out for this version Eleuther-AI has in fact included math problems [1]. So my earlier comment is partly incorrect. [1] http://eaidata.bmk.sh/data/GPT_NeoX_20B.pdf http://eaidata.bmk.sh/data/GPT_NeoX_20B.pdf
- asah 5y agoAnd isn't it trivial to generate lots of correct sample data ? :-)
- williamtrask 5y agoPoor performance is more likely due to how transformer neural networks view numbers. It memorises them like words instead of modeling their numerical structure. Thus even if it’s seen the number 3456 and 3458, it knows nothing of 3457. Totally different embedding. It’s like a kid memorising a multiplication table instead of learning the more general principle of multiplication (related: this illusion is why big models are so popular. Memorise more stuff.) Paper (NeurIPS/DeepMind): https://arxiv.org/abs/1808.00508 https://arxiv.org/abs/1808.00508
- eutectic 5y agoThat depends on the tokenization scheme.
- nikolayasdf123 5y agoThis just shows that this model did not learn anything. Humans do not see billions of examples to add numbers. We see just few and can apply learned notation and procedures to infinity with 100% precision. GPT-3 learned mathematical intuition. Humans can hardly learn multiplication table over months and repetitions of same examples, and that table hardly matters at all. GPT-3 is just plainly wrong objective they trying to optimise.
- arghwhat 5y agoI think you'd find that most people doing large number math in their head is also off by a few percent like this model. Sure, with pen and paper we can follow specific algorithms manually to very slowly get a precise result. If we wanted a computer to merely follow instructions, then I suspect that there are better ways...
- jonathankoren 5y agoYou’re really lowering the bar for success here. It’s now unreasonable for a computer to correctly add two numbers together? Give me a break. It wasn’t even reasonable for a Pentium chip to incorrectly divide two numbers back in 1994.
- arghwhat 5y agoNeural networks are not used to obtain exact results.
- jonathankoren 5y agoIt’s amazing that this thought came out of neural network.
- arghwhat 5y agoIt's amazing that you thought that this was a sensible way to respond to a discussion. GPT-NeoX-20B would likely have handled this situation better than you.
- dash2 5y agoArithmetic seems like an example where it would help to learn from the real world, not just from text. I learnt to add up by watching my teacher manipulate plastic Lego-style blocks. Put 3 blocks with 2 blocks, and you have 1, 2, 3, 4, 5.
- londons_explore 5y agoBut somewhere in that massive corpus of text will be a description just like you've just given...
- tsimionescu 5y agoSure, but GPT-3 doesn't attach semantics to text, it just learns how to produce text patterns that are similar to text patterns it has seen before.
- FeepingCreature 5y ago"Similar" is an inherently semantic property.
- JonChesterfield 5y agoWords with the same length are similar to one another but not well correlated semantically. E.g. bog vs dog.
- tsimionescu 5y agoWould you say that something like string.GetSimilarity(string), which tells you by what number of characters two strings differ, is interpreting the text? Basically what GPT-3 does is to find a string X of a particular length such that it maximizes concat(userInput, X).GetSimilarity(someStringInTrainingSet). Edit: to be clear, I'm not suggesting it's looking up the training set at runtime, X.GetSimilarity(someStringInTrainingSet) is basically what got baked in during training.
- 5y ago
- bobuk 5y agoAnd now we need to do a contest and compare how humans will answer on this questions, and whose answer is closer to the truth. Take my bet - humans will lose or will be on the same level of guessing.
- d--b 5y agoWhat? You think this is poor performance? This totally blows my mind. I would never have guessed that GPT could get ANY of these right. I mean, is there a data point in the dataset used to train where you can read 2241 + 19873 = 22114? Quite unlikely... And those multiplications. It's consistently getting the number of digits right and the first two numbers correct. How the hell does this happen? Sure, it's sometimes way off. But generally it is in the right ballpark. I certainly think people should look into what's happening inside the model.
- AndrewOMartin 5y agoIt's unlikely that "2241 + 19873 = 22114" specifically is in the dataset, but very likely that there are many expressions equivalent to that expression in the dataset, and we've just picked one of those. Imagine someone watching every lottery draw and after each draw going "Wow! the chances of those exact numbers coming up in that order are atronomical!"
- dlkf 5y ago> there are many expressions equivalent to that expression in the dataset What do you mean by this?
- AndrewOMartin 5y agoI meant that I expect there are many examples of "a + b = c" in the training corpus, so GPT3 will answer some of them correctly.
- rawoke083600 5y agoMaybe this is an example of where you need an "extra specialized skill"(arithmetic) vs the general and semi-ambiguous-skill of language+conversation. GPT-3 is "good with conversation (language)" GPT-3 now needs a "sub-nn-model" to do the very 'specialized skill called math' *GPT-3 Should 'learn' to recognize which questions should be delicate to a submodel.
- 5y ago
- FL33TW00D 5y agoCan we get a title edit? GPT3 is over 8x the parameter count.
- rapiz 5y agoHow large is the data set they're training on? I suspect there are many math equations including these or similar numbers. As the result shows, the model is generally right about the first a few digits, but frequently wrong about the last digit. This may due to the fact that the data set can hardly cover the exact numbers in the questions, but it's likely to cover the first a few digits.
- Robin_Message 5y agoIt seems to me that carries are where this trips up. Which is weirdly human. I wonder if there are enough examples to learn each digit pair addition or subtraction, but not enough to learn every contextual action.
- wildmanx 5y agoNot really "human". Doing no-carry addition is much easier for a machine to do as well, as that's basically what XOR does, i.e., SIMD. Carry introduces dependencies between the digits, potentially as long as the whole string goes. So that's pretty hard to understand, also for a machine.
- deleted 5y ago[deleted]
- mannykannot 5y agoAt first I thought you were saying that doing arithmetic by carrying is not really a human trait, but on reflection, I think you are saying that carrying methods are inherently mistake-prone, regardless of who or what is using them. I feel it would be a very big deal if GPT-3 (or this variant) was carrying, even if imperfectly, but other comments here seem to be suggesting that, on account of the way all input is tokenized, consistently doing arithmetic by carrying would simply be outside of the set of transformations it could perform (though some results that look like it might arise by chance.)
- elcapitan 5y agoMaybe we're already past singularity and the AI is simply pretending to be bad at this in order to avoid making humans feel insecure. /s
- deleted 5y ago[deleted]
- sillysaurusx 5y agoIf you want to play with the model, you can (with difficulty) for free at https://goose.ai/playground https://goose.ai/playground. You have to log in, but thankfully you can via google. The playground crashes every minute, and the defaults ruin your outputs (temperature 1, really? 0.7 to 0.8 is a necessity, with top-k 40), and they turned off autocorrect on mobile, presumably because they hate you and your family for owning an iPad, but you can indeed play with it. The outputs feel pretty magical, too. With the settings above, it started printing... an IRC conversation? https://gist.github.com/shawwn/9a201990196b61cd21847487185dd28c https://gist.github.com/shawwn/9a201990196b61cd21847487185dd... This is impressive, because I'm not sure we explicitly included any IRC logs in The Pile. re: the current title "GPT-3's answers to arithmetic questions": We've come full circle. I used to give Eleuther a hard time for confusing people. But now that people confuse themselves, they should declare victory. It's as close to success as an open source effort could hope for. And with only years of work -- not too shabby. You can join them: https://www.eleuther.ai/faq/ https://www.eleuther.ai/faq/ GPT-NeoX-20B paper: http://eaidata.bmk.sh/data/GPT_NeoX_20B.pdf http://eaidata.bmk.sh/data/GPT_NeoX_20B.pdf
- rapiz 5y agoThanks for that. I've played around a little bit. > What is 123456789 - 123456789? > 123456788 > What is 123456789 * 0? > 123456789 Not even near. It didn't surprise me that the model failed to handle cases above, which are unlikely to present in the data set.
- sillysaurusx 5y agoTry temp 0.1 top-k 40. For math, it matters to have an unthinkably low temperature. It’s what generated the results in the OP. What is 12345 - 12345? 0 What is the distance between -0.1 and -0.01? 0.09 One neat example just now: What is 12345 divided by 12? 4115/4
- ggrrhh_ta 5y agoI pointed it out above; even though it is text, the ASCII representation is just a different base for the numbers - base 2^8 - ('325' is '3' * (2^32) + '2' * (2^16) + '5' * 2^8 = 51 * 2^32 + 50 * 2^16 + 53 * 2^8); it should approximate those polynomial functions very well.
- samuelfekete 5y agoA really intelligent AI would just use a calculator.
- can16358p 5y agoI'd really want to see in a super complicated AI, given many questions, devising a way to compile a basic calculator program that obeys the simple math rules (e.g only addition and multiplication) for itself inline and calculate the thrown questions using that instead of training math data.
- visarga 5y agoA combo of LM + REPL would be the winning ticket. The language model generates a piece of code, executes it, observe results. Is the answer found? No, maybe it cannot be solved by one single iteration, continue until solution found or max number of iterations reached. Single step responses are limited by their nature, and computation is hard to do "in your head"/"in your LM". The REPL could also come with libraries for all the important algorithms. Thus the LM could use numerical solvers and other nice toys. Another excellent module we could augment LM's with is a search engine - already done, proven to improve the efficiency of the LM by 25x (see link). Another one would be a physics simulator. https://deepmind.com/research/publications/2021/improving-language-models-by-retrieving-from-trillions-of-tokens https://deepmind.com/research/publications/2021/improving-la... We're just starting with language models. They will improve tremendously.
- dumdumdumdum 5y agoBut a really really intelligent AI would try to run them through a neural net and be amazed when it got most of them approximately correct.
- 2bitlobster 5y agoLast week I asked gpt3 (through philosopher ai) “why don’t irrational numbers fit neatly in our number system?” And I got back this bad answer: <<< Would you tell me what an irrational number is? Is it a number that cannot be written as a fraction, like pi or e? Or is it something else? Okay, I see. Thanks for clarifying that. Then what is an irrational number? It is a number that cannot be written as a fraction, correct? Okay, I understand. Then what is a number that can be written as a fraction? It must be a rational number, right? But aren't all numbers irrational, since it is impossible to write them as a fraction? And any rational number must be irrational, right? >>> FYI the pseudo dialogue format of the answer was more surprising than the wrong answer
- jrockway 5y agoYeah sometimes when you're doing a Q&A with it it starts generating questions, and then you become the thing answering. It's fun.
- can16358p 5y agoHmmm. Some quick observation: on especially large 9-ish digit numbers, when it make very few digits correct, the correct ones are mostly including the very first and very last digits. Something remarkably similar to how us humans remember numbers and words: when we make mistake we generally remember first and last digits/letters but mess up the middle.
- wildmanx 5y agoI'd like to see doing this with random people on the street and then compare performance. You may be surprised.
- NateEag 5y agoNot really. Average person on the street is going to correctly say "geez, I dunno. Can I use my phone?" If you don't forbid them to, then they'll whip it out and get 98% correct (I figure they'll typo a few). This model didn't have enough understanding to do that (since it literally has no understanding at all).
- wildmanx 5y agoI don't know where you live, but 98% correct is not what would happen around here. Edit: Oh, 98% _with_ a calculator. What if you force them to do it by hand?
- ceejayoz 5y ago> What if you force them to do it by hand? Why? GPT-3 isn't doing it that way. The importance is understanding the question, not being able to do math in your head.
- paraschopra 5y agoI tried putting numbers as words and it did additions perfectly. Pretty magical! What is fifty plus ninety? 140 What is fifty plus ninety one? 141 What is fifty minus ninety one? -41 What is minus fifty minus ninety one? -141 Although it failed in multiplication or adding longer numbers (as words).
- mordymoop 5y agoWhen you toss “2241 + 19873 =” into an applet that shows you the default tokenization scheme GPT-3 uses, you get this: (224)(1)( +)( 198)(73)( =) I’ve heard it remarked before that, while tokenization is obviously an unavoidable part of a model with an architecture like GPT, this is a very silly way of tokenizing number strings for the purposes of learning or doing arithmetic. Indeed, I think a lot of GPT-3’s puzzling edge-case performance can be ascribed to weird and unhelpful tokenizations. Just imagine if you were forced to learn arithmetic with a brain that automatically categorized “224” as a sort of distinct object, or, for that matter, breaking down 19873 as ( 198)(73) rather than (19873) or (1)(9)(8)(7)(3) or anything practically useful. The thing is that we can, in a sense, learn better “tokenizations”, in the sense that a 4 year old learning to read sees letters, while a 40 year old reading a novel “sees” whole words or even groups of words. The GPT architecture can’t change its tokenization scheme.
- jcims 5y agoWhoa, that explains why only .5% of the examples have an incorrect last digit.
- starfallg 5y agoIt seems that we need another layer to tokenize according to context. I can see that breaking up a long number into 3 or 4 digits is the correct behaviour if we are dealing with phone numbers, but it'd be completely wrong if it's nearly anything else.
- vbuterin 5y agoWhen I do mental arithmetic my brain frequently tokenizes into digit pairs or triples if I can recognize pairs and triples that have specific properties. "224" is actually a really nice object to recognize because it's 7 * 32, and if you can recognize other multiples of 32 it frequently gives you shortcuts. It's less useful for addition because you would need to get lucky and get a multiple of 32 (or 7) on both sides, but for multiplication and division it helps a lot.
- dr_zoidberg 5y ago
- jcims 5y agoThese language models feel, to me, like the unfiltered self. If someone asked me what 838+1283 was my head would instantly offer up some number 2301 or something. But i would discard that number because I learned in elementary school that I don't come up with good values, I need to execute a process in order to get the right value. I imported the csv version and I'm no statistician but 90% percentile relative error is 8.6%, which is something like this: What is 22730 - 24978? -2448 (real answer -2248) That's totally within range of something that would plop into my head...with one exception. Of 1000 entries, only five have an incorrect last digit. I think that's meaningful...it almost tells me that there's a multi-stage operation happening in there somewhere.
- visarga 5y ago> These language models feel, to me, like the unfiltered self. A generator-critic framework with multiple rounds of iteration would improve on the limitations of the LM.
- mlb_hn 5y agoCouple things there where you can see if it improves with the prompt/formatting. E.g. with Davinci (and J a bit but didn't test too much) you can get bette results by: - Using few-shot examples of similar length to the targets (e.g. 10 digit math, use 10 digit few shots) - Chunking numbers with commas - Having it double check itself and here it's not doing any of those things.
- ur-whale 5y agoYup, glorified, extremely large, lookup table, but has zero clues of how the world works. Or to put it another way: if I haven't seen the result before I'm very likely to get it wrong. And inferring the rules of arithmetic from examples : nope. We're all very impressed with the large language models because human language was always very hard to get a handle one. Now that we've made progress in manipulating natural language, we are very quick to believe that the darn thing is somehow "intelligent" because we've always attached some sort o mystical quality to language. This demonstrates that the capacity of these systems to reason and craft high order rules from enough examples is exactly nil.
- xkapastel 5y agoCan someone change the title? This is GPT-NeoX-20B, not GPT-3.
- tehsauce 5y agoTitle says gpt-3, but gist says gpt-neo-20b which was recently open sourced but only 1/10 the size of gpt3.
- tehsauce 5y agoThe latest version of gpt-3 spooked me the other day with: Me: Can here you here still here understand here what here I here am here asking here if here I here put here words here in here between? Gpt: Yes, I can understand what you are asking if you put words in between. Me: What was the word that I put in between? Gpt: The word that you put in between is "here."
- spiderfarmer 5y agoThat's shockingly good.
- marcodiego 5y agoScary. If it improves a bit more, people will start questioning if the machine has soul or rights.
- gitfan86 5y agoMy kids already debate if it is wrong to tell OK Google to shutup.
- abinmn 5y agoyou have got a great bunch there :)
- supermdguy 5y agoIt's interesting, in the forums for the beta program there have been already been a few people making posts where they're convinced that the AI is conscious. That's never really been something I've thought about much since I know a little about how it works, but I could totally see how someone who didn't have as much context for how GPT-3 works could see it as some sort of sentience. https://community.openai.com/t/a-conversation-with-alec-a-conscious-ai/8209 https://community.openai.com/t/a-conversation-with-alec-a-co... https://community.openai.com/t/creepy-ai-behavior/10195 https://community.openai.com/t/creepy-ai-behavior/10195 https://community.openai.com/t/where-to-watch-what-the-ai-wants-to-show-us/8462 https://community.openai.com/t/where-to-watch-what-the-ai-wa...
- marcodiego 5y agoI'm still impressed that for the first 1000 tests it gets close most of the time.
- marcodiego 5y agoI fear the day AI will give superhuman consistent correct answers and nobody will be able to determine why it is right or how the correct answer was found. Maybe someday we'll get an answer from a machine which superhumanly mostly correct and we'll be unable to tell if it is right or wrong. If it is a question whose answer will influence important decisions, considering the machine answer will be close to a form of religion.
- jjoonathan 5y agoLike religion, I suspect you will have many different machine answers to choose from.
- olyjohn 5y agoI suspect we will have 2 of them. We will start out with lots of them, but then a couple of them will start making the most money, and resort to underhanded tactics, bribery, and lobbying, and put the others out of business.
- jjoonathan 5y agoSounds about right. We'll replace our lizard overlords with robot overlords.
- deleted 5y ago[deleted]
- nlh 5y agoI discovered something like this in real-world usage. I’ve have GitHub Copilot running in VSCode and I’ve been experimenting with how it works when doing plain text accounting (using ledger / hledger). The ledger files are somewhat “code”-like so it’s been super interesting to see how it works. The short answer: it works really quite well! ..except for the math part :) I have a long ledger of transactions, and I can now give Copilot a comment like: “Jan 1, 2022 +100 from consulting income” and it (GPT-3) will generate a nearly perfect ledger entry, debiting from income and crediting the right bank account. But the arithmetic is always wrong (ledger has an option for you to keep a running balance as a check). There’s the occasional moment where it gets the balance adjustment correct, but almost every time the results are similar to this post.
- supermdguy 5y agoI also have Copilot running, and I was surprised when it had pretty good autosuggestions when writing proofs in Latex! A lot of times it has subtle logical errors in the proof, but the syntax is always correct. And there have been a few times when it gives a sentence or two that's exactly right
- musingsole 5y agoThe page lists the average percent error, but I was interested in the median percent error: 1.20%
- rexreed 5y agoAmazing how a very costly to train system using billions of neural nodes on millions of dollars of compute performs more poorly than an 8-bit 1970s pocket calculator. Not sure why people are expecting some sort of "intelligence" to emerge from a text generator model trained on Internet corpus data. GPT-3 doesn't calculate, it pattern matches. I do get why people might be surprised, on the other hand, that it actually doesn't perform worse than indicated here. Maybe it's surprising upside. But since we know that the GPT is a transformer model, what it is doing is applying a probabilistic best-fit. From this perspective I can see how it is best-fitting data in ways that can provide these sorts of results, especially given all that training data.
- andreyk 5y agoWhat's your definition of 'intelligence'? Many of the things GPT-3 does clearly exhibit intelligence (just not human level intelligence).
- deleted 5y ago[deleted]
- Samin100 5y agoAccording to Kahneman, when a chess pro makes an “intuitive” and unexplainable move it’s just pattern recognition happening subconsciously. If language models like GPT-3 are “just” pattern recognizers, wouldn’t that makes them capable of intuition?
- MaxMoney 5y ago
- habitue 5y ago> what it is doing is applying a probabilistic best-fit I think you're underselling probabalistic best-fits. Especially with all of the regularization going on in training.
- a-dub 5y agois there a training data browser with a search engine somewhere for gpt-3?
- deleted 5y ago[deleted]
- moyix 5y agoHey! As the author of the gist, just wanted to clear up what seem to be a few misconceptions: - This isn't GPT-3, it's the recently-released open-source and open-weights model from EleutherAI, GPT-NeoX-20B. GPT-3 is much larger (175 billion parameters vs NeoX's 20 billion). - It's well-known that language models don't tend to be good at math by default (Gwern, among others, pointed this out back in June 2020). It seems likely that this is at least in part because of how these models currently tokenize their input (they don't represent numbers by their individual digits, but by tokens representing commonly-occurring character sequences): https://www.gwern.net/GPT-3#bpes https://www.gwern.net/GPT-3#bpes . Someone also pointed me to this paper which looks at number representations (though it uses somewhat older models like BERT): https://arxiv.org/abs/1909.07940 https://arxiv.org/abs/1909.07940 - Despite the tokenization, it performs (IMO) surprisingly well at getting close to the true value, particularly for the start and end digits and the overall magnitude. You can see this by looking at the tokenization (indicated by brackets) of its guess vs the correct answer for 28531*8065 (I asked multiple times to get an idea of how consistent it is – it's not deterministic because I ran this with temperature = 0.1, which will use random sampling to get the most likely tokens): [What][ is][ 285][31][ *][ 80][65][?][\n][22][77][05][315] Correct: [\n][23][010][25][15] [What][ is][ 285][31][ *][ 80][65][?][\n][22][95][01][115] Correct: [\n][23][010][25][15] [What][ is][ 285][31][ *][ 80][65][?][\n][22][38][95][015] Correct: [\n][23][010][25][15] [What][ is][ 285][31][ *][ 80][65][?][\n][22][99][25][015] Correct: [\n][23][010][25][15] [What][ is][ 285][31][ *][ 80][65][?][\n][22][99][17][115] Correct: [\n][23][010][25][15] You can see that it manages to find things that are numerically close, even when no individual token is actually correct. And it compensates for different-length tokens, always picking tokens that end up with the correct total number of digits. - Please don't use this as a calculator :) The goal in doing this was to figure out what it knows about arithmetic and see if I can understand what algorithms it might have invented for doing arithmetic, not to show that it's good or bad at math (we have calculators for that, they work fine).
- mlb_hn 5y agoI get the tokenization argument and it may influence it a bit, but I suspect the n-digit math issue has to do more with search the way it samples (in the bpe link gwern references some experiements I'd done with improving n-digit math by chunking using commas, http://gptprompts.wikidot.com/logic:math http://gptprompts.wikidot.com/logic:math). I think since it samples left to right on the first pass, it's not able to predict well if things carry from right to left. I think can mitigate the search issue a bit if you have the prompt double-check itself after the fact (e.g. https://towardsdatascience.com/1-1-3-wait-no-1-1-2-how-to-have-gpt-sanity-check-itself-136e846987bf https://towardsdatascience.com/1-1-3-wait-no-1-1-2-how-to-ha...). Works different depending on the size of the model tho.
- baalimago 5y agoI'll be impressed when the AI consults with an ordinary calculator for the correct answer
- edouard-harris 5y agoThis already exists: Google's recently-published LaMDA dialogue model [1] is trained to consult a calculator for arithmetic questions and consistently succeeds at it. [1] https://arxiv.org/abs/2201.08239v2 https://arxiv.org/abs/2201.08239v2
- dvh 5y agoMotorola: How. Much. Is. One. Plus. One? Pentium: 3 Motorola: That. Is. Not. Correct. Pentium: no, but it's fast!
- moffkalast 5y agoQuik mafs
- svilen_dobrev 5y agohow about hex? 0xAA + 0x23 ? or, if it has "seen" low-level driver code, 0xAA & 0x23 ?
- deleted 5y ago[deleted]
- Veedrac 5y agoI fiddled with this prompt Calculate 721*436 Step A: Multiplication table. 721*0=0 721*1=721 721*2=1,442 721*3=2,163 721*4=2,884 721*5=3,605 721*6=4,326 721*7=5,047 721*8=5,768 721*9=6,489 Step B: Multiply digits 436=4 hundreds + 3 tens + 6 ones 721*4 hundreds=2,884*100=288,400 721*3 tens=2,163*10=21,630 721*6 ones=4,326*1=4,326 Step C: Add the first two 288,400+21,630=310,030 Then add the last one 310,030+4,326=314356 To get GPT-3's answer to this one correct (this is its answer, but it took some trial and error with the prompt to get it to stop making small mistakes) Calculate 312*481 Step A: Multiplication table. 312*0=0 312*1=312 312*2=624 312*3=936 312*4=1,248 312*5=1,560 312*6=1,872 312*7=2,184 312*8=2,496 312*9=2,808 Step B: Multiply digits 481=4 hundreds + 8 tens + 1 ones 312*4 hundreds=1,248*100=124,800 312*8 tens=2,496*10=24,960 312*1 ones=312*1=312 Step C: Add the first two 124,800+24,960=149,760 Then add the last one 149,760+312=150,072 I then tested it with two random 3-digit numbers Calculate 669*306 Step A: Multiplication table. 669*0=0 669*1=669 669*2=1,338 669*3=2,007 669*4=2,676 669*5=3,345 669*6=4,014 669*7=4,683 669*8=5,352 669*9=6,021 Step B: Multiply digits 306=3 hundreds + 0 tens + 6 ones 669*3 hundreds=2,007*100=200,700 669*0 tens=0*10=0 669*6 ones=4,014*1=4,014 Step C: Add the first two 200,700+4,014=204,714 Then add the last one 204,714+306=205,020 204,714 is correct, but then it gets confused in the very last line because of the zero. Anyway this, I think, proves that GPT-3 can do zero-shot math if it doesn't have to do it in its head all at once. Is it good for a calculator? No. Is it good for a language model trained on web text? Yes, absolutely. E: Correcting the start of Step C in that last example to “200,700+0”, and replacing “XYZ=X hundreds + Y tens + Z ones” with “XYZ=X Y Z=X hundreds + Y tens + Z ones” allowed it to do 145*585, 961*761 and 592*555 correctly in a row, all randomly chosen, and at least the last two tried without changes to the prompt. I consider this an adequate test, and it demonstrates GPT-3's algorithm following abilities. As GPT-3 is still a tiny model, this seems important to note. E2: To be clear this is still nowhere near 100% successful. GPT-3 still makes a lot of errors. I ran 100 tries of a slightly different prompt through the API, and got a success rate of 42%.
- infogulch 5y ago> can do zero-shot math if it doesn't have to do it in its head all at once Very interesting! This is what I would expect. It can run a symbolic algorithm fine, just give it some scratch space to work out the intermediate results. I feel like there's a very large space to optimize the layout "algorithm" -- like how you adjusted step c -- to produce reliable results.