7 ms·
Thank you for this. Technically it's not GPT-3, but GPT-NeoX-20B, although they are based on a similar architecture. The poor performance is most likely due to
by spupe 5y ago
Thank you for this. Technically it's not GPT-3, but GPT-NeoX-20B, although they are based on a similar architecture.
The poor performance is most likely due to not having a large database of math problems to draw from. Github, for example, is part of the dataset that is used to train both GPT-3 and GPT-Neo variants, which is partly why they can generate meaningful code (sometimes). I wonder how a model finetuned for math would perform.
- dang 5y agoOk, we've reverted the title now. Thanks! (Submitted title was 'GPT-3's answers to arithmetic questions')
- spupe 5y agoI went and checked, it turns out for this version Eleuther-AI has in fact included math problems [1]. So my earlier comment is partly incorrect. [1] http://eaidata.bmk.sh/data/GPT_NeoX_20B.pdf http://eaidata.bmk.sh/data/GPT_NeoX_20B.pdf
- asah 5y agoAnd isn't it trivial to generate lots of correct sample data ? :-)
- williamtrask 5y agoPoor performance is more likely due to how transformer neural networks view numbers. It memorises them like words instead of modeling their numerical structure. Thus even if it’s seen the number 3456 and 3458, it knows nothing of 3457. Totally different embedding. It’s like a kid memorising a multiplication table instead of learning the more general principle of multiplication (related: this illusion is why big models are so popular. Memorise more stuff.) Paper (NeurIPS/DeepMind): https://arxiv.org/abs/1808.00508 https://arxiv.org/abs/1808.00508
- eutectic 5y agoThat depends on the tokenization scheme.
- plutonorm 5y agoIt's recently been shown that even though the numbers are represented with different tokens, the network learns to form an internal representation that understands the progression from one token to the next.
- nikolayasdf123 5y agoThe idea that each number has to be inside ones Brain or Neural Network or Token is plainly wrong. Network has to grasp the "abstract" number, but it clearly did not grasp that concept.
- Isinlor 5y agoHow would you test if it grasped the concept?
- catach 5y agoBeing able to extrapolate to numbers that were not in the training set, perhaps? At least that'd be a basic part of the requirement.
- Isinlor 5y agoSure: Deep Symbolic Regression for Recurrent Sequences https://arxiv.org/abs/2201.04600 https://arxiv.org/abs/2201.04600 (Interactive demo: http://recur-env.eba-rm3fchmn.us-east-2.elasticbeanstalk.com/ http://recur-env.eba-rm3fchmn.us-east-2.elasticbeanstalk.com... ) Grokking: Generalization Beyond Overfitting on Small Algorithmic Datasets https://arxiv.org/abs/2201.02177 https://arxiv.org/abs/2201.02177 Both of these models can generalize to numbers it have not seen.
- throwaway4good 5y agoNo. The poor performance comes from the overall approach of using neural nets to solve basic math problems.
- FL410 5y agoThe cool part comes when the model can make the connection that multiply 12345 by 87654 is the same as def multiply_two_numbers(x, y): return x * y Which of course produces the desired result. The interesting part is that github copilot wrote the above with only the prompt "def multiply_two" as the prompt.
- andreyk 5y agoOh, that's a pretty big difference, would be nice if post title was altered...
- deleted 5y ago[deleted]