4 ms·
But the example the poster you are replying to gave shows that GPT-3 can take the square root of a user-defined function composed with itself. It clearly can pe
by dddbbb 6y ago
But the example the poster you are replying to gave shows that GPT-3 can take the square root of a user-defined function composed with itself. It clearly can perform arithmetic, and we don't need to trust OpenAI now that users can interact with the model.
- YeGoblynQueenne 6y agoAnd yet it can't always correctly subtract two two-digit numbers. Does that sound like a system that can perform arithmetic? As I say in previous comments, no. It sounds much more like a system that can reproduce resutls it has seen in trainig, but has no general concept of arithmetic. This would also explain the square root example example easily. Also, the examples in the OP's linked tweet are very simple examples of square roots and function composition that are very likely to have been lifted verbatim from some textbook, or who knows what ...and that's the problem, because who knows what the model has flat memorised and what it's composing from smaller components. >> It clearly can perform arithmetic, and we don't need to trust OpenAI now that users can interact with the model. The paper I link above performed a systematic evaluation of GPT-3's arithmetic ability. Playing around with the OpenAI API and eyballing a few results is not going to give a clearer understaning of its abilities. In general, hitting a language model with a few queries is never going to give any clear understanding of its capabilities. Systematic evaluation is always necessary and the average user (or the non-average user) is not going to be able to do that.
- loxali 6y agoI think it's pretty clear that there is something more than memorising examples from the training set going on. Look at this: https://www.johnfaben.com/blog/gpt-3-arithmetic https://www.johnfaben.com/blog/gpt-3-arithmetic In which GPT-3 answers the question "what is one hundred and five divided by three?" with "35.7". It also gave several other close-but-not-correct answers. It seems pretty unlikely these are all present in the training set, and surely can't all have been lifted verbatim. I agree systematic testing is probably more useful, but find it really hard to believe this is all happening without any sort of model of arithmetic.