4 ms·
So this may mean that language models like GPT-3 are not best suited for code generation tasks. And networks trained on languages may not learn the feature embe
by nutanc 6y ago
So this may mean that language models like GPT-3 are not best suited for code generation tasks. And networks trained on languages may not learn the feature embeddings needed for programming languages.
- magusdei 6y agoWouldn't the empirical success of GPT-3 in simple programming tasks itself be evidence against this interpretation? Furthermore, GPT-3 is only a language model because it is trained on textual data. Transformer architectures simply map sequences to other sequences. It doesn't particularly matter what those sequences represent. GPT-2 has been used to complete images, for example: https://openai.com/blog/image-gpt/ https://openai.com/blog/image-gpt/
- nutanc 6y agoEmpirical success shows that the GPT-3 model has seen the sequence before(maybe many times). Transformer architectures do map sequences to sequences. What is not known is that the task of programming is a sequence problem. This experiment seems to suggest that maybe its not a sequence problem.