3 ms·
I don't think it's possible with GPT-3, and that's mostly due to how the text is parsed into token before being fed to the network [1]. It breaks down the text
by nowahe 5y ago
I don't think it's possible with GPT-3, and that's mostly due to how the text is parsed into token before being fed to the network [1]. It breaks down the text in ~4 words token, which allows to effectively quadruple the max input size, at the cost of loosing fine details on the input data. It leads to issues like not being able to create rhymes, not understanding humor or not being able to parse fine structures. Gwern has a nice article talking about the limitations introduced by it [2].
[1] https://beta.openai.com/docs/introduction/tokens https://beta.openai.com/docs/introduction/tokens
[2] https://www.gwern.net/GPT-3#bpes https://www.gwern.net/GPT-3#bpes
- gwern 5y agoInterestingly, Codex/Copilot might have even more extreme BPE issues than GPT-3 does. They mention that they manage to expand the window in terms of characters greatly (to 4096?), which I take as implying they recomputed the BPEs on the source corpus to make it much more tailored and relevant to source code (which makes sense, because I would expect source code to be far more verbose and repetitive, in terms of syntax and vocab, than Internet-wide natural language, and so simply using the old GPT-2/GPT-3 BPEs won't work well). Which is fantastic if you're doing Python, and that wide window is part of how they do tricks like being able to put in parts of debugging sessions, but if you are going to use Copilot for non-source-code generation, I wonder what the consequences might be...? BPE issues can be hard to notice even when you are looking for them and have the vocab at hand to check.