3 ms·
I am not educated on this matter, but I have to ask for your clarification. Would that not be just a pre-emptive lookup? Akin to keeping a cache of "results" pe
by throwaway413 5y ago
I am not educated on this matter, but I have to ask for your clarification. Would that not be just a pre-emptive lookup? Akin to keeping a cache of "results" per input token that are essentially memorized and regurgitated?
Sounds like there is still a db lookup, just not at runtime and instead at build time of the NN. Can you clarify this please?
- freeone3000 5y agoGPT-3's raw output is "logits", or indexes into an encoding space. The encoding space contains individual tokens; for generation, it would be words, or even word pieces. The pieces are as small as "for", or "if". Constructing code from an embedding space, even if it is more specialized, is like constructing sentences by using a dictionary -- it is a lookup table, but it's not a database. Generation works by looking at the existing document (or portion), and based on what is already present, generating a token. Then repeating until some condition is met (such as length, end of sentence, something else). The issue here is that certain sentences (code segments) are memorized, and reproduced -- much like a language learner who completes every sentence which begins with "Mi nombre" with the phrase "Mi nombre es Mark". The regurgitation is based on high probability built into the priors, not an explicit lookup. A different logit sampling method (instead of taking the likeliest) reduces regurgitation, without changing anything else about the network. (It also makes nonsense happen more often, since nonsense items are inherently less likely!)
- nl 5y agoThis response is correct.