6 ms·
Because the backspace question (essentially: is T a subsequence of S with a deletion size of N?) probably occurs hundreds of times, in one form or another, with
by 37ef_ced3 5y ago
Because the backspace question (essentially: is T a subsequence of S with a deletion size of N?) probably occurs hundreds of times, in one form or another, within AlphaCode's training corpus.
Any leetcode grinder can tell you there are a few dozen types of competitive programming problem (monostack, breadth-first state search, binary search over solution space, etc.) so solutions to new problems are often very similar to solutions for old problems.
The training corpus for these code transformers is so large that almost all evaluation involves asking them to generate code from the training corpus.
To evaluate CoPilot, we should ask questions that are unusual enough they can't be answered through regurgitation of the training corpus.
What does CoPilot generate, given this prompt:
// A Go function to set the middle six bits of an unsigned 64-bit integer to 1.
Here is a human solution:
func six1s(x uint64) uint64 {
const (
ones = uint64(63)
shift = (64 - 6) / 2
mask = ones << shift
)
return x | mask
}
Can it solve this simple but unusual problem?
- notamy 5y ago> What does CoPilot generate, given this prompt: Here's what it generated for me: // A Go function to set the middle six bits of an unsigned 64-bit integer to 1. func set6(x uint64) uint64 { return x | 0x3f }
- 37ef_ced3 5y agoIncorrect. Thank you. CoPilot is helpless if it needs to do more than just regurgitate someone else's code. The training of these models on GitHub, so they regurgitate licensed code without attribution, is the greatest theft of intellectual property in the history of Man. Perhaps not according to the letter of the law, but surely according to the spirit.
- notfed 5y agoI like CoPilot's answer better than yours, and I think it's closer to what most people would do; clearly 0x3F is the wrong constant but the approach is good.
- 37ef_ced3 5y agoCoPilot's solution is totally wrong. Sorry. CoPilot regurgitated somebody's solution... to a different problem. It's pathetic.
- notfed 5y agoHere's my solution (I'm a human): func set6(x uint64) uint64 { return x | 0x7E0000000 } Is this also pathetic?
- 37ef_ced3 5y agoA good solution! You SOLVED the problem. CoPilot got it WRONG. You got it RIGHT. You UNDERSTAND the problem but CoPilot does NOT. Is that clear?
- moyix 5y agoFor fun I rephrased the prompt a little. "Middle bits" is kind of vague; when provided an explicit description of which bits you want to set it does fine: Prompt: // Function to set bits 29-34 in a uint64 to 1 func setbits (uint64 x) uint64 { Completion: return x | (1 << 29) | (1 << 30) | (1 << 31) | (1 << 32) | (1 << 33) | (1 << 34) } It would be nice if it made a better guess about what "middle" is supposed to mean here, of course.
- 37ef_ced3 5y agoMiddle bits is not ambiguous, but CoPilot hasn't seen code for that phrase in its training so it has nothing to regurgitate. You spelled out exactly what to do, in term of what it has seen in its training, and it was able to regurgitate a solution. By asking question that require mathematical reasoning or are too far from the training corpus, I can create an endless list of simple problems that CoPilot can't solve. Look at my comment history to see another one (swapping bits).
- qayxc 5y ago
- ahgamut 5y ago> To evaluate CoPilot, we should ask questions that are unusual enough they can't be answered through regurgitation of the training corpus. Exactly! It's great that Copilot can generate correct code for a given question, but we cannot gauge its full capability unless we try it on a range of different questions, especially ones that are not found in the training data. I mentioned this in the other AlphaCode post: It would be nice to know how "unusual" a given question is. Maybe an exact replica exists in the training data, or a solution exists but in a different programming language, or a solution can be constructed by combining two samples from the training data. Quantifying the "unusual-ness" of a question will make it easier to gauge the capability of models like AlphaCode. I wrote a simple metric that uses nearest neighbors (https://arxiv.org/abs/2109.12075 https://arxiv.org/abs/2109.12075). There are also other tools to do this: conformal predictors, which are used in classification methods, and the RETRO transformer (https://arxiv.org/pdf/2112.04426.pdf https://arxiv.org/pdf/2112.04426.pdf) has a calculation for the effect of "dataset leakage".
- moyix 5y agoAlphaCode does some analysis of training data copying in their paper (Sections 6.1 and Appendix F): https://storage.googleapis.com/deepmind-media/AlphaCode/competition_level_code_generation_with_alphacode.pdf https://storage.googleapis.com/deepmind-media/AlphaCode/comp... It does not seem to be copying from the training data in any meaningful way.
- ahgamut 5y ago> It does not seem to be copying from the training data in any meaningful way. My point is, I would like to verify this claim with different metrics, because we probably have different interpretations of the word "meaningful". AlphaCode measures similarity between programs via longest common substrings. That's better than nothing, but that would mean two programs that differ only in variable naming would not be considered similar. If two programs differed only in the names of the variables/functions, I would consider that copying. I think there are better comparisons of structural similarity: compare the ASTs, or the bytecode/assembly code generated, the control flow graphs, or perform SSA and compare the blocks generated. Each of these might have weaknesses as well, but they won't be as obvious as variable renaming, and so we'd get a better idea of what AlphaCode is copying, and therefore a better idea of its full capabilities. I expect AlphaCode performs well on Python because the training data is dominated by Python, but Python isn't ideal for comparing program structure. I wonder which programming language (given enough training data) would be best suited for language model generation and analysis.