3 ms·
Interesting concept, now to rush and lookup word2vec applications and see if they make sense for code :p Also, props for an academic work having an extensive R
by gcommer 8y ago
Interesting concept, now to rush and lookup word2vec applications and see if they make sense for code :p
Also, props for an academic work having an extensive README! https://github.com/tech-srl/code2vec https://github.com/tech-srl/code2vec
I do wonder if some sort of AST normalization would improve the input signal. Example 8 on the website shows their system correctly identifying an isPrime function. However some irrelevant perturbations can break it. If you swap the if statement condition around from `n % i == 0` to `0 == n % i` the proposed names are totally different and make no sense.
- namibj 8y agoWell, than maybe train it on random AST equivalency transformations? I'd assume that to work better.
- yahave 8y agoYes, that indeed works better
- faizshah 8y agoI wonder if this could be used to help detect plagiarism when auto grading CS homework.
- toolslive 8y agoMaybe, but in the past, we've detected plagiarism in a different way: just look at the assembly output of the compiled code. The compiler is a good normalizer that removes artificial differences like naming of variables and functions.
- ychen306 8y agoAssuming I understand their approach correctly, they are using little (one could argue none at all) semantic information.
- Karlozkiller 8y agoIf the task you want to solve is automatic function naming I definitely think normalizing would be an improvement. But I'm not sure it would be the right thing for all applications. Don't have any examples though.