3 ms·
The data was collected for a paper that was attempting to find a lexical difference between human contributions and LLM contributions. The reason for this was b
by tay1or 17d ago
The data was collected for a paper that was attempting to find a lexical difference between human contributions and LLM contributions. The reason for this was because repository maintainers need to make sure submitted code has been at least reviewed by a human. By detecting unattributed AI submissions, it can tell maintainers that they need to pay extra attention to the repositories.
- prologic 17d agoAnd does that work? Can you genuinely tell whether a human has reviewed Ai-assisted/generated code?
- tay1or 17d agoYou can't tell if a human has reviewed AI-assisted / generated code. But you can tell if an AI wrote a comment or commit message. The evaluation section of the paper goes over in-depth the findings: https://github.com/ActuallyTaylor/strata/blob/main/paper/Taylor%20Lineman%20-%20Changes%20in%20Vocabulary%20Between%20Human-created%20and%20LLM-assisted%20Code%20Repositories.pdf https://github.com/ActuallyTaylor/strata/blob/main/paper/Tay...